Resource configuration method and system, electronic equipment and storage medium

By dynamically configuring general servers and graphics card resource pools and utilizing graphics processors that support online plug-in and unplugging, the problems of low resource utilization and poor overall machine flexibility in traditional heterogeneous instances are solved, achieving efficient resource utilization and flexible configuration.

CN120723322APending Publication Date: 2025-09-30HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410361498.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Traditional heterogeneous instances have high GPU failure rates and long maintenance cycles for complete machine failures. Furthermore, they are limited by the fixed CPU and GPU configurations in the complete machine configuration, making it difficult to meet diverse CPU and GPU resource ratio requirements, resulting in resource fragmentation and low utilization.

Method used

By obtaining the configuration requirement information of the computing power service, resources are configured for general servers and graphics card resource pools. By using multiple graphics processors that support online plug-in and unplug functions, CPU and GPU resources are dynamically adjusted to provide flexible computing instances.

Benefits of technology

It improves resource utilization and overall machine flexibility in elastic computing scenarios, enables flexible configuration of processor resources, and solves the problems of low resource utilization and poor overall machine flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723322A_ABST
    Figure CN120723322A_ABST
Patent Text Reader

Abstract

The invention discloses a resource configuration method and system, electronic equipment and a storage medium. The method comprises the following steps: acquiring configuration demand information corresponding to computing power service; according to the configuration demand information, resource configuration is carried out on a universal server and a graphics card resource pool, a configuration result is obtained, the universal server is obtained based on processor resource configuration, and the graphics card resource pool comprises a plurality of graphics processors supporting an online plugging function; and providing a first instance for the computing power service based on the configuration result, the first instance comprising processor resources allocated by the general server and processor resources allocated by the graphics card resource pool. According to the method and the device, the technical problems of low resource utilization rate and poor whole machine flexibility caused by the fact that heterogeneous machine types are limited by fixed configuration of whole machine processor resources in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and more specifically, to a resource configuration method, system, electronic device, and storage medium. Background Art

[0002] With the continuous development of cloud computing technology, heterogeneous instances are becoming increasingly popular in computing power service products. However, the graphics processing units (GPUs) used in traditional heterogeneous instances have high failure rates and long maintenance cycles for complete machine failures. Furthermore, traditional heterogeneous instances are limited by the fixed configuration of CPUs and GPUs within the overall machine configuration, making it difficult to meet the diverse CPU and GPU resource allocation requirements in today's complex cloud scenarios. Consequently, traditional heterogeneous instances correspond to a wide variety of heterogeneous machine models, resulting in significant resource fragmentation and low heterogeneous resource utilization.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a resource configuration method, system, electronic device and storage medium to at least solve the technical problems in the related art of low resource utilization and poor flexibility of the entire machine due to the fact that heterogeneous models are limited by the fixed configuration of the processor resources of the entire machine.

[0005] According to one aspect of an embodiment of the present application, a resource configuration method is provided, including: obtaining configuration requirement information corresponding to a computing power service; performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on a processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0006] According to another aspect of an embodiment of the present application, a resource configuration method is also provided, including: obtaining a resource configuration request through a first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; returning a resource configuration response through a second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of a first instance provided for the computing power service, the first instance is determined based on a configuration result, and the configuration result is obtained by performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0007] According to another aspect of an embodiment of the present application, a resource configuration method is also provided, including: obtaining a currently input resource configuration request, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; in response to the resource configuration request, returning a resource configuration reply, wherein the information carried in the resource configuration reply includes: instance information of a first instance provided for the computing power service, the first instance is determined based on a configuration result, and the configuration result is obtained by performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool; and displaying the instance information in a graphical user interface.

[0008] According to another aspect of an embodiment of the present application, a resource configuration method is provided, comprising: obtaining configuration requirement information corresponding to a computing power service; using the configuration requirement information and a resource configuration model to perform resource configuration on a general server and a graphics card resource pool to obtain a configuration result, wherein the resource configuration model is a neural network model pre-trained by machine learning using multiple sets of training data, the general server is configured based on processor resources, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplugging functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool. According to another aspect of an embodiment of the present application, a resource configuration system is provided, comprising: a general server, configured based on processor resources, memory resources, and service manager resources; a graphics card resource pool, including multiple graphics processors that support online plug-in and unplugging functions; a graphics card interface box, connected to the general server and the graphics card resource pool, for performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information corresponding to the computing power service to obtain a configuration result, wherein the configuration result is used to support the general server in providing a first instance for the computing power service, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0009] According to another aspect of an embodiment of the present application, an electronic device is further provided, including: a memory storing an executable program; and a processor for running the program, wherein any one of the above-mentioned resource configuration methods is executed when the program is running.

[0010] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned resource configuration methods.

[0011] According to another aspect of an embodiment of the present application, a computer program product is provided, including a computer program, which implements any of the above resource configuration methods when executed by a processor.

[0012] In an embodiment of the present application, by obtaining configuration requirement information corresponding to a computing power service; further configuring resources for a general-purpose server and a graphics card resource pool according to the configuration requirement information, a configuration result is obtained, wherein the general-purpose server is obtained based on processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; based on the configuration result, a first instance is provided for the computing power service, wherein the first instance includes processor resources allocated by the general-purpose server and processor resources allocated by the graphics card resource pool. Thus, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall flexibility of heterogeneous models in elastic computing scenarios, thereby solving the technical problem of low resource utilization and poor overall flexibility of the related art due to the fact that heterogeneous models are limited by the fixed configuration of the processor resources of the overall machine.

[0013] It is easy to notice that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0015] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a resource configuration method is shown;

[0016] Figure 2 is a flow chart of a resource configuration method according to Example 1 of the present application;

[0017] Figure 3 is a schematic diagram of a computing node in an elastic computing system according to the prior art;

[0018] Figure 4 is a schematic diagram of an optional elastic computing system according to Example 1 of the present application;

[0019] Figure 5 is a flow chart of a resource configuration method according to Example 2 of the present application;

[0020] Figure 6 is a flowchart of a resource configuration method according to Example 3 of the present application;

[0021] Figure 7is a flowchart of a resource configuration method according to Example 4 of the present application;

[0022] Figure 8 is a structural diagram of a resource configuration system according to Example 5 of the present application;

[0023] Figure 9 is a structural diagram of a resource configuration device according to Example 6 of the present application;

[0024] Figure 10 is a structural diagram of another resource configuration device according to Example 6 of the present application;

[0025] Figure 11 is a structural diagram of another resource configuration device according to Example 6 of the present application;

[0026] Figure 12 This is a structural block diagram of an electronic device according to Example 7 of the present application. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:

[0030] A graphics processing unit box (GPU BOX) typically contains multiple GPU cards. Connecting a GPU BOX to a computer provides graphics processing capabilities, for example, accelerating graphics processing, data processing, or deep learning tasks. The GPU BOX can be connected to one or more general-purpose computing servers as a head unit via high-speed data transmission interface cables (such as PCIe cables) to provide GPU instances.

[0031] GPU instance: refers to a virtual server instance provided in a cloud computing environment, which is equipped with a graphics processing unit (GPU) specifically designed to accelerate graphics processing and scientific computing tasks. GPU instances are typically used for workloads that require a large amount of parallel computing, such as machine learning, deep learning, data analysis, scientific computing and other fields. Because GPUs have significant advantages over traditional central processing units (CPUs) in parallel computing, GPUs can provide higher performance and efficiency when processing workloads in the above fields. Therefore, many cloud computing service providers provide GPU instances as part of their cloud computing services to meet users' needs for high-performance computing and parallel computing.

[0032] A Peripheral Component Interconnect Express Switch (PCIE SWITCH) is a device that provides a high-speed data transmission interface connecting various hardware devices within a computer. For example, a PCIE SWITCH is used to switch between the CPU and PCIE devices or expand ports. PCIE SWITCH features multi-host functionality (also known as MultiHost, meaning a network service or platform can simultaneously host multiple hosts or websites), port fault identification and isolation, and flexible port allocation.

[0033] Instance: refers to a group of computing resources consisting of CPU, memory, network, GPU, or storage, sold by elastic computing products. For example, instances can include virtual machine instances, container instances, GPU instances, etc.

[0034] Mobile Operations Center (MOC): used to implement virtualization functions such as input and output, computing, etc. under cloud computing. In this application, MOC can refer to smart network card.

[0035] Example 1

[0036] According to an embodiment of the present application, an embodiment of a resource configuration method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a resource configuration method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (illustrated as 102a, 102b, ..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microcontroller unit (MCU) or a programmable logic device (FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, the computer terminal 10 may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a computer bus), a network interface, a cursor control device (such as a mouse, a touchpad, etc.), a keyboard, a power supply, and / or a camera.

[0038] It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the resource configuration method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned resource configuration method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0041] The transmission device 106 is configured to connect to a network via a network interface to receive or transmit data. Specific examples of the aforementioned network may include a wired and / or wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one embodiment, the transmission device 106 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0042] like Figure 1 The display shown may be, for example, a touch screen liquid crystal display (LCD), which enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0043] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the aforementioned computer device (or mobile device).

[0044] Under the above operating environment, this application provides Figure 2 The resource configuration method shown. Figure 2 is a flow chart of a resource configuration method according to Example 1 of the present application, such as Figure 2 As shown, the resource configuration method includes:

[0045] Step S21: Obtain configuration requirement information corresponding to the computing power service;

[0046] Step S22: performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplugging functions;

[0047] Step S23: providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0048] The aforementioned computing power service is a cloud computing service that dynamically scales computing resources based on the needs of specific scenarios. This service allows users to automatically increase or decrease computing resources based on actual demand, and even adjust the ratio of different resource types to meet different workloads and business requirements. This flexible resource allocation improves resource utilization, reduces costs, and provides faster response times and higher availability.

[0049] In the application scenario, the computing power service can be an elastic computing service. In this scenario, according to the resource configuration method provided in the embodiment of the present application, the scale and performance of computing resources can be dynamically adjusted according to user needs, thereby providing flexible and variable computing instance resources for the elastic computing service. In particular, the method provided in the embodiment of the present application can flexibly select the central processing unit (CPU) resources and graphics processing unit (GPU) resources to be allocated when configuring resources for the elastic computing service.

[0050] In the embodiments of the present application, the resource configuration method corresponding to the above-mentioned computing power service can be used to assist in realizing scientific computing tasks, machine learning tasks (such as training deep learning models, image recognition, natural language processing, etc.), cloud gaming tasks (such as providing high-performance game screen rendering and real-time interaction), big data analysis tasks, etc. in preset application scenarios. The above-mentioned preset application scenarios may be, but are not limited to, scenarios involving cloud computing in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation.

[0051] The above configuration requirement information can be determined by user input data or by the computing service and preset configuration rules. The above configuration requirement information can represent the quantity, type, specifications, etc. of processor resources required by the computing service.

[0052] Based on the above configuration requirement information, the specific implementation method for resource configuration of the above-mentioned general-purpose server and graphics card resource pool can be: performing resource selection, resource combination, and resource packaging operations on the resources that can be provided by the general-purpose server and the graphics processor resources provided by the graphics card resource pool, thereby obtaining the above-mentioned configuration result. The above-mentioned configuration result is used at least to determine the processor resources to be used when providing computing instances for computing power services. The above-mentioned configuration result can also be used to determine the memory resources, service manager resources, etc. to be used when providing computing instances for computing power services. For example, the above-mentioned configuration result can be: the packaging result of the above-mentioned processor resources, the packaging result of the above-mentioned processor resources with memory resources and service manager resources, or the resource identifier of the processor resources, memory resources, or service manager resources.

[0053] It should be noted that the types of processor resources configured in the above-mentioned general server and the processor resources configured in the above-mentioned graphics card resource pool may be different. For example, the processor resources configured in the general server are central processing unit resources, and the processor resources in the graphics card resource pool are graphics processor resources. In particular, the multiple graphics processors in the above-mentioned graphics card resource pool can also support online plug-in and unplugging functions.

[0054] Furthermore, the first instance provided for the computing power service based on the configuration results can be a GPU instance, which can include CPU resources allocated by the general server and GPU resources allocated by the graphics card resource pool. In particular, the GPU resources allocated by the graphics card resource pool can be single GPU card resources, 2GPU card resources, 4GPU card resources, 8GPU card resources, etc.

[0055] Therefore, this application configures resources for the general server and graphics card resource pool according to the configuration requirement information, and provides a first instance of the configuration result, which can realize the reuse of processor resources in the general server and the flexible configuration of multiple graphics processors in the graphics card resource pool (for example, flexibly selecting at least one graphics processor to be mounted on the general server to provide a GPU instance), thereby obtaining rich and diverse configuration results to meet the different needs of computing power services in different types of scenarios.

[0056] In an embodiment of the present application, by obtaining configuration requirement information corresponding to a computing power service; further configuring resources for a general-purpose server and a graphics card resource pool according to the configuration requirement information, a configuration result is obtained, wherein the general-purpose server is obtained based on processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; based on the configuration result, a first instance is provided for the computing power service, wherein the first instance includes processor resources allocated by the general-purpose server and processor resources allocated by the graphics card resource pool. Thus, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall flexibility of heterogeneous models in elastic computing scenarios, thereby solving the technical problem of low resource utilization and poor overall flexibility of the related art due to the fact that heterogeneous models are limited by the fixed configuration of the processor resources of the overall machine.

[0057] In the application scenario, the existing elastic computing system provides several types of computing nodes (Compute Node, referred to as CN) such as Figure 3 shown. Figure 3 In the example, node 1 is a CPU server that supports elastic computing CPU instances, node 2 is a single-GPU server that supports elastic computing GPU instances, node 3 is a 4-GPU server that supports elastic computing GPU instances (that is, 4 GPU cards are mounted on the CPU), and node 4 is an 8-GPU server that supports elastic computing GPU instances (that is, 8 GPU cards are mounted on the CPU).

[0058] In the overall topology of the aforementioned elastic computing system, GPU servers, 4-GPU servers, or 8-GPU servers incorporate GPU cards compared to CPU servers. However, due to the high power consumption of GPU cards and the complex structure and heat dissipation design of GPU servers, the failure rate of GPU servers is much higher than that of CPU servers. Furthermore, existing GPU servers typically utilize the current standard PCIE card form factor and do not support online repair. If any component in the entire server (such as the SmartNIC, memory, CPU, GPU, etc.) fails, the entire server must be taken offline for repair, resulting in poor resource availability for GPU servers.

[0059] Furthermore, different cloud service instances have significantly different requirements for CPU and GPU configuration ratios. For example, search and recommendation scenarios require higher CPUs, while large-model inference scenarios require higher GPU computing power or memory. Traditional AI (such as image recognition) has relatively balanced CPU and GPU requirements. This requires cloud elastic computing systems to flexibly configure CPU and GPU resources to meet the resource requirements of different scenarios. When considering the combination of multiple CPU platforms and multiple GPUs, the number of GPU instance models will increase significantly (typically, the number of models is calculated by multiplying the number of CPU platforms, GPU models, and GPU cards), which greatly increases the difficulty of developing and operating cloud elastic computing systems and resource operations.

[0060] In this regard, based on the above existing elastic computing system, according to the embodiment of the present application, the following is provided: Figure 4 The elastic computing system architecture shown includes multiple general-purpose servers, switches, and a graphics processor resource pool (also called a graphics card resource pool).

[0061] It should be noted that in the elastic computing system architecture provided in the embodiment of the present application, the above-mentioned general-purpose server is used to reuse some models corresponding to the CPU instances in the computing power service, and the general-purpose server reserves the expansion capability of the PCIE interface. In particular, in the embodiment of the present application, the multiple general-purpose servers in the elastic computing system can be the same type of server configured with the same CPU platform, memory and service manager (in this case, the smart network card MOC), or they can be different types of server configured with different CPU platforms or different CPU models and memory and service managers. Therefore, the above-mentioned elastic computing system provided in the embodiment of the present application can realize the functions of all types of models in the traditional elastic computing system (that is, the number of CPU platforms, GPU models and the number of GPU cards multiplied together), thereby enhancing resource availability.

[0062] It is easy to find that the above-mentioned elastic computing system architecture provided in the embodiment of the present application decouples different types of processor resources at the hardware level. For example, the central processing unit is configured in a general-purpose server, and the graphics processor is configured in a graphics processor resource pool. The general-purpose server and the graphics processor resource pool are connected through a data transmission switch with multi-host (mutihost) function (for example, the switch can be a PCIE switch). Therefore, the present application can not only realize the flexible configuration of different types of processor resources according to configuration requirements, but also can perform online fault identification and fault repair of processor resources through the software functions of the data transmission switch, thereby improving the overall resource availability of the elastic computing system.

[0063] In an optional embodiment, in step S21, obtaining configuration requirement information corresponding to the computing power service includes the following method steps:

[0064] Step S211, obtaining the computing task category corresponding to the computing power service;

[0065] Step S212: Generate configuration requirement information based on the computing task category, wherein the configuration requirement information is used to characterize the proportion of different types of processor resources to be configured required for the computing service.

[0066] In the above optional embodiments, the computing task categories corresponding to the computing power service can be the first category, the second category, and the third category, wherein the configuration requirement information corresponding to the first category indicates that the CPU resources in the processor resources to be configured are more than the GPU resources (or the ratio of the CPU resource amount to the GPU resource amount is higher than the first threshold), the configuration requirement information corresponding to the second category indicates that the CPU resources in the processor resources to be configured are less than the GPU resources (or the ratio of the CPU resource amount to the GPU resource amount is lower than the second threshold), and the configuration requirement information corresponding to the third category indicates that the CPU resources and GPU resources in the processor resources to be configured are relatively balanced (or the ratio of the CPU resource amount to the GPU resource amount is within the specified numerical range). For example, the search recommendation scenario has a high demand for the CPU, so the elastic computing tasks in the search recommendation scenario usually belong to the above first category; the large model reasoning scenario has a high demand for the computing power or video memory of the GPU, so the elastic computing tasks in the large model reasoning scenario usually belong to the above second category; traditional artificial intelligence (such as image recognition, etc.) has a relatively balanced demand for CPU and GPU, so the elastic computing tasks in the traditional artificial intelligence scenario usually belong to the above third category.

[0067] In an application scenario, the configuration requirement information may be generated according to preset configuration rules and the computing task types. The preset configuration rules are used to determine the proportion to be configured corresponding to each computing task type.

[0068] The process of automatically generating configuration requirement information based on the computing task category corresponding to the computing power service and then determining the proportion of different types of processor resources to be configured can be implemented by computer code in the built-in chip of the PCIE switch (i.e., the switch). Therefore, through the above method steps, the embodiment of the present application can adapt to the resource configuration requirements corresponding to the computing power service in real time and perform subsequent resource configuration steps, with high flexibility and high resource availability for the entire elastic computing system.

[0069] In an optional embodiment, in step S22, resource configuration is performed on the general server and graphics card resource pool according to the configuration requirement information to obtain a configuration result, including the following method steps:

[0070] Step S221: determining at least one target graphics processor to be mounted to the general server from a plurality of graphics processors in the graphics card resource pool according to the configuration requirement information;

[0071] Step S222 : performing resource combination and packaging processing on the general server and the target graphics processor to obtain a configuration result.

[0072] like Figure 4 As shown, in the elastic computing system provided in the embodiment of the present application, the graphics card resource pool includes multiple graphics processors (also known as pluggable GPU cards). According to the above configuration requirement information, it is possible to determine the type, quantity, specifications, etc. of the processor resources that need to be connected in the instance configured for the current computing task of the computing power service. On this basis, according to the configuration requirement information, at least one GPU card to be mounted to the general server is selected from the multiple GPU cards in the graphics card resource pool. For example, the current computing task of the computing power service requires 4 GPU servers. After the computing task is connected to the general server 1, it is uploaded to the switch by the general server 1. The software module in the switch will select the four GPU cards to be mounted to the general server 1 from the graphics card resource pool based on the configuration requirement information corresponding to the computing task (also known as the above target graphics processor). The specific method of selection can be to determine the identifiers of the four GPU cards. Further, the general server 1 and the above four GPU cards are resource combined and packaged to obtain a configuration result. The configuration result is used to provide a GPU instance (also known as the first instance) for the above computing task. The computing resources corresponding to the GPU instance are 4 GPU servers.

[0073] In an optional embodiment, in step S222, resource combination and packaging processing is performed on the general server and the target graphics processor to obtain a configuration result, including the following method steps:

[0074] Step S2221: Send a first plug-in / plug-out instruction to the target graphics processor, wherein the first plug-in / plug-out instruction is used to control the target graphics processor to adjust the plug-in / plug-out state so as to be mounted to the general server;

[0075] Step S2222: Encapsulate the general server and the target graphics processor to obtain a configuration result.

[0076] like Figure 4As shown, the above-mentioned method steps of the embodiment of the present application can be implemented by a computer program in the built-in chip of the switch. In the application scenario, the multiple graphics processors in the above-mentioned graphics card resource pool are notification-type pluggable GPU cards. During the program operation of the built-in chip of the PCIE switch, after determining at least one target graphics processor to be mounted on the general server, a first plug-in instruction is generated and sent to each target graphics processor in the at least one target graphics processor, so that each target graphics processor adjusts the plug-in state in response to the first plug-in instruction to be mounted on the general server. Furthermore, according to the preset target packaging form, the general server and at least one target graphics processor are packaged to obtain a configuration result.

[0077] In an optional embodiment, the resource configuration method further includes the following method steps:

[0078] Step S24: providing a second instance for the computing service according to the configuration requirement information, wherein the second instance includes processor resources provided by a general server.

[0079] In an application scenario, if the configuration requirement information determines that the current computing task of the computing service only requires a CPU instance and does not require a GPU card, a second instance (i.e., a CPU instance) is provided for the computing service based on the general-purpose server. This second instance may include CPU resources, memory resources, or service manager resources allocated by the general-purpose server.

[0080] In an optional embodiment, the resource configuration method further includes the following method steps:

[0081] Step S251: performing fault identification on a plurality of graphics processors in a general server and a graphics card resource pool to obtain a fault identification result, wherein the fault identification result is used to determine a faulty component;

[0082] Step S252: Based on the fault identification result, the fault component is repaired.

[0083] In the above optional embodiment, since the general server and the graphics card resource pool are decoupled at the hardware level, after identifying that the CPU resources in the general server or the GPU card in the graphics card resource pool have failed, there is no need to take the entire machine offline for repair. Instead, the faulty resources can be isolated, replaced, and reallocated through the software system of the built-in chip of the switch.

[0084] Specifically, real-time fault identification is performed on general-purpose servers and graphics card resource pools, resulting in a fault identification result. This result can indicate that the entire machine is fault-free or include the component identifier of a faulty component. This faulty component can be the CPU, memory, or service manager (MOC) component in a general-purpose server, or a GPU card in a graphics card resource pool. The fault repair process for the faulty component based on the fault identification result is completed using the fault repair tool (such as pre-programmed fault repair code) in the switch's built-in chip.

[0085] In an optional embodiment, in step S252, based on the fault identification result, the fault component is repaired, including the following method steps:

[0086] Step S2521: Based on the fault identification result, determine a replacement component corresponding to the faulty component from the general server and graphics card resource pool, wherein the replacement component is an available component of the same category as the faulty component and has not experienced any faults;

[0087] Step S2522: Send a second plug-in / plug-out instruction to the faulty component and the replacement component, wherein the second plug-in / plug-out instruction is used to control the faulty component to be isolated offline and to control the replacement component to adjust the plug-in / plug-out state so as to be mounted to the general server.

[0088] Through the above method and steps, a faulty component (such as a CPU component or GPU card) in a processor resource that has been packaged into an instance can be replaced with a usable component of the same category as the faulty component and that has not experienced a fault. The usable component can be a component in an available state, which is the state of a component that has not been packaged into an instance.

[0089] In an exemplary application scenario, based on Figure 4 The elastic computing system shown provides an instance 1 consisting of a general server 1 and GPU0. When fault identification is performed on the general server and graphics card resource pool, it is found that the general server 1 is down due to a failure of at least one of the CPU, memory or smart network card, and GPU0 is not faulty. Based on the preset fault isolation strategy, the general server 1 is isolated and taken offline and enters the replacement and maintenance process (such as notifying technicians for repair or waiting for replacement). Based on the preset resource reallocation strategy, GPU0 is quickly mounted to the general server 2 (assuming that the configuration category of the general server 2 is the same as that of the general server 1 and is in an available state) to provide the above-mentioned instance 1, thereby improving the online time of GPU0, that is, improving the availability of the GPU resources.

[0090] In another exemplary application scenario, based on Figure 4The elastic computing system shown provides instance 2, which consists of a general-purpose server 4 and GPU 7. When fault identification is performed on the general-purpose server and graphics card resource pool, if GPU 7 fails but general-purpose server 4 does not, based on a preset fault isolation strategy, GPU 7 is marked as faulty and a replacement repair process is initiated (e.g., notifying a technician for repair or waiting for replacement). Based on a preset resource reallocation strategy, a replacement component, such as GPU 6, is determined from unallocated GPU cards in the graphics card resource pool and mounted and mapped to general-purpose server 4 to form instance 2. In particular, if no unallocated GPU cards exist in the graphics card resource pool when GPU 7 is detected as faulty, a new CPU instance can be provided based on general-purpose server 4 to improve the resource availability of general-purpose server 4.

[0091] It should be noted that, depending on the application scenario requirements, general server 1 can be combined with GPU0 to form a single-card GPU instance to meet the computing and processing requirements of the more complex service logic in the search recommendation scenario; general server 1 can be combined with GPU1 and GPU2 to form a dual-card GPU instance to be suitable for traditional artificial intelligence scenarios such as image recognition where the CPU and GPU resource requirements are relatively balanced; general server 2 can be combined with GPU4, GPU5, GPU6 and GPU7 to form a four-card GPU instance to meet the computing and processing requirements in large model inference scenarios; general server 3 can provide a CPU instance alone, or it can be combined with all GPU cards in the graphics card resource pool to form a GPU instance. In these cases, fault identification and fault repair can be performed on each component in each instance.

[0092] It should be noted that in the application scenario, each general-purpose server can be an independent entity and can be configured with different CPU platforms and different memory specifications (such as different capacities, different speeds, etc.). Each general-purpose server can be interconnected with a switch (PCIE SWITCH in this case) via a PCIE cable. Each general-purpose server can independently sense a CPU instance or can be combined with a GPU card to provide a GPU instance.

[0093] In addition, the above-mentioned switch and graphics processor resource pool can also be configured as a GPU BOX model. That is to say, in addition to the basic functions of providing power supply, heat dissipation and structural support for multiple GPU cards, the GPU BOX model has the switch (PCIE Switch in this case) as a key component in the GPU BOX. The uplink port of the switch has a multi-host (mutihost) function to support the PCIE interface access of multiple general servers, and the downlink port of the switch can be connected to the graphics card resource pool. The multiple GPU cards in the graphics card resource pool can be GPU cards of the same manufacturer and the same model, or GPU cards of different manufacturers and different models. The built-in chip of the above-mentioned switch can be a commercial device or a self-developed chip (for example, obtained after compatible design of part of the logic of the smart chip in the self-developed cloud service platform).

[0094] For example, a typical configuration for the aforementioned GPU BOX model might include four general-purpose servers as the server head, a GPU BOX with a PCIE switch as the base for GPU resources, and eight GPU cards with online swap functionality. In specific application scenarios, the number of general-purpose servers and GPU cards can be flexibly adjusted based on the specifications and performance of the PCIE switch and GPU BOX, as well as the requirements of the cloud application scenario.

[0095] From the above, this application realizes real-time fault identification, online fault repair and online resource reallocation of processor resources by decoupling CPU resources and GPU resources at the hardware level, and also improves the configuration flexibility and resource availability of the entire machine's CPU resources and GPU resources.

[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0097] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0098] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0099] Example 2

[0100] In the operating environment as in Example 1, the present application provides Figure 5 Another resource configuration method is shown. Figure 5 is a flow chart of a resource configuration method according to Example 2 of the present application, such as Figure 5 As shown, the resource configuration method includes:

[0101] Step S51: obtaining a resource configuration request through a first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service;

[0102] Step S52: Return a resource configuration response through the second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of the first instance provided for the computing power service, the first instance is determined based on the configuration result, the configuration result is obtained by performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information, the general server is obtained based on the processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0103] According to the above method steps, a method for implementing a resource configuration service is provided, which is run on a cloud server. The cloud server obtains a resource configuration request issued by a service caller through a first application programming interface (API), executes a resource configuration process based on request data carried in the resource configuration request, and thereby obtains instance information of a first instance corresponding to the configuration result. Furthermore, the cloud server returns a resource configuration response to the service caller through a second application programming interface, thereby providing the instance information of the first instance corresponding to the configuration result to the service caller.

[0104] In the embodiments of the present application, the resource configuration method corresponding to the above-mentioned computing power service can be used to assist in realizing scientific computing tasks, machine learning tasks (such as training deep learning models, image recognition, natural language processing, etc.), cloud gaming tasks (such as providing high-performance game screen rendering and real-time interaction), big data analysis tasks, etc. in preset application scenarios. The above-mentioned preset application scenarios may be, but are not limited to, scenarios involving cloud computing in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation.

[0105] In the application scenario, the computing power service can be an elastic computing service. In this scenario, according to the resource configuration method provided in the embodiment of the present application, the scale and performance of computing resources can be dynamically adjusted according to user needs, thereby providing flexible and variable computing instance resources for the elastic computing service. In particular, the method provided in the embodiment of the present application can flexibly select the central processing unit (CPU) resources and graphics processing unit (GPU) resources to be allocated when configuring resources for the elastic computing service.

[0106] The above configuration requirement information can be determined by user input data or by the computing service and preset configuration rules. The above configuration requirement information can represent the quantity, type, specifications, etc. of processor resources required by the computing service.

[0107] Based on the above configuration requirement information, the specific implementation method for resource configuration of the above-mentioned general-purpose server and graphics card resource pool can be: performing resource selection, resource combination, and resource packaging operations on the resources that can be provided by the general-purpose server and the graphics processor resources provided by the graphics card resource pool, thereby obtaining the above-mentioned configuration result. The above-mentioned configuration result is used at least to determine the processor resources to be used when providing computing instances for computing power services. The above-mentioned configuration result can also be used to determine the memory resources, service manager resources, etc. to be used when providing computing instances for computing power services. For example, the above-mentioned configuration result can be: the packaging result of the above-mentioned processor resources, the packaging result of the above-mentioned processor resources with memory resources and service manager resources, or the resource identifier of the processor resources, memory resources, or service manager resources.

[0108] It should be noted that the types of processor resources configured in the above-mentioned general server and the processor resources configured in the above-mentioned graphics card resource pool may be different. For example, the processor resources configured in the general server are central processing unit resources, and the processor resources in the graphics card resource pool are graphics processor resources. In particular, the multiple graphics processors in the above-mentioned graphics card resource pool can also support online plug-in and unplugging functions.

[0109] Furthermore, the first instance provided for the computing power service based on the configuration results can be a GPU instance, which can include CPU resources allocated by the general server and GPU resources allocated by the graphics card resource pool. In particular, the GPU resources allocated by the graphics card resource pool can be single GPU card resources, 2GPU card resources, 4GPU card resources, 8GPU card resources, etc.

[0110] Therefore, this application configures resources for the general server and graphics card resource pool according to the configuration requirement information, and provides a first instance of the configuration result, which can realize the reuse of processor resources in the general server and the flexible configuration of multiple graphics processors in the graphics card resource pool (for example, flexibly selecting at least one graphics processor to be mounted on the general server to provide a GPU instance), thereby obtaining rich and diverse configuration results to meet the different needs of computing power services in different types of scenarios.

[0111] In an embodiment of the present application, a resource configuration request is obtained through a first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; a resource configuration response is returned through a second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of a first instance provided for the computing power service, the first instance is determined based on the configuration result, the configuration result is obtained by configuring resources for a general server and a graphics card resource pool according to the configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool. Thus, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall machine flexibility in heterogeneous models in elastic computing scenarios, thereby solving the technical problems of low resource utilization and poor overall machine flexibility caused by the fact that heterogeneous models are limited by the fixed configuration of the overall machine processor resources in the related art.

[0112] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.

[0113] Example 3

[0114] In the operating environment as in Example 1, the present application provides Figure 6 Another resource configuration method is shown. Figure 6 is a flow chart of a resource configuration method according to Example 3 of the present application, such as Figure 6 As shown, the resource configuration method includes:

[0115] Step S61: Acquire the currently input resource configuration request, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service;

[0116] Step S62: In response to the resource configuration request, a resource configuration reply is returned, wherein the resource configuration reply carries information including: instance information of a first instance provided for the computing power service, the first instance being determined based on a configuration result, the configuration result being obtained by configuring resources of a general-purpose server and a graphics card resource pool according to the configuration requirement information, the general-purpose server being configured based on processor resources, the graphics card resource pool including multiple graphics processors supporting online pluggable functionality, the first instance including processor resources allocated by the general-purpose server and processor resources allocated by the graphics card resource pool;

[0117] Step S63: Display the instance information in the graphical user interface.

[0118] According to the above method steps, a visualization solution for resource configuration function is provided. The terminal device provides a graphical user interface, and at least one resource configuration scenario is displayed in the graphical user interface. The display content of the graphical user interface also includes an input component (such as a text input box, a voice input control, etc.) and a display component (such as a text display window, an image display window, etc.). The user inputs a resource configuration request through the input component to specify the configuration requirement information corresponding to the computing power service. After detecting the user's input behavior, the resource configuration process is executed based on the configuration requirement information to obtain the instance information of the first instance corresponding to the configuration result. Further, the instance information of the first instance corresponding to the configuration result is displayed through the display component in the graphical user interface.

[0119] In the embodiments of the present application, the resource configuration method corresponding to the above-mentioned computing power service can be used to assist in realizing scientific computing tasks, machine learning tasks (such as training deep learning models, image recognition, natural language processing, etc.), cloud gaming tasks (such as providing high-performance game screen rendering and real-time interaction), big data analysis tasks, etc. in preset application scenarios. The above-mentioned preset application scenarios may be, but are not limited to, scenarios involving cloud computing in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation.

[0120] In the application scenario, the computing power service can be an elastic computing service. In this scenario, according to the resource configuration method provided in the embodiment of the present application, the scale and performance of computing resources can be dynamically adjusted according to user needs, thereby providing flexible and variable computing instance resources for the elastic computing service. In particular, the method provided in the embodiment of the present application can flexibly select the central processing unit (CPU) resources and graphics processing unit (GPU) resources to be allocated when configuring resources for the elastic computing service.

[0121] The above configuration requirement information can be determined by user input data or by the computing service and preset configuration rules. The above configuration requirement information can represent the quantity, type, specifications, etc. of processor resources required by the computing service.

[0122] Based on the above configuration requirement information, the specific implementation method for resource configuration of the above-mentioned general-purpose server and graphics card resource pool can be: performing resource selection, resource combination, and resource packaging operations on the resources that can be provided by the general-purpose server and the graphics processor resources provided by the graphics card resource pool, thereby obtaining the above-mentioned configuration result. The above-mentioned configuration result is used at least to determine the processor resources to be used when providing computing instances for computing power services. The above-mentioned configuration result can also be used to determine the memory resources, service manager resources, etc. to be used when providing computing instances for computing power services. For example, the above-mentioned configuration result can be: the packaging result of the above-mentioned processor resources, the packaging result of the above-mentioned processor resources with memory resources and service manager resources, or the resource identifier of the processor resources, memory resources, or service manager resources.

[0123] It should be noted that the types of processor resources configured in the above-mentioned general server and the processor resources configured in the above-mentioned graphics card resource pool may be different. For example, the processor resources configured in the general server are central processing unit resources, and the processor resources in the graphics card resource pool are graphics processor resources. In particular, the multiple graphics processors in the above-mentioned graphics card resource pool can also support online plug-in and unplugging functions.

[0124] Furthermore, the first instance provided for the computing power service based on the configuration results can be a GPU instance, which can include CPU resources allocated by the general server and GPU resources allocated by the graphics card resource pool. In particular, the GPU resources allocated by the graphics card resource pool can be single GPU card resources, 2GPU card resources, 4GPU card resources, 8GPU card resources, etc.

[0125] Therefore, this application configures resources for the general server and graphics card resource pool according to the configuration requirement information, and provides a first instance of the configuration result, which can realize the reuse of processor resources in the general server and the flexible configuration of multiple graphics processors in the graphics card resource pool (for example, flexibly selecting at least one graphics processor to be mounted on the general server to provide a GPU instance), thereby obtaining rich and diverse configuration results to meet the different needs of computing power services in different types of scenarios.

[0126] In an embodiment of the present application, a resource configuration request currently input is obtained, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; in response to the resource configuration request, a resource configuration reply is returned, wherein the information carried in the resource configuration reply includes: instance information of the first instance provided for the computing power service, the first instance is determined based on the configuration result, the configuration result is obtained by configuring resources for the general server and the graphics card resource pool according to the configuration requirement information, the general server is obtained based on the processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool; the instance information is displayed in a graphical user interface. Thus, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall machine flexibility in heterogeneous models in elastic computing scenarios, thereby solving the technical problems in related technologies of low resource utilization and poor overall machine flexibility caused by the fact that heterogeneous models are limited by the fixed configuration of the overall machine processor resources.

[0127] It should be noted that the preferred implementation of this embodiment can refer to the relevant description in Example 1 or Example 2, and will not be repeated here.

[0128] Example 4

[0129] In the operating environment as in Example 1, the present application provides Figure 7 Another resource configuration method is shown. Figure 7 is a flow chart of a resource configuration method according to Example 4 of the present application, such as Figure 7 As shown, the resource configuration method includes:

[0130] Step S71: Obtain configuration requirement information corresponding to the computing power service;

[0131] Step S72: Using the configuration requirement information and the resource configuration model, perform resource configuration on the general-purpose server and the graphics card resource pool to obtain a configuration result. The resource configuration model is a neural network model pre-trained by machine learning using multiple sets of training data. The general-purpose server is configured based on processor resources, and the graphics card resource pool includes multiple graphics processors that support online plug-in functionality.

[0132] Step S73: Provide a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0133] Each of the multiple sets of training data corresponding to the resource configuration model described above includes configuration requirement samples and corresponding real-world mount information. The initial neural network model is used to predict the input configuration requirement information to obtain the preset mount information. The training loss is calculated based on the real-world and predicted mount information. The network parameters of the initial neural network model are adjusted based on the training loss to obtain the resource configuration model.

[0134] The above-mentioned resource configuration method provided in the embodiment of the present application also includes other optional implementation methods, and you can refer to the relevant descriptions in the aforementioned embodiments, which will not be repeated here.

[0135] In an optional embodiment, in step S72, resource configuration is performed on a general server and a graphics card resource pool using the configuration requirement information and the resource configuration model to obtain a configuration result, including the following method steps:

[0136] Step S721: Input the configuration requirement information into the resource configuration model to obtain information to be mounted, wherein the information to be mounted is used to determine at least one target graphics processor to be mounted to the general server from multiple graphics processors in the graphics card resource pool;

[0137] Step S722: Combining and packaging resources of the general server and the target GPU to obtain a configuration result.

[0138] The above-mentioned computing power service can be an elastic computing service. Based on this, in the elastic computing system provided in the embodiment of the present application, the graphics card resource pool includes multiple graphics processors (that is, pluggable GPU cards). In the application scenario, a pre-trained resource configuration model is used. According to the above-mentioned configuration requirement information, the type, quantity, specifications, etc. of the processor resources that need to be accessed in the instance configured for the current computing task of the computing power service can be determined. On this basis, according to the configuration requirement information, at least one GPU card to be attached to the general server is selected from the multiple GPU cards in the graphics card resource pool.

[0139] In an embodiment of the present application, a resource configuration request currently input is obtained, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; in response to the resource configuration request, a resource configuration reply is returned, wherein the information carried in the resource configuration reply includes: instance information of the first instance provided for the computing power service, the first instance is determined based on the configuration result, the configuration result is obtained by configuring resources for the general server and the graphics card resource pool according to the configuration requirement information, the general server is obtained based on the processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool; the instance information is displayed in a graphical user interface. Thus, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall machine flexibility in heterogeneous models in elastic computing scenarios, thereby solving the technical problems in related technologies of low resource utilization and poor overall machine flexibility caused by the fact that heterogeneous models are limited by the fixed configuration of the overall machine processor resources.

[0140] It should be noted that the preferred implementation of this embodiment can be found in the relevant descriptions in Example 1, Example 2 or Example 3, and will not be repeated here.

[0141] Example 5

[0142] According to an embodiment of the present application, a system embodiment for implementing the above-mentioned resource configuration method is also provided. Figure 8 This is a schematic diagram of the structure of a resource configuration system according to Example 5 of the present application. Figure 8 As shown, the resource configuration system includes:

[0143] General purpose server 801 is configured based on processor resources, memory resources and service manager resources;

[0144] Graphics card resource pool 802, including multiple graphics processors supporting online plug-in and unplug functionality;

[0145] The graphics card interface box 803 is connected to the general server and the graphics card resource pool, and is used to configure resources for the general server and the graphics card resource pool according to the configuration requirement information corresponding to the computing power service, and obtain the configuration result. The configuration result is used to support the general server to provide a first instance for the computing power service. The first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0146] Optionally, in the above resource configuration system, the graphics card interface box is further used to identify and repair faults of the general server and graphics card resource pool.

[0147] Optionally, in the above resource configuration system, the general server is also used to provide a second instance for the computing power service, wherein the second instance includes processor resources provided by the general server.

[0148] Based on the above resource configuration system, the resource configuration methods provided in the above embodiments 1, 2, 3, or 4 can be implemented. The resource configuration methods corresponding to the above computing power services can be applied to assist in implementing scientific computing tasks, machine learning tasks (such as training deep learning models, image recognition, natural language processing, etc.), cloud gaming tasks (such as providing high-performance game screen rendering and real-time interaction), big data analysis tasks, etc. in preset application scenarios. The above preset application scenarios can include, but are not limited to: scenarios involving cloud computing in fields such as e-commerce, education, medical care, conferences, social networks, financial products, logistics, and navigation.

[0149] The above configuration requirement information can be determined by user input data or by the computing service and preset configuration rules. The above configuration requirement information can represent the quantity, type, specifications, etc. of processor resources required by the computing service.

[0150] Based on the above configuration requirement information, the specific implementation method for resource configuration of the above-mentioned general-purpose server and graphics card resource pool can be: performing resource selection, resource combination, and resource packaging operations on the resources that can be provided by the general-purpose server and the graphics processor resources provided by the graphics card resource pool, thereby obtaining the above-mentioned configuration result. The above-mentioned configuration result is used at least to determine the processor resources to be used when providing computing instances for computing power services. The above-mentioned configuration result can also be used to determine the memory resources, service manager resources, etc. to be used when providing computing instances for computing power services. For example, the above-mentioned configuration result can be: the packaging result of the above-mentioned processor resources, the packaging result of the above-mentioned processor resources with memory resources and service manager resources, or the resource identifier of the processor resources, memory resources, or service manager resources.

[0151] It should be noted that the types of processor resources configured in the above-mentioned general server and the processor resources configured in the above-mentioned graphics card resource pool may be different. For example, the processor resources configured in the general server are central processing unit resources, and the processor resources in the graphics card resource pool are graphics processor resources. In particular, the multiple graphics processors in the above-mentioned graphics card resource pool can also support online plug-in and unplugging functions.

[0152] Furthermore, the first instance provided for the computing power service based on the configuration results can be a GPU instance, which can include CPU resources allocated by the general server and GPU resources allocated by the graphics card resource pool. In particular, the GPU resources allocated by the graphics card resource pool can be single GPU card resources, 2GPU card resources, 4GPU card resources, 8GPU card resources, etc.

[0153] Therefore, this application configures resources for the general server and graphics card resource pool according to the configuration requirement information, and provides a first instance of the configuration result, which can realize the reuse of processor resources in the general server and the flexible configuration of multiple graphics processors in the graphics card resource pool (for example, flexibly selecting at least one graphics processor to be mounted on the general server to provide a GPU instance), thereby obtaining rich and diverse configuration results to meet the different needs of computing power services in different types of scenarios.

[0154] In an embodiment of the present application, a resource configuration system is proposed, including: a general server, which is configured based on processor resources, memory resources and service manager resources; a graphics card resource pool, including multiple graphics processors that support online plug-in functions; a graphics card interface box, which is connected to the general server and the graphics card resource pool, and is used to configure resources for the general server and the graphics card resource pool according to the configuration requirement information corresponding to the computing power service, and obtain a configuration result, which is used to support the general server to provide a first instance for the computing power service, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool. As a result, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall machine flexibility in heterogeneous models in elastic computing scenarios, and further solving the technical problems of low resource utilization and poor overall machine flexibility caused by the fact that heterogeneous models are limited by the fixed configuration of the processor resources of the overall machine in the related art.

[0155] It should be noted that the preferred implementation of this embodiment can refer to the relevant descriptions in Example 1, Example 2, Example 3 or Example 4, and will not be repeated here.

[0156] Example 6

[0157] According to an embodiment of the present application, a device embodiment for implementing the above-mentioned resource configuration method is also provided. Figure 9 is a structural diagram of a resource configuration device according to Example 6 of the present application, such as Figure 9 As shown, the device includes:

[0158] Acquisition module 901, used to obtain configuration requirement information corresponding to the computing power service;

[0159] Configuration module 902, configured to configure resources for a general server and a graphics card resource pool according to the configuration requirement information, and obtain a configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in function;

[0160] The first instance module 903 is used to provide a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0161] Optionally, the above-mentioned acquisition module 901 is also used to: obtain the computing task category corresponding to the computing power service; generate configuration requirement information based on the computing task category, wherein the configuration requirement information is used to characterize the proportion of different types of processor resources to be configured required for the computing power service.

[0162] Optionally, the configuration module 902 is further configured to: determine, based on configuration requirement information, at least one target graphics processor to be mounted to a general server from a plurality of graphics processors in a graphics card resource pool; and perform resource combination and packaging processing on the general server and the target graphics processor to obtain a configuration result.

[0163] Optionally, the configuration module 902 is further configured to: send a first plug-in instruction to the target graphics processor, wherein the first plug-in instruction is configured to control the target graphics processor to adjust the plug-in state so as to be mounted on the general server; and perform packaging processing on the general server and the target graphics processor to obtain a configuration result.

[0164] Optionally, in addition to all the above-mentioned modules, the above-mentioned resource configuration device also includes: a second instance module 904 (not shown in the figure), which is used to: provide a second instance for the computing power service according to the configuration requirement information, wherein the second instance includes processor resources provided by a general server.

[0165] Optionally, in addition to all the above-mentioned modules, the above-mentioned resource configuration device also includes: a fault module 905 (not shown in the figure), which is used to: identify faults of multiple graphics processors in the general server and graphics card resource pool to obtain fault identification results, wherein the fault identification results are used to determine the faulty component where the fault occurs; based on the fault identification results, the fault component is repaired.

[0166] Optionally, the above-mentioned fault module 905 is also used to: determine a replacement component corresponding to the fault component from the general server and graphics card resource pool based on the fault identification result, wherein the replacement component is an available component that is of the same category as the fault component and has not failed; send a second plug-in and unplug-out instruction to the faulty component and the replacement component, wherein the second plug-in and unplug-out instruction is used to control the offline isolation of the faulty component and control the replacement component to adjust the plug-in and unplugging status so as to be mounted on the general server.

[0167] In an embodiment of the present application, the configuration requirement information corresponding to the computing power service is obtained through the acquisition module; the configuration module is further used to configure the resources of the general server and the graphics card resource pool according to the configuration requirement information to obtain the configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; on this basis, the first instance module is used to provide the computing power service with a first instance based on the configuration result, wherein the first instance includes the processor resources allocated by the general server and the processor resources allocated by the graphics card resource pool. Thus, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall flexibility of heterogeneous models in elastic computing scenarios, and further solving the technical problems of low resource utilization and poor overall flexibility caused by the fact that heterogeneous models are limited by the fixed configuration of the processor resources of the overall machine in the related technology.

[0168] It should be noted that the acquisition module 901, configuration module 902, and first instance module 903 correspond to steps S21 to S23 in Example 1. The three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be part of a device and can be run in the computer terminal 10 provided in Example 1.

[0169] According to an embodiment of the present application, a device embodiment for implementing the resource configuration method in the above-mentioned embodiment 2 is also provided. Figure 10 is a structural diagram of another resource configuration device according to Example 6 of the present application, such as Figure 10 As shown, the device includes:

[0170] The request module 1001 is configured to obtain a resource configuration request through a first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service;

[0171] Response module 1002 is used to return a resource configuration response through a second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of a first instance provided for computing power services, the first instance is determined based on a configuration result, the configuration result is obtained by performing resource configuration on a general server and a graphics card resource pool according to configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0172] It should be noted that the request module 1001 and the response module 1002 correspond to steps S51 to S52 in Example 2. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 2. It should be noted that the modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be part of the device and can be run in the computer terminal 10 provided in Example 1.

[0173] According to an embodiment of the present application, a device embodiment for implementing the resource configuration method in the above-mentioned embodiment 3 is also provided. Figure 11 is a structural diagram of another resource configuration device according to Example 6 of the present application, such as Figure 11 As shown, the device includes:

[0174] The request module 1101 is used to obtain the currently input resource configuration request, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service;

[0175] A reply module 1102 is configured to respond to the resource configuration request and return a resource configuration reply, wherein the resource configuration reply includes: instance information of a first instance provided for the computing service, the first instance being determined based on a configuration result, the configuration result being obtained by configuring resources of a general-purpose server and a graphics card resource pool according to the configuration requirement information, the general-purpose server being configured based on processor resources, the graphics card resource pool including multiple graphics processors supporting online pluggable functionality, the first instance including processor resources allocated by the general-purpose server and processor resources allocated by the graphics card resource pool;

[0176] The display module 1103 is used to display instance information in a graphical user interface.

[0177] It should be noted that the request module 1101, reply module 1102, and display module 1103 correspond to steps S61 to S63 in Example 3. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 3. It should be noted that the modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be run as part of the device in the computer terminal 10 provided in Example 1.

[0178] It should be noted that the preferred implementation of this embodiment can refer to the relevant descriptions in Example 1, Example 2, Example 3 or Example 4, and will not be repeated here.

[0179] Example 7

[0180] According to an embodiment of the present application, an electronic device is further provided, which can be any computer device in a computer device group. Optionally, in this embodiment, the electronic device can also be replaced by a terminal device such as a mobile terminal.

[0181] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0182] In this embodiment, the above-mentioned electronic device can execute the program code of the following steps in the resource configuration method: obtaining configuration requirement information corresponding to the computing power service; performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0183] Optionally, Figure 12 is a structural block diagram of an electronic device according to embodiment 7 of the present application, such as Figure 12 As shown, the electronic device 120 may include: one or more (only one is shown in the figure) processors 1202, a memory 1204, a storage controller 1206, and a peripheral interface 1208, wherein the peripheral interface 1208 is connected to the radio frequency module, the audio module and the display.

[0184] Among them, the memory 1204 can be used to store software programs and modules, such as the program instructions / modules corresponding to the resource configuration method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned resource configuration method. The memory 1204 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1204 may further include a memory remotely located relative to the processor, and these remote memories can be connected to the electronic device 120 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0185] Processor 1202 can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain configuration requirement information corresponding to the computing power service; configure resources for the general server and the graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; provide a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0186] Optionally, the processor 1202 may also execute the program code of the following steps: obtaining the computing task category corresponding to the computing power service; generating configuration requirement information based on the computing task category, wherein the configuration requirement information is used to characterize the proportion of different types of processor resources to be configured required for the computing power service.

[0187] Optionally, the processor 1202 may further execute program code for the following steps: determining, based on configuration requirement information, at least one target graphics processor to be mounted to a general-purpose server from a plurality of graphics processors in a graphics card resource pool; and performing resource combination and packaging processing on the general-purpose server and the target graphics processor to obtain a configuration result.

[0188] Optionally, the processor 1202 may further execute program code for the following steps: sending a first plug-in / plug-out instruction to the target graphics processor, wherein the first plug-in / plug-out instruction is used to control the target graphics processor to adjust the plug-in / plug-out state so as to be mounted on the general server; and performing packaging processing on the general server and the target graphics processor to obtain a configuration result.

[0189] Optionally, the processor 1202 may further execute program code of the following steps: providing a second instance for the computing service according to the configuration requirement information, wherein the second instance includes processor resources provided by a general server.

[0190] Optionally, the processor 1202 may also execute program code for the following steps: performing fault identification on a plurality of graphics processors in a general server and a graphics card resource pool to obtain a fault identification result, wherein the fault identification result is used to determine a faulty component where the fault occurs; and performing fault repair on the faulty component based on the fault identification result.

[0191] Optionally, the processor 1202 may also execute the following program code: based on the fault identification result, determining a replacement component corresponding to the faulty component from the general server and graphics card resource pool, wherein the replacement component is an available component that is of the same category as the faulty component and has not failed; sending a second plug-in instruction to the faulty component and the replacement component, wherein the second plug-in instruction is used to control the offline isolation of the faulty component and to control the replacement component to adjust the plug-in status so as to be mounted on the general server.

[0192] The processor 1202 can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a resource configuration request through the first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; return a resource configuration response through the second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of the first instance provided for the computing power service, the first instance is determined based on the configuration result, and the configuration result is obtained by configuring resources for the general server and the graphics card resource pool according to the configuration requirement information. The general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in functions. The first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0193] Processor 1202 can call the information and application stored in the memory through the transmission device to perform the following steps: obtain the currently input resource configuration request, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; in response to the resource configuration request, return a resource configuration reply, wherein the information carried in the resource configuration reply includes: instance information of the first instance provided for the computing power service, the first instance is determined based on the configuration result, and the configuration result is obtained by configuring resources for the general server and the graphics card resource pool according to the configuration requirement information, the general server is obtained based on the processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool; and display the instance information in the graphical user interface.

[0194] Processor 1202 can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain configuration requirement information corresponding to the computing power service; use the configuration requirement information and resource configuration model to perform resource configuration on the general server and the graphics card resource pool to obtain a configuration result, wherein the resource configuration model is a neural network model pre-trained by machine learning using multiple sets of training data, the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; based on the configuration result, a first instance is provided for the computing power service, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0195] Optionally, the processor 1202 may further execute program code for the following steps: inputting configuration requirement information into a resource configuration model to obtain information to be mounted, wherein the information to be mounted is used to determine at least one target graphics processor to be mounted to a general-purpose server from a plurality of graphics processors in a graphics card resource pool; and performing resource combination and packaging processing on the general-purpose server and the target graphics processor to obtain a configuration result.

[0196] According to an embodiment of the present application, an electronic device for implementing the above-mentioned resource configuration method is provided. By obtaining the configuration requirement information corresponding to the computing power service; further configuring the resources of the general server and the graphics card resource pool according to the configuration requirement information, a configuration result is obtained, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in functions; based on the configuration result, a first instance is provided for the computing power service, wherein the first instance includes the processor resources allocated by the general server and the processor resources allocated by the graphics card resource pool. Thus, the present application achieves the purpose of flexibly configuring the processor resources based on the configuration requirements, thereby realizing the technical effect of improving the resource utilization and the flexibility of the whole machine in the heterogeneous models in the elastic computing scenario, thereby solving the technical problems of low resource utilization and poor flexibility of the whole machine caused by the fact that the heterogeneous models are limited by the fixed configuration of the processor resources of the whole machine in the related technology.

[0197] It can be understood by those skilled in the art that Figure 12 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile Internet device (Mobile Internet Devices, MID for short). Figure 12 It does not limit the structure of the above electronic device. For example, the electronic device 120 may also include Figure 12 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 12 Different configurations shown.

[0198] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which can include: a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.

[0199] Example 8

[0200] According to an embodiment of the present application, a computer-readable storage medium is further provided. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the resource configuration method provided in the above embodiment 1, embodiment 2, embodiment 3 or embodiment 4.

[0201] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0202] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining configuration requirement information corresponding to the computing power service; performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0203] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining the computing task category corresponding to the computing power service; generating configuration requirement information based on the computing task category, wherein the configuration requirement information is used to characterize the proportion of different types of processor resources to be configured required for the computing power service.

[0204] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: determining, based on configuration requirement information, at least one target graphics processor to be mounted to a general server from a plurality of graphics processors in a graphics card resource pool; and performing resource combination and packaging processing on the general server and the target graphics processor to obtain a configuration result.

[0205] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: sending a first plug-in / plug-out instruction to the target graphics processor, wherein the first plug-in / plug-out instruction is used to control the target graphics processor to adjust the plug-in / plug-out state so as to be mounted to the general server; and performing packaging processing on the general server and the target graphics processor to obtain a configuration result.

[0206] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: providing a second instance for the computing power service according to the configuration requirement information, wherein the second instance includes processor resources provided by the general server.

[0207] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing fault identification on multiple graphics processors in a general server and a graphics card resource pool to obtain fault identification results, wherein the fault identification results are used to determine a faulty component where the fault occurs; and performing fault repair on the faulty component based on the fault identification results.

[0208] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: based on the fault identification result, determining a replacement component corresponding to the faulty component from the general server and graphics card resource pool, wherein the replacement component is an available component that is of the same category as the faulty component and has not failed; sending a second plug-in instruction to the faulty component and the replacement component, wherein the second plug-in instruction is used to control the offline isolation of the faulty component and control the replacement component to adjust the plug-in status so as to be mounted on the general server.

[0209] Optionally, in this embodiment, a computer-readable storage medium is configured to store program code for executing the following steps: obtaining a resource configuration request through a first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; returning a resource configuration response through a second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of a first instance provided for the computing power service, the first instance is determined based on a configuration result, and the configuration result is obtained by performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0210] Optionally, in this embodiment, a computer-readable storage medium is configured to store program code for executing the following steps: obtaining a currently input resource configuration request, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; returning a resource configuration reply in response to the resource configuration request, wherein the information carried in the resource configuration reply includes: instance information of a first instance provided for the computing power service, the first instance is determined based on a configuration result, and the configuration result is obtained by performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool; and displaying the instance information in a graphical user interface.

[0211] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining configuration requirement information corresponding to the computing power service; using the configuration requirement information and the resource configuration model to perform resource configuration on the general server and the graphics card resource pool to obtain a configuration result, wherein the resource configuration model is a neural network model pre-trained by machine learning using multiple sets of training data, the general server is obtained based on processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0212] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: inputting configuration requirement information into a resource configuration model to obtain information to be mounted, wherein the information to be mounted is used to determine at least one target graphics processor to be mounted to a general-purpose server from a plurality of graphics processors in a graphics card resource pool; performing resource combination and packaging processing on the general-purpose server and the target graphics processor to obtain a configuration result. Using an embodiment of the present application, a computer-readable storage medium for implementing the above-mentioned resource configuration method is provided. By obtaining configuration requirement information corresponding to a computing power service; further performing resource configuration on the general-purpose server and the graphics card resource pool according to the configuration requirement information, a configuration result is obtained, wherein the general-purpose server is obtained based on the processor resource configuration, and the graphics card resource pool includes a plurality of graphics processors that support online plug-in and unplugging functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general-purpose server and processor resources allocated by the graphics card resource pool. Therefore, the present application achieves the purpose of flexibly configuring processor resources based on configuration requirements, thereby realizing the technical effect of improving resource utilization and overall machine flexibility in heterogeneous models in elastic computing scenarios, and further solving the technical problems in related technologies of low resource utilization and poor overall machine flexibility caused by heterogeneous models being limited by the fixed configuration of the overall machine processor resources.

[0213] According to an embodiment of the present application, a computer program product is further provided. Optionally, in this embodiment, the computer program product can provide resource configuration services based on the resource configuration method provided in the above embodiment 1, embodiment 2, embodiment 3 or embodiment 4.

[0214] Optionally, in this embodiment, the computer program product may be a set of instructions and codes pre-written according to the resource configuration method. The computer program product may run on various computer platforms, including personal computers, servers, mobile devices, etc.

[0215] Optionally, in this embodiment, the instructions and codes corresponding to the computer program product are used to implement the following method steps: obtaining configuration requirement information corresponding to the computing power service; performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on the processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplug functions; providing a first instance for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

[0216] Through the above-mentioned computer program product, resource configuration services can be provided in application scenarios involving processor resource configuration for cloud computing services, thereby achieving the purpose of flexibly configuring processor resources based on configuration requirements, thereby achieving the technical effect of improving resource utilization and overall machine flexibility in heterogeneous models in elastic computing scenarios, and thus solving the technical problems in related technologies of low resource utilization and poor overall machine flexibility caused by heterogeneous models being limited by the fixed configuration of overall machine processor resources.

[0217] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0218] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0219] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0220] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0221] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, ROM, RAM, mobile hard drives, magnetic disks or optical disks.

[0223] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A resource configuration method, characterized in that: include: Obtain configuration requirements information corresponding to the computing power service; Performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information to obtain a configuration result, wherein the general server is obtained based on processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in and unplugging functions; A first instance is provided for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

2. The resource allocation method according to claim 1, characterized in that: Obtaining the configuration requirement information corresponding to the computing power service includes: Obtain the computing task category corresponding to the computing power service; The configuration requirement information is generated based on the computing task category, wherein the configuration requirement information is used to characterize the proportion of different types of processor resources to be configured required by the computing power service.

3. The resource allocation method according to claim 1, characterized in that: Performing resource configuration on the general server and the graphics card resource pool according to the configuration requirement information, and obtaining the configuration result includes: Determining, according to the configuration requirement information, at least one target graphics processor to be mounted to the general server from the plurality of graphics processors in the graphics card resource pool; Resource combination and packaging processing is performed on the general server and the target graphics processor to obtain the configuration result.

4. The resource allocation method according to claim 3, characterized in that: Combining and packaging resources of the general server and the target graphics processor to obtain the configuration result includes: Sending a first plug-in / plug-out instruction to the target graphics processor, wherein the first plug-in / plug-out instruction is used to control the target graphics processor to adjust the plug-in / plug-out state so as to be mounted to the general server; The general server and the target graphics processor are packaged to obtain the configuration result.

5. The resource allocation method according to claim 1, characterized in that: The resource configuration method further includes: A second instance is provided for the computing power service according to the configuration requirement information, wherein the second instance includes processor resources provided by the general server.

6. The resource allocation method according to claim 1, characterized in that: The resource configuration method further includes: Performing fault identification on the general server and the plurality of graphics processors in the graphics resource pool to obtain a fault identification result, wherein the fault identification result is used to determine a faulty component; Based on the fault identification result, the fault component is repaired.

7. The resource allocation method according to claim 6, characterized in that: Based on the fault identification result, repairing the fault component includes: Based on the fault identification result, determining a replacement component corresponding to the faulty component from the general server and the graphics card resource pool, wherein the replacement component is an available component of the same category as the faulty component and has not experienced a fault; A second plug-in / plug-out instruction is sent to the faulty component and the replacement component, wherein the second plug-in / plug-out instruction is used to control the faulty component to be isolated offline and to control the replacement component to adjust the plug-in / plug-out state so as to be mounted to the general server.

8. A resource allocation method, characterized in that: include: Obtaining a resource configuration request through a first application programming interface, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; A resource configuration response is returned through a second application programming interface, wherein the response data carried in the resource configuration response includes: instance information of a first instance provided for the computing power service, the first instance is determined based on a configuration result, the configuration result is obtained by performing resource configuration on a general server and a graphics card resource pool according to the configuration requirement information, the general server is obtained based on processor resource configuration, the graphics card resource pool includes multiple graphics processors that support online plug-in functions, and the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

9. A resource allocation method, characterized in that: include: Obtaining a currently input resource configuration request, wherein the request data carried in the resource configuration request includes: configuration requirement information corresponding to the computing power service; In response to the resource configuration request, returning a resource configuration reply, wherein information carried in the resource configuration reply includes: instance information of a first instance provided for the computing power service, the first instance being determined based on a configuration result, the configuration result being obtained by performing resource configuration on a general-purpose server and a graphics card resource pool according to the configuration requirement information, the general-purpose server being obtained based on processor resource configuration, the graphics card resource pool including multiple graphics processors supporting online pluggable functionality, the first instance including processor resources allocated by the general-purpose server and processor resources allocated by the graphics card resource pool; The instance information is displayed in a graphical user interface.

10. A resource allocation method, characterized in that: include: Obtain configuration requirements information corresponding to the computing power service; Using the configuration requirement information and the resource configuration model, resource configuration is performed on a general-purpose server and a graphics card resource pool to obtain a configuration result, wherein the resource configuration model is a neural network model pre-trained by machine learning using multiple sets of training data, the general-purpose server is obtained based on processor resource configuration, and the graphics card resource pool includes multiple graphics processors that support online plug-in functionality; A first instance is provided for the computing power service based on the configuration result, wherein the first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

11. The resource allocation method according to claim 10, characterized in that: Using the configuration requirement information and resource configuration model, resources are configured for the general server and graphics card resource pool, and the configuration results obtained include: Inputting the configuration requirement information into the resource configuration model to obtain information to be mounted, wherein the information to be mounted is used to determine at least one target graphics processor to be mounted to the general server from the plurality of graphics processors in the graphics card resource pool; Resource combination and packaging processing is performed on the general server and the target graphics processor to obtain the configuration result.

12. A resource allocation system, characterized in that: include: General purpose servers are configured based on processor resources, memory resources, and service manager resources; Graphics resource pool, including multiple graphics processors that support online plug-in and unplug functionality; A graphics card interface box is connected to the general server and the graphics card resource pool, and is used to configure resources for the general server and the graphics card resource pool according to the configuration requirement information corresponding to the computing power service to obtain a configuration result. The configuration result is used to support the general server to provide a first instance for the computing power service. The first instance includes processor resources allocated by the general server and processor resources allocated by the graphics card resource pool.

13. The resource allocation system according to claim 12, characterized in that: The graphics card interface box is also used to identify and repair faults of the general server and the graphics card resource pool.

14. The resource allocation system according to claim 12, characterized in that: The general server is further configured to provide a second instance for the computing service, wherein the second instance includes processor resources provided by the general server.

15. An electronic device, characterized in that: include: a memory storing an executable program; A processor is used to run the program, wherein the program executes the resource configuration method according to any one of claims 1 to 9 when running.

16. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the resource configuration method according to any one of claims 1 to 11.

17. A computer program product, characterized in that The method comprises a computer program, which implements the resource configuration method according to any one of claims 1 to 11 when executed by a processor.