Resource allocation method, machine learning platform, equipment and storage medium

By implementing resource allocation methods on the machine learning platform, dynamically scheduling and allocating computing resources, the problem of unbalanced resource utilization is solved and the resource utilization is improved.

CN120144328AInactive Publication Date: 2025-06-13SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510631884.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There is an imbalance in computing resource utilization in resource allocation by machine learning platforms, resulting in low resource utilization.

Method used

By implementing the resource allocation method on the machine learning platform, receiving resource scheduling requests, analyzing requests determine the target module, and calling the computing nodes in the target cluster according to the module collection relationship to achieve balance of resource calls.

Benefits of technology

The resource utilization rate is improved, and through dynamic scheduling and allocation of computing resources, the efficient utilization of resources is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144328A_ABST
    Figure CN120144328A_ABST
Patent Text Reader

Abstract

The invention provides a resource allocation method, a machine learning platform, equipment and a storage medium. The method comprises the following steps: receiving a resource scheduling request; analyzing the resource scheduling request, and determining a target module associated with the resource scheduling request; under the condition that the target module belongs to the module set, calling a computing node in a first node set in a target cluster in response to the resource scheduling request; and under the condition that the target module does not belong to the module set, in response to the resource scheduling request, calling a computing node in a second node set in the target cluster. In the embodiment of the invention, under the condition that the resource scheduling request is received, the target module associated with the resource scheduling request is determined, and then the corresponding computing resource in the target cluster is called based on the relationship between the target module and the module set, so that the balance of computing power resource calling is realized, and the resource utilization rate is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of machine learning, and particularly to a resource allocation method, a machine learning platform, a device, and a storage medium. Background Art

[0002] The general process of AI model service from requirement submission, to model development and training, and then to delivery and online operation is as follows: The data processor processes the training data, the algorithm engineer develops the machine learning model and trains the machine learning model using the training data. After the machine learning model is trained, the operation and maintenance engineer deploys the environment and deploys the trained machine learning model, and then performs online inference through the machine learning model to achieve model delivery.

[0003] The above-mentioned data processor, algorithm engineer, and operation and maintenance engineer work through their respective development platforms, and the operation of the machine learning platform requires a large amount of computing resources. This leads to cross-platform resource calls in the machine learning platform, which easily causes uneven utilization of computing resources and thus low resource utilization. Summary of the Invention

[0004] The main objective of this application is to provide a resource allocation method, a machine learning platform, a device, and a storage medium, aiming to solve the technical problem of uneven utilization of computing resources, which in turn leads to low resource utilization.

[0005] To achieve the above objective, this application provides a resource allocation method applied to a machine learning platform, where the machine learning platform includes a data processing module, a model development module, a model training module, a model deployment module, and a model inference module; The method includes: Receiving a resource scheduling request; Parsing the resource scheduling request to determine the target module associated with the resource scheduling request; the target module is any one of the data processing module, the model development module, the model training module, the model deployment module, and the model inference module; When the target module belongs to a module set, in response to the resource scheduling request, calling the computing nodes in the first node set of the target cluster; the machine learning platform is deployed on the target cluster, and the module set includes the model development module, the model training module, and the model inference module; When the target module does not belong to the module set, in response to the resource scheduling request, calling the computing nodes in the second node set of the target cluster.

[0006] Optionally, the first node set includes a development node pool, a training node pool, and an inference node pool; When the target module belongs to the module set, in response to the resource scheduling request, invoking the computing nodes in the first node set in the target cluster includes: When the target module is a model development module, in response to the resource scheduling request, invoking the computing nodes included in the development node pool; When the target module is a model training module, in response to the resource scheduling request, invoking the computing nodes included in the training node pool; When the target module is a model inference module, in response to the resource scheduling request, invoking the computing nodes included in the inference node pool.

[0007] Optionally, the second node set includes at least some of the computing nodes in the first node set.

[0008] Optionally, after receiving the resource scheduling request, the method further includes: Parsing the resource scheduling request to determine multiple model services represented by the resource scheduling request; When the computing resources scheduled for the multiple model services are less than or equal to a preset threshold, in response to the resource scheduling request, invoking 1 computing node in the target cluster.

[0009] Optionally, after receiving the resource scheduling request, the method further includes: Parsing the resource scheduling request to determine 1 model service represented by the resource scheduling request; When the computing resources scheduled for the 1 model service are greater than the preset threshold, in response to the resource scheduling request, invoking multiple computing nodes in the target cluster according to the computing resources scheduled for the 1 model service.

[0010] Optionally, the method further includes: Real-time detecting the occupancy of each computing node in the target cluster; Determining the target computing nodes that are occupied but not running in the target cluster; Releasing the target computing nodes every preset running duration.

[0011] Optionally, the method further includes: Real-time obtaining the log file of the machine learning platform; Parsing the log file; When the log file represents an abnormal operation, sending an alarm message.

[0012] In addition, to achieve the above object, the present application further provides a machine learning platform, including a data processing module, a model development module, a model training module, a model deployment module, and a model inference module; The machine learning platform is used to receive a resource scheduling request; Parse the resource scheduling request to determine the target module associated with the resource scheduling request; the target module is any one of the data processing module, the model development module, the model training module, the model deployment module, and the model inference module; When the target module belongs to the module set, in response to the resource scheduling request, call the computing nodes in the first node set of the target cluster; the machine learning platform is deployed on the target cluster, and the module set includes the model development module, the model training module, and the model inference module; When the target module does not belong to the module set, in response to the resource scheduling request, call the computing nodes in the second node set of the target cluster.

[0013] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solution: The computer device includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps of any one of the resource allocation methods proposed in the embodiments of the present application.

[0014] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution: A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the steps of any one of the resource allocation methods proposed in the embodiments of the present application.

[0015] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: The present application provides a resource allocation method, a machine learning platform, a device, and a storage medium. The above method is applied to a machine learning platform, which includes a data processing module, a model development module, a model training module, a model deployment module, and a model inference module; the above method includes: receiving a resource scheduling request; parsing the resource scheduling request to determine the target module associated with the resource scheduling request; the target module is any one of the data processing module, the model development module, the model training module, the model deployment module, and the model inference module; when the target module belongs to the module set, in response to the resource scheduling request, call the computing nodes in the first node set of the target cluster; the machine learning platform is deployed on the target cluster, and the module set includes the model development module, the model training module, and the model inference module; when the target module does not belong to the module set, in response to the resource scheduling request, call the computing nodes in the second node set of the target cluster. In the embodiments of the present application, when a resource scheduling request is received, the target module associated with the resource scheduling request is determined, and then, based on the relationship between the target module and the module set, the corresponding computing resources in the target cluster are called, so as to achieve the balance of computing power resource calls, and further improve the low resource utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is an exemplary system architecture diagram to which the present application can be applied; Figure 2 is a flowchart of the resource allocation method provided by the embodiments of the present application; Figure 3 is a schematic structural diagram of an embodiment of the machine learning platform provided by the embodiments of the present application; Figure 4 is a basic structural block diagram of a computer device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The resource allocation method provided by the embodiments of this application is applied to a machine learning platform. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0019] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0020] To enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0021] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0022] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social online platform software, etc.

[0023] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.

[0024] The server 105 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal devices 101, 102, and 103.

[0025] It should be noted that the resource allocation method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the machine learning platform is generally set in the server / terminal device.

[0026] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0027] Please refer to Figure 2 , which shows a flowchart of an embodiment of the resource allocation method proposed in the present application. The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology.

[0028] The resource allocation method provided by the embodiments of the present application is applied to a machine learning platform, which includes a data processing module, a model development module, a model training module, a model deployment module, and a model inference module; The resource allocation method includes the following steps: S210, receive a resource scheduling request.

[0029] In this step, the machine learning platform receives a resource scheduling request.

[0030] As described above, the machine learning platform includes a data processing module, a model development module, a model training module, a model deployment module, and a model inference module. The above resource scheduling request can be sent by any of the above modules.

[0031] S220, analyze the resource scheduling request to determine the target module associated with the resource scheduling request.

[0032] In this step, after receiving a resource scheduling request, parse the resource scheduling request to determine the target module associated with the resource scheduling request.

[0033] Optionally, the resource scheduling request carries a module identifier. After parsing the resource scheduling request, obtain the module identifier to determine the target module associated with the resource scheduling request.

[0034] Among them, the target module is any one of a data processing module, a model development module, a model training module, a model deployment module, and a model inference module.

[0035] For example, if the target module associated with the resource scheduling request is a data processing module, it means that the sender of the resource scheduling request is the data processing module.

[0036] S230, when the target module belongs to the module set, in response to the resource scheduling request, call the computing nodes in the first node set of the target cluster.

[0037] It should be understood that the machine learning platform is deployed on the target cluster, and the machine learning platform is pre-set with a module set, and the above module set includes a model development module, a model training module, and a model inference module.

[0038] In this step, if the target module associated with the resource scheduling request belongs to the module set, call the computing nodes in the first node set of the target cluster.

[0039] S240, when the target module does not belong to the module set, in response to the resource scheduling request, call the computing nodes in the second node set of the target cluster.

[0040] In this step, if the target module associated with the resource scheduling request does not belong to the module set, call the computing nodes in the second node set of the target cluster.

[0041] For the specific implementation manner of calling the computing nodes in response to the resource scheduling request, please refer to the subsequent embodiments.

[0042] In an embodiment of the present application, a resource scheduling request is received; the resource scheduling request is parsed to determine a target module associated with the resource scheduling request; the target module is any one of a data processing module, a model development module, a model training module, a model deployment module, and a model inference module; in the case where the target module belongs to a module set, in response to the resource scheduling request, a computing node in a first node set of a target cluster is called; a machine learning platform is deployed on the target cluster, and the module set includes a model development module, a model training module, and a model inference module; in the case where the target module does not belong to the module set, in response to the resource scheduling request, a computing node in a second node set of the target cluster is called. In the embodiment of the present application, when a resource scheduling request is received, the target module associated with the resource scheduling request is determined, and then, based on the relationship between the target module and the module set, the corresponding computing resources in the target cluster are called, so as to achieve the balance of computing power resource calls, and further improve the low resource utilization rate.

[0043] Optionally, the first node set includes a development node pool, a training node pool, and an inference node pool; The step of, in the case where the target module belongs to the module set, in response to the resource scheduling request, calling a computing node in a first node set of a target cluster includes: In the case where the target module is a model development module, in response to the resource scheduling request, call the computing nodes included in the development node pool; In the case where the target module is a model training module, in response to the resource scheduling request, call the computing nodes included in the training node pool; In the case where the target module is a model inference module, in response to the resource scheduling request, call the computing nodes included in the inference node pool.

[0044] Optionally, the second node set includes at least some of the computing nodes in the first node set.

[0045] It should be noted that the first node set includes a development node pool, a training node pool, and an inference node pool, and the computing nodes included in the development node pool, the training node pool, and the inference node pool are all different, that is, there is physical resource isolation among the development node pool, the training node pool, and the inference node pool.

[0046] The second node set includes at least some of the computing nodes in the first node set.

[0047] That is to say, in an optional implementation manner, the second node set includes some of the computing nodes in the first node set. In this implementation manner, if the target module does not belong to the module set, that is, the target module is a data processing module or a model deployment module, the computing nodes included in the development node pool, the training node pool, or the inference node pool will be called; or, the computing nodes not included in the development node pool, the training node pool, and the inference node pool will be called.

[0048] In another optional implementation manner, the second node set is the same as the first node set. In this implementation manner, if the target module does not belong to the module set, that is, the target module is a data processing module or a model deployment module, the computing nodes included in the development node pool, the training node pool, or the inference node pool will also be called.

[0049] It should be understood that the first node set includes the development node pool, the training node pool, and the inference node pool.

[0050] If the target module is a model development module, in response to the resource scheduling request, the computing nodes included in the development node pool will be called; if the target module is a model training module, in response to the resource scheduling request, the computing nodes included in the training node pool will be called; if the target module is a model inference module, in response to the resource scheduling request, the computing nodes included in the inference node pool will be called.

[0051] In this embodiment, different node pools are divided for the model development module, the model training module, and the model inference module, and for the resource scheduling requests sent by the model development module, the model training module, and the model inference module, the computing nodes in the corresponding node pools are scheduled, which not only realizes the overall management of computing power resources but also ensures the physical isolation of resources in the development, training, and inference links.

[0052] Optionally, after receiving the resource scheduling request, the method further includes: Parsing the resource scheduling request to determine multiple model services represented by the resource scheduling request; In the case where the computing resources scheduled for the multiple model services are less than or equal to a preset threshold, in response to the resource scheduling request, 1 computing node in the target cluster is called.

[0053] A possible situation is that the resource scheduling request may represent multiple model services.

[0054] In this embodiment, the resource scheduling request is parsed to determine multiple model services represented by the resource scheduling request, and the computing resources scheduled for the multiple model services are determined. In the case where the computing resources scheduled for the multiple model services are less than or equal to a preset threshold, it means that calling 1 computing node can support the scheduling of multiple model services, so in response to the resource scheduling request, 1 computing node in the target cluster is called.

[0055] In this embodiment, for the application scenario of the business small model, the machine learning platform realizes fine-grained computing resource management and allocation, supports calling 1 computing node to process multiple services. Optionally, the above computing node is a 128Mi video memory unit, and the above video memory unit can be a graphics card.

[0056] Optionally, after receiving the resource scheduling request, the method further includes: Analyze the resource scheduling request to determine 1 model service represented by the resource scheduling request; In the case where the computing resources scheduled for the 1 model service are greater than a preset threshold, in response to the resource scheduling request, call multiple computing nodes in the target cluster according to the computing resources scheduled for the 1 model service.

[0057] Another possible situation is that the resource scheduling request can represent 1 model service.

[0058] In this embodiment, analyze the resource scheduling request to determine 1 model service represented by the resource scheduling request, determine the computing resources scheduled for the 1 model service. In the case where the computing resources scheduled for the 1 model service are greater than a preset threshold, it means that multiple computing nodes need to be called to support the scheduling of this model service, then in response to the resource scheduling request, call multiple computing nodes in the target cluster.

[0059] In this embodiment, for the application scenario of the large model, when the computing resources required to schedule 1 model service are relatively large, the machine learning platform supports calling multiple computing nodes. Optionally, the above computing nodes can be graphics cards.

[0060] Optionally, the method further includes: Real-time detect the occupancy of each computing node in the target cluster; Determine the target computing nodes that are occupied but not running in the target cluster; Release the target computing nodes every preset running duration.

[0061] In this embodiment, the machine learning platform can also real-time detect the occupancy of each computing node in the target cluster and determine the target computing nodes that are occupied but not running in the target cluster. For the above target computing nodes, release the target computing nodes every preset running duration.

[0062] In this embodiment, for the computing nodes that are idle and not released in time, set a preset running duration. In the case where it has not been used after the preset running duration has elapsed, release the above computing nodes, so as to avoid waste of computing resources and improve the utilization rate of computing resources.

[0063] Optionally, the method further includes: Obtain the log file of the machine learning platform in real time; Parse the log file; In the case where the log file indicates an abnormal operation, send an alarm message.

[0064] In this embodiment, the machine learning platform can also obtain the log file in real time, parse the above log file, and if the log file indicates an abnormal operation of the machine learning platform, an alarm message can be sent externally.

[0065] Specifically, the machine learning platform can quickly perceive abnormal problems and send alarm messages in a timely manner by uniformly collecting and analyzing the monitoring and log data of the cluster, model service, and gateway system, ensuring that problems are perceived and processed in a timely manner.

[0066] Please refer to Figure 3 , a machine learning platform 300 provided by an embodiment of the present application. The machine learning platform 300 includes a data processing module 310, a model development module 320, a model training module 330, a model deployment module 340, and a model inference module 350; The machine learning platform is used to receive a resource scheduling request; Parse the resource scheduling request to determine the target module associated with the resource scheduling request; the target module is any one of the data processing module 310, the model development module 320, the model training module 330, the model deployment module 340, and the model inference module 350; In the case where the target module belongs to the module set, in response to the resource scheduling request, call the computing nodes in the first node set of the target cluster; the machine learning platform is deployed on the target cluster, and the module set includes the model development module 320, the model training module 330, and the model inference module 350; In the case where the target module does not belong to the module set, in response to the resource scheduling request, call the computing nodes in the second node set of the target cluster.

[0067] Optionally, the first node set includes a development node pool, a training node pool, and an inference node pool; The machine learning platform 300 is further used for: In the case where the target module is the model development module 320, in response to the resource scheduling request, call the computing nodes included in the development node pool; In the case where the target module is the model training module 330, in response to the resource scheduling request, call the computing nodes included in the training node pool; When the target module is the model inference module 350, in response to the resource scheduling request, call the computing nodes included in the inference node pool.

[0068] Optionally, the second node set includes at least some of the computing nodes in the first node set.

[0069] Optionally, the machine learning platform 300 is further configured to: Parse the resource scheduling request to determine multiple model services represented by the resource scheduling request; When the computing resources scheduled for the multiple model services are less than or equal to a preset threshold, in response to the resource scheduling request, call 1 computing node in the target cluster.

[0070] Optionally, the machine learning platform 300 is further configured to: Parse the resource scheduling request to determine 1 model service represented by the resource scheduling request; When the computing resources scheduled for the 1 model service are greater than the preset threshold, in response to the resource scheduling request, call multiple computing nodes in the target cluster according to the computing resources scheduled for the 1 model service.

[0071] Optionally, the machine learning platform 300 is further configured to: Real-time detect the occupancy of each computing node in the target cluster; Determine the target computing nodes that are occupied but not running in the target cluster; Release the target computing nodes every preset running duration.

[0072] Optionally, the machine learning platform 300 is further configured to: Real-time obtain the log file of the machine learning platform 300; Parse the log file; When the log file indicates an abnormal operation, send an alarm message.

[0073] To solve the above technical problems, an embodiment of the present application further provides a computer device. Specifically, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment. The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0074] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0075] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as the program code of the resource allocation method. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.

[0076] In some embodiments, the processor 42 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the program code stored in the memory 41 or process data, such as running the program code of the resource allocation method.

[0077] The network interface 43 may include a wireless network interface or a wired network interface, and the network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0078] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing the resource allocation program, and the resource allocation program can be executed by at least one processor to enable the at least one processor to execute the steps of the resource allocation method as described above.

[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general-purpose hardware online platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0080] The present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0081] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are shown in the drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure made by using the content of this application's specification and drawings, directly or indirectly applied in other related technical fields, is equally within the scope of patent protection of this application.

Claims

1. A resource allocation method, characterized in that: The method is applied to a machine learning platform, which includes a data processing module, a model development module, a model training module, a model deployment module, and a model reasoning module; The method comprises: receiving resource scheduling requests; Parsing the resource scheduling request, and determining a target module associated with the resource scheduling request; the target module is any one of the data processing module, the model development module, the model training module, the model deployment module, and the model reasoning module; In the case where the target module belongs to a module set, in response to the resource scheduling request, calling a computing node in a first node set in a target cluster; the target cluster is deployed with the machine learning platform, and the module set includes the model development module, the model training module, and the model reasoning module; In a case where the target module does not belong to the module set, in response to the resource scheduling request, a computing node in the second node set in the target cluster is called.

2. The method according to claim 1, characterized in that The first node set includes a development node pool, a training node pool and an inference node pool; When the target module belongs to a module set, in response to the resource scheduling request, calling a computing node in a first node set in the target cluster includes: In the case where the target module is a model development module, in response to the resource scheduling request, calling a computing node included in the development node pool; In a case where the target module is a model training module, in response to the resource scheduling request, calling a computing node included in the training node pool; In the case where the target module is a model reasoning module, in response to the resource scheduling request, a computing node included in the reasoning node pool is called.

3. The method according to claim 1, characterized in that The second node set includes at least some of the computing nodes in the first node set.

4. The method according to claim 1, characterized in that: After receiving the resource scheduling request, the method further includes: Parsing the resource scheduling request, and determining a plurality of model services represented by the resource scheduling request; In a case where the computing resources scheduled by the multiple model services are less than or equal to a preset threshold, in response to the resource scheduling request, one computing node in the target cluster is called.

5. The method according to claim 1, characterized in that After receiving the resource scheduling request, the method further includes: Parsing the resource scheduling request, and determining a model service represented by the resource scheduling request; In a case where the computing resources scheduled by the one model service are greater than a preset threshold, in response to the resource scheduling request, multiple computing nodes in the target cluster are called according to the computing resources scheduled by the one model service.

6. The method according to claim 1, characterized in that The method further comprises: Real-time detection of the occupancy of each computing node in the target cluster; Determine an occupied but unused target computing node in the target cluster; The target computing node is released at intervals of a preset running time.

7. The method according to claim 1, characterized in that The method further comprises: Obtaining log files of the machine learning platform in real time; Parsing the log file; When the log file indicates abnormal operation, an alarm message is sent.

8. A machine learning platform, characterized in that: It includes data processing module, model development module, model training module, model deployment module and model reasoning module; The machine learning platform is used to receive a resource scheduling request; Parsing the resource scheduling request, and determining a target module associated with the resource scheduling request; The target module is any one of the data processing module, the model development module, the model training module, the model deployment module and the model reasoning module; In the case where the target module belongs to a module set, in response to the resource scheduling request, calling a computing node in a first node set in a target cluster; the target cluster is deployed with the machine learning platform, and the module set includes the model development module, the model training module, and the model reasoning module; In a case where the target module does not belong to the module set, in response to the resource scheduling request, a computing node in the second node set in the target cluster is called.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the resource allocation method according to any one of claims 1 to 5 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the resource allocation method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Implementation method, device and equipment for integration of training and reasoning, storage medium and product

    CN118796465A

  • Model reasoning scheduling method and device and server cluster

    CN118897736A

  • Intelligent computing power scheduling method and device for intelligent common computing power computing center

    CN119938340A