Method and device for intelligent computing center to consume computing power by reasoning Serverless
By dynamically adjusting the number of inference units of the inference cluster in the Serverless module of the Intelligent Computing Center and using the PrefixHas algorithm for load optimization, the problems of waste of computing resources and low utilization rate are solved, and efficient computing resource management and user experience are improved.
Patent Information
- Application Number
- CN202510637659.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
AI Technical Summary
In the process of applying the computing power resources of the intelligent computing center to conduct inference of large language models, unbalanced computing power load leads to serious waste of resources and low utilization rate.
In the Serverless module of the Intelligent Computing Center, the number of expected inference units in the inference cluster is checked at a preset time, and the capacity is expanded or reduced, and the load optimization is used to ensure the dynamic scaling of computing power resources.
It realizes efficient utilization of computing power resources, avoids waste, improves the utilization rate of computing power resources and the stability of the reasoning process, and improves the user experience.
Smart Images

Figure CN120494103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and in particular to a method and device for an intelligent computing center to absorb computing power through Serverless reasoning. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.
[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] Since the emergence of intelligent computing centers, the high barrier to entry for using cloud computing power has been a challenge. When using computing resources from intelligent computing centers to reason about large language models, the computing power cannot be dynamically scaled and the computing load is poor, resulting in significant waste and low utilization. Summary of the Invention
[0008] The present invention provides a method and device for absorbing computing power of an intelligent computing center through Serverless reasoning, so as to solve the problem of unbalanced computing load, serious waste of computing power resources and low utilization of computing power resources in the process of using the computing power resources of the intelligent computing center to reason about large language models.
[0009] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0010] In a first aspect, the present invention provides a method for absorbing computing power in an intelligent computing center through Serverless inference. The method is applied to a Serverless module, which is used to manage the inference cluster of a large language model deployed in the intelligent computing center. The method for absorbing computing power includes:
[0011] Step S1: every preset check time, check whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually used by the inference cluster, and obtain a check result;
[0012] Step S2: If the check result indicates that the expected number of inference units is greater than or equal to the actual number of inference units, expand the inference cluster to increase the actual number of inference units;
[0013] Step S3: If the check result indicates that the expected number of inference units is less than the actual number of inference units, return to step S1 until the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, and scale down the inference cluster to reduce the actual number of inference units.
[0014] Step S4: using a target load algorithm to perform load optimization processing on the inference cluster; the target load algorithm includes a PrefixHas algorithm.
[0015] Optionally, step S3 includes:
[0016] Step S31: When the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, the inference cluster is scaled down after a preset scaling delay time to reduce the actual number of inference units.
[0017] In a second aspect, the present invention provides a device for absorbing computing power in an intelligent computing center through Serverless inference. The device is applied to a Serverless module. The Serverless module is used to manage the inference cluster of a large language model deployed in the intelligent computing center. The device for absorbing computing power includes:
[0018] A checking module, configured to check, at every preset checking time, whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually used by the inference cluster, and obtain a checking result;
[0019] an execution module, configured to expand the inference cluster to increase the actual number of inference units if the checking result indicates that the expected number of inference units is greater than or equal to the actual number of inference units;
[0020] The execution module is further configured to, if the check result indicates that the expected number of inference units is less than the actual number of inference units, return to the step of checking whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually applied by the inference cluster at every preset check time, and obtain the check result, until the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, and scale down the inference cluster to reduce the actual number of inference units;
[0021] The execution module is further configured to perform load optimization processing on the inference cluster using a target load algorithm; the target load algorithm includes a PrefixHas algorithm.
[0022] Optionally, the execution module is also used to scale down the inference cluster after a preset scaling-down delay time to reduce the actual number of inference units when the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold.
[0023] In a third aspect, the present invention provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method for absorbing computing power through Serverless reasoning in an intelligent computing center as described in any one of the first aspects are implemented.
[0024] In a fourth aspect, the present invention provides a readable storage medium storing a program or instruction. When the program or instruction is executed by a processor, the steps of the method for absorbing computing power by reasoning Serverless in an intelligent computing center as described in any one of the first aspects are implemented.
[0025] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the method for absorbing computing power by reasoning Serverless in an intelligent computing center as described in any one of the first aspects.
[0026] In the present invention, step S1: at every preset inspection time, check whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually applied by the inference cluster, and obtain the inspection result; step S2: if the inspection result indicates that the expected number of inference units is greater than or equal to the actual number of inference units, expand the inference cluster to increase the actual number of inference units; step S3: if the inspection result indicates that the expected number of inference units is less than the actual number of inference units, return to step S1 until the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, shrink the inference cluster to reduce the actual number of inference units, step S4: adopt the target load algorithm to perform load optimization processing on the inference cluster; the target load algorithm includes the PrefixHas algorithm. The present invention can quickly and elastically scale the computing power of the intelligent computing center in the inference process of the inference unit, route the same type of user requests to the same inference unit, avoid computing power waste, and improve the utilization rate of computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0028] Figure 1 This is a flowchart of a method for absorbing computing power through Serverless reasoning in the intelligent computing center of the present invention;
[0029] Figure 2 This is a schematic diagram of the principle of Serverless absorbing computing power;
[0030] Figure 3 This is a block diagram of the principle of the device for absorbing computing power through Serverless reasoning in the intelligent computing center of the present invention;
[0031] Figure 4 It is a principle block diagram of the electronic device of the present invention. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0033] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable under appropriate circumstances, so that the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "or" in the present invention represents at least one of the connected objects. For example, "A or B" covers three options, namely, option one: including A but not including B; option two: including B but not including A; option three: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0034] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0035] First, the technical terms involved in the present invention are briefly explained below.
[0036] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0037] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, supercomputing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP=CP general + CP intelligent + CP super.
[0038] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0039] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.
[0040] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.
[0041] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0042] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.
[0043] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0044] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.
[0045] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.
[0046] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0047] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".
[0048] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0049] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0050] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0051] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.
[0052] The "large language model" mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0053] The "Serverless" mentioned in this invention means: tasks are completed by running in stateless reasoning units triggered by events. These reasoning units may only exist in one call, or the number of reasoning units may be automatically adjusted according to the load.
[0054] The "Inference Unit" mentioned in the present invention refers to a component or module used for reasoning and logical deduction in computer science, artificial intelligence or logic.
[0055] An "inference cluster," as used in this document, refers to a group of computing resources (e.g., servers, compute nodes, and inference units) that work together to perform inference tasks, particularly in machine learning and artificial intelligence applications. Such clusters can be used to process large amounts of data, perform model inference, and execute real-time predictions.
[0056] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.
[0057] This invention provides a method for absorbing computing power of intelligent computing center through inference serverless, which is applied to serverless module. Serverless module is used to manage the inference cluster of large language model deployed in intelligent computing center. Figure 1 As shown, Figure 1 This is a flow chart of the method for absorbing computing power through Serverless reasoning in the intelligent computing center of the present invention. The method for absorbing computing power includes:
[0058] Step S1: every preset check time, check whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently used by the inference cluster, and obtain a check result;
[0059] Step S2: If the check result indicates that the expected number of inference units is greater than or equal to the actual number of inference units, the inference cluster is expanded to increase the actual number of inference units;
[0060] Step S3: If the check result indicates that the expected number of inference units is less than the actual number of inference units, return to step S1 until the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, and scale down the inference cluster to reduce the actual number of inference units;
[0061] Step S4: Use a target load algorithm to optimize the load of the inference cluster; the target load algorithm includes the PrefixHas algorithm.
[0062] It can be understood that the inference cluster can be the computing power resources allocated by the intelligent computing center to realize the inference operation of the large language model deployed in the intelligent computing center. The present invention regulates the load of the above-mentioned computing power resources by setting a Serverless module to run the method of absorbing computing power of the present invention, thereby avoiding waste of computing power resources or insufficient computing power resources and improving the operation efficiency of computing power resources. It should be noted that the expected number of inference units can refer to the number of inference units currently required to meet the preset indicators. The preset indicators, for example, the average load of the inference units in the inference cluster does not exceed the preset load threshold and the timeliness threshold of the inference results. In some optional embodiments, the inference unit may refer to a GPU (Graphics Processing Unit).
[0063] In the present invention, the check time can be set by the user based on actual needs and is not limited in this invention. For example, the check time can be 30 seconds. That is, every 30 seconds, the expected number of inference units in the inference cluster is checked to see if it is less than the actual number of inference units currently used in the inference cluster, and the check result is obtained.
[0064] In the present invention, expanding the inference cluster refers to increasing the number of actual inference units in the inference cluster; shrinking the inference cluster refers to reducing the number of actual inference units in the inference cluster.
[0065] In some embodiments of the present invention, optionally, the number of inference units that are expanded or shrunk each time can be determined by the difference between the expected number of inference units and the actual number of inference units. Exemplarily, if the expected number of inference units is greater than or equal to the actual number of inference units, and the expected number of inference units is 3 more than the actual number of inference units, then the inference cluster is expanded by adding 3 inference units. Exemplarily, if the inspection result indicates that the expected number of inference units is less than the actual number of inference units, and the expected number of inference units is 3 less than the actual number of inference units, then if the shrinking condition is met (i.e., the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units is greater than or equal to a preset number threshold), then the inference cluster is shrunk by reducing 3 inference units.
[0066] In some embodiments of the present invention, the number of inference units for each expansion or contraction can optionally be set by the user based on actual needs. For example, if the user sets the number of inference units for each expansion or contraction to 1, each expansion of the inference cluster adds one inference unit, and each contraction reduces one inference unit.
[0067] It should be noted that the number threshold can be set by the user according to actual needs, and the present invention does not impose any limitation on this.
[0068] In the present invention, the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units represents the number of times step S3 is continuously entered after step S1. For example, after executing step S1 for the first time, if the inspection result indicates that the expected number of inference units is less than the actual number of inference units, the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units is 1. Then, the process returns to step S1 and executes step S1 for the second time. If the inspection result indicates that the expected number of inference units is less than the actual number of inference units, the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units is 2. Returning to step S1 again, if step S1 is executed for the third time and the inspection result indicates that the expected number of inference units is greater than or equal to the actual number of inference units (i.e., the process enters step S2), the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units is 0. That is, because the inspection result after executing step S1 for the third time indicates that the expected number of inference units is greater than or equal to the actual number of inference units, the continuity of the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units is interrupted, and the number of times the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units is reset to zero and counted again.
[0069] In the present invention, the scaling-down condition is based on the number of times that the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units, which is greater than or equal to a preset number threshold. This can avoid the interference of occasional situations on the change of inference cluster capacity, improve the stability of the inference process, and improve the user experience of using large language models.
[0070] See also Figure 2 As shown, Figure 2 This is a schematic diagram of the principle of Serverless absorbing computing power. Each "pod" represents an inference unit, and multiple "pods" represent an inference cluster. Under the control of Serverless, capacity can be flexibly expanded and reduced, which can effectively reduce the waste of computing power resources used for inference, improve inference efficiency, and save costs.
[0071] It should be noted that the PrefixHas algorithm, also known as string prefix hashing, is a hashing technique for efficiently comparing strings. Its basic principle is to treat strings as P-based numbers, typically with P being 131 or 13331. This reduces the probability of hash collisions. The main features of the PrefixHas algorithm include: using the FNV-1a hash algorithm to calculate hash values; supporting hashing of Pod specifications and strings; securely encoding hash results to ensure URL security; and using it for Pod versioning and consistent hashing. PrefixHash is suitable for scenarios requiring session persistence, high cache hit rates, and limited computing resources. Using PrefixHash enables more accurate hash hits, improves return times, and saves computing power.
[0072] Advantages of the PrefixHash algorithm: low computational overhead, requests with the same prefix are always routed to the same endpoint, and suitable for scenarios requiring session affinity.
[0073] In the present invention, the load of the inference cluster is optimized by adopting a target load algorithm; the target load algorithm includes a PrefixHas algorithm, which can route user requests of the same type to the same inference unit to avoid wasting computing power.
[0074] In the present invention, step S1: at every preset inspection time, check whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually applied by the inference cluster, and obtain the inspection result; step S2: if the inspection result indicates that the expected number of inference units is greater than or equal to the actual number of inference units, expand the inference cluster to increase the actual number of inference units; step S3: if the inspection result indicates that the expected number of inference units is less than the actual number of inference units, return to step S1 until the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, shrink the inference cluster to reduce the actual number of inference units, step S4: adopt the target load algorithm to perform load optimization processing on the inference cluster; the target load algorithm includes the PrefixHas algorithm. The present invention can quickly and elastically scale the computing power of the intelligent computing center in the inference process of the inference unit, route the same type of user requests to the same inference unit, avoid computing power waste, and improve the utilization rate of computing power.
[0075] In some embodiments of the present invention, optionally, step S3 includes:
[0076] Step S31: When the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, the inference cluster is scaled down after a preset scaling-down delay time to reduce the actual number of inference units.
[0077] In the present invention, the scaling-down delay time can be set by the user according to actual needs, and the present invention does not impose any limitation on this.
[0078] It should be noted that if the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to the preset threshold, the scaling-down condition has been met. In this case, a preset scaling-down delay is set to avoid a sudden scaling-down that may cause a negative experience for users using large language models (for example, a significant increase in waiting time for inference results), thereby improving the user experience.
[0079] In some embodiments of the present invention, optionally, the target load algorithm may further include a Least Load algorithm.
[0080] The Least Load algorithm is a load balancing strategy used to distribute requests or tasks across multiple servers or compute nodes to achieve a more even load distribution. The basic idea is to assign new requests to the server with the least load, thereby preventing any one server from becoming overloaded while others remain idle. This approach is widely used in distributed systems, cloud computing, and network services.
[0081] The main features of LeastLoad are: using the atomic counter inFlight to track the number of requests currently being processed by each endpoint; selecting the endpoint with the smallest inFlight value as the best choice; supporting adapter filtering, selecting only endpoints with corresponding adapters; using read-write locks to ensure concurrency safety; and returning a decrement function to reduce the count when the request is completed.
[0082] Advantages of the LeastLoad algorithm: It can achieve true load balancing, dynamically adapt to the load conditions of the endpoint, and is suitable for processing scenarios with large differences in request times.
[0083] Usage scenarios:
[0084] LeastLoad is suitable for stateless services, scenarios where request processing times vary greatly, and precise load balancing is required.
[0085] targetRequest: the maximum number of requests at the same time;
[0086] In some embodiments of the present invention, the expected number of instances is optionally equivalent to the expected number of inference units. Expected number of instances = ceil(avgActiveRequests / targetRequest), i.e., ceil(average number of active requests / set number of requests);
[0087] If the expected number of instances is greater than the actual number of instances (i.e., equivalent to the actual number of inference units), the capacity will be expanded immediately;
[0088] If the expected number of instances is less than the actual number of instances, the scaling-down condition is recorded as met. If the number of consecutive checks is met (ceil(scaleDownDelaySeconds / interval)), scaling-down is initiated.
[0089] scaleDownDelaySeconds: scaling delay time;
[0090] interval: how often to check;
[0091] Number of checks: ceil(scaleDownDelaySeconds / interval) # If the scaling-down conditions are met after this many checks, then scale down.
[0092] The present invention provides a device for absorbing computing power through Serverless reasoning in an intelligent computing center, which is applied to a Serverless module. The Serverless module is used to manage the reasoning cluster of a large language model deployed in the intelligent computing center. Figure 3 As shown, Figure 3 This is a block diagram of the principle of a device for absorbing computing power by reasoning Serverless in the intelligent computing center of the present invention. The device 30 for absorbing computing power includes:
[0093] A checking module 31 is configured to check, at a preset checking time interval, whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually used by the inference cluster, and obtain a checking result;
[0094] an execution module 32 configured to expand the inference cluster to increase the actual number of inference units if the checking result indicates that the expected number of inference units is greater than or equal to the actual number of inference units;
[0095] The execution module 32 is further configured to, if the check result indicates that the expected number of inference units is less than the actual number of inference units, return to the step of checking whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually applied by the inference cluster at every preset check time, and obtain the check result, until the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, and scale down the inference cluster to reduce the actual number of inference units;
[0096] The execution module 32 is further configured to perform load optimization processing on the inference cluster using a target load algorithm; the target load algorithm includes a PrefixHas algorithm.
[0097] In some embodiments of the present invention, optionally, the execution module 32 is also used to scale down the inference cluster after a preset scaling-down delay time to reduce the actual number of inference units when the inspection result continuously indicates that the number of times the expected number of inference units is less than the actual number of inference units is greater than or equal to a preset number threshold.
[0098] The intelligent computing center provided by the present invention can achieve the following by reasoning about the serverless computing power consumption device: Figures 1 to 2 The various processes implemented by the method embodiment achieve the same technical effect and are not described here again to avoid repetition.
[0099] The present invention provides an electronic device 40, see Figure 4 As shown, Figure 4 This is a principle block diagram of the electronic device 40 of the present invention, including a processor 41, a memory 42, and a program or instruction stored in the memory 42 and executable on the processor 41. When the program or instruction is executed by the processor, any step of the method of absorbing computing power through serverless reasoning in an intelligent computing center of the present invention is implemented.
[0100] The present invention provides a readable storage medium on which programs or instructions are stored. When the programs or instructions are executed by a processor, the various processes of the embodiments of the method for absorbing computing power by reasoning Serverless in an intelligent computing center as described above are implemented, and the same technical effects can be achieved. To avoid repetition, they are not described here.
[0101] The readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0102] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the various processes of any of the above-mentioned embodiments of the method for an intelligent computing center to absorb computing power through serverless reasoning can be implemented, and the same technical effects can be achieved. To avoid repetition, they are not described here.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0104] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for absorbing computing power through Serverless reasoning in an intelligent computing center, characterized by: Applied to the Serverless module, which is used to manage the inference cluster of a large language model deployed in an intelligent computing center, the method for absorbing computing power includes: Step S1: every preset check time, check whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually used by the inference cluster, and obtain a check result; Step S2: If the check result indicates that the expected number of inference units is greater than or equal to the actual number of inference units, expand the inference cluster to increase the actual number of inference units; Step S3: If the check result indicates that the expected number of inference units is less than the actual number of inference units, return to step S1 until the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, and scale down the inference cluster to reduce the actual number of inference units. Step S4: using a target load algorithm to perform load optimization processing on the inference cluster; the target load algorithm includes a PrefixHas algorithm.
2. The method for absorbing computing power through Serverless reasoning in an intelligent computing center according to claim 1 is characterized in that: The step S3 comprises: Step S31: When the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, the inference cluster is scaled down after a preset scaling delay time to reduce the actual number of inference units.
3. A device for absorbing computing power through Serverless reasoning in an intelligent computing center, characterized by: Applied to the Serverless module, which is used to manage the inference cluster of a large language model deployed in an intelligent computing center. The device for absorbing computing power includes: A checking module, configured to check, at every preset checking time, whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually used by the inference cluster, and obtain a checking result; an execution module, configured to expand the inference cluster to increase the actual number of inference units if the checking result indicates that the expected number of inference units is greater than or equal to the actual number of inference units; The execution module is further configured to, if the check result indicates that the expected number of inference units is less than the actual number of inference units, return to the step of checking whether the expected number of inference units of the inference cluster is less than the actual number of inference units currently actually applied by the inference cluster at every preset check time, and obtain the check result, until the check result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, and scale down the inference cluster to reduce the actual number of inference units; The execution module is further configured to perform load optimization processing on the inference cluster using a target load algorithm; the target load algorithm includes a PrefixHas algorithm.
4. The device for absorbing computing power through Serverless reasoning in an intelligent computing center according to claim 3 is characterized in that: The execution module is also used to, when the inspection result continuously indicates that the expected number of inference units is less than the actual number of inference units for a number greater than or equal to a preset number threshold, scale down the inference cluster after a preset scaling-down delay time to reduce the actual number of inference units.
5. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method for absorbing computing power by reasoning Serverless in an intelligent computing center as described in any one of claims 1 to 2 are implemented.
6. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, which, when executed by a processor, implement the steps of the method for absorbing computing power by reasoning Serverless in an intelligent computing center as described in any one of claims 1 to 2.
7. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method for absorbing computing power by reasoning Serverless in an intelligent computing center as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Method and device for carrying out online optimization scheduling on AI reasoning cluster
CN117971502A
Inference service management method, equipment, medium and computer program product
CN119415273A
Customer service optimization method based on computing power implementation of intelligent computing center and related device
CN119990148A