Intelligent adaptive AI large model dynamic activation and scheduling method and device, and storage medium

Through the dynamic activation and scheduling method of intelligent adaptive AI big model, the number of big model service instances and request distribution strategy are dynamically adjusted, and the problem of resource waste in AI big model services in a low load state is solved, achieving more efficient resource utilization and request processing.

CN119987978AInactive Publication Date: 2025-05-13SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Patent Information

Application Number
CN202510464951.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN119987978A_ABST
    Figure CN119987978A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent adaptive AI large model dynamic activation and scheduling method and device and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: receiving an external request through a request router, and determining the routing direction of the external request according to the access state of a large model activator; if the access state of the large model activator is open, routing the external request to the large model activator, and caching the external request through the large model activator; based on the number of the external requests cached in the large model activator, triggering the automatic expansion device to adjust the number of large model service instances; and distributing the external request cached in the large model activator to each large model service instance for processing through a flow distribution algorithm. The technical effect of improving the resource utilization efficiency of the AI large model service is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device and storage medium for dynamic activation and scheduling of intelligent adaptive AI large models. Background Art

[0002] In the current AI large model service deployment and management framework, in order to ensure the availability of the service and to be able to flexibly scale resources as needed, large model services usually require at least one instance to be running. However, when the system is under low load or has no user traffic, the at least one model instance that is running continuously does not actually play its due processing role, but occupies computing resources and memory space, resulting in unnecessary waste of resources.

[0003] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0004] The main purpose of this application is to provide an intelligent adaptive AI large model dynamic activation and scheduling method, device and storage medium, aiming to solve the technical problem of low resource utilization efficiency of AI large model services.

[0005] To achieve the above objectives, the present application proposes an intelligent adaptive AI large model dynamic activation and scheduling method, which is applied to an AI large model service operation platform. The AI ​​large model service operation platform includes a request router, a large model activator, an automatic scaler, a Kubernetes API server, a request queue agent, and a large model service. The intelligent adaptive AI large model dynamic activation and scheduling method includes: Receiving an external request through the request router, and determining a routing direction of the external request according to a path state of the large model activator; If the access state of the large model activator is open, routing the external request to the large model activator, and caching the external request through the large model activator; Based on the number of the external requests cached in the large model activator, trigger the automatic scaler to adjust the number of large model service instances; Through the traffic distribution algorithm, the external requests cached in the large model activator are distributed to each of the large model service instances for processing.

[0006] In one embodiment, the step of receiving the external request through the request router and determining the routing direction of the external request according to the path state of the large model activator further includes: If the access state of the large model activator is closed, the external request is directly routed to the large model service for processing.

[0007] In one embodiment, before the step of receiving the external request through the request router and determining the routing direction of the external request according to the path state of the large model activator, the step includes: If the number of the large model service instances is not zero, marking the access state of the large model activator as closed; If the large model service does not receive a request within a preset time window, the number of instances of the large model service is reduced by the automatic scaler until it is reduced to zero; If the number of large model service instances is zero, the channel state of the large model activator is marked as open.

[0008] In one embodiment, the step of triggering the automatic scaler to adjust the number of large model service instances based on the number of external requests cached in the large model activator includes: If the cache quantity of the external request reaches the cache threshold, a trigger signal is sent to the automatic scaler through the large model activator to instruct the automatic scaler to increase the number of large model service instances; The trigger signal is received by the automatic scaler, and the demand quantity of the large model service is calculated according to the cache quantity; Based on the demand quantity, the autoscaler interacts with the Kubernetes API server to adjust the number of large model service instances to reach the demand quantity.

[0009] In one embodiment, the step of triggering the autoscaler to adjust the number of large model service instances based on the number of external requests cached in the large model activator further includes: Monitor the number of requests for inference in the large model service as the number of concurrent requests; Based on the number of concurrent requests and the processing capacity of the large model service instance, the number of the large model service instances required to process the request is evaluated by the automatic scaler as the demand quantity; According to the demand quantity, the number of the large model service instances is adjusted through the Kubernetes API server.

[0010] In one embodiment, the step of distributing the external request cached in the large model activator to each of the large model service instances for processing by using a traffic distribution algorithm includes: Monitoring the load of each of the large model service instances; Selecting a target large model service instance according to the load and processing capacity of the large model service instance by using the traffic distribution algorithm; The external request is forwarded to the target large model service instance for processing through the request queue agent.

[0011] In one embodiment, the step of selecting a target large model service instance according to the load and processing capacity of the large model service instance through the traffic distribution algorithm includes: Monitor the number of requests being processed by each of the large model service instances, evaluate the load of the large model service instances, sort the large model service instances from low to high according to the load, and determine the load sequence; Analyze the processing capabilities of each of the large model service instances, and on the basis of the load sequence, sort the large model service instances from strong to weak according to the processing capabilities to determine a comprehensive queue; Based on the comprehensive queue, the large model service instance with the highest ranking is selected as the target large model service instance.

[0012] In one embodiment, the step of forwarding the external request to the target large model service instance for processing through the request queue proxy further includes: Receiving the external request through the request queue agent and evaluating the processing capability of each of the large model service instances; If the processing capacity of each of the large model service instances has reached the upper limit, or the external request cannot be processed immediately, the external request is placed in the buffer area of ​​the request queue agent and sorted according to a preset priority; According to the result of the sorting and the load condition of the large model service instance, the external requests in the cache area are distributed to the large model service in sequence for processing.

[0013] In addition, to achieve the above-mentioned objectives, the present application also proposes an intelligent adaptive AI large model dynamic activation and scheduling device, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the intelligent adaptive AI large model dynamic activation and scheduling method as described above.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the intelligent adaptive AI large model dynamic activation and scheduling method as described above are implemented.

[0015] The present application provides an intelligent adaptive AI large model dynamic activation and scheduling method. The present application receives external requests through a request router, and determines the routing direction of the external request according to the path state of the large model activator; if the path state of the large model activator is open, the external request is routed to the large model activator, and the external request is cached by the large model activator; based on the number of external requests cached in the large model activator, the automatic scaler is triggered to adjust the number of large model service instances; through the traffic distribution algorithm, the external requests cached in the large model activator are distributed to each large model service instance for processing. The present application first activates the service instance only when needed through a dynamic routing mechanism to reduce unnecessary resource occupation. By monitoring the number of external requests cached in the large model activator, the automatic scaler is triggered to adjust the number of large model service instances. During periods of no or low traffic, the number of running instances can be reduced, thereby releasing unnecessary computing resources and memory space. Through the traffic distribution algorithm, the cached requests are intelligently distributed to each running large model service for processing, which helps to achieve load balancing, avoid overloading of a single service, improve request processing efficiency, and then realize on-demand allocation of resources, thereby improving resource utilization. This application achieves the technical effect of improving the resource utilization efficiency of AI large model services. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0018] Figure 1 A flow chart of the first embodiment of the method for dynamic activation and scheduling of an intelligent adaptive AI large model provided in this application; Figure 2 A flow chart of the second embodiment of the method for dynamic activation and scheduling of an intelligent adaptive AI large model provided in this application; Figure 3 A flow chart of the third embodiment of the method for dynamic activation and scheduling of an intelligent adaptive AI large model provided in this application; Figure 4 A flow chart of the fourth embodiment of the method for dynamic activation and scheduling of an intelligent adaptive AI large model provided in this application; Figure 5 A flow chart of the fourth embodiment of the method for dynamic activation and scheduling of an intelligent adaptive AI large model provided in this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent adaptive AI large model dynamic activation and scheduling method in the embodiment of the present application.

[0019] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0022] The main solutions of the embodiments of this application are: At present, in the AI ​​large model service deployment and management framework, in order to ensure the availability of the service and to be able to flexibly scale resources as needed, large model services usually require at least one instance to be running. However, when the system is under low load or has no user traffic, at least one model instance that is running continuously does not actually play its due processing role, but occupies computing resources and memory space, resulting in unnecessary waste of resources.

[0023] This application uses a dynamic routing mechanism to activate service instances only when needed, reducing unnecessary resource usage. By monitoring the number of external requests cached in the large model activator, the automatic scaler is triggered to adjust the number of large model service instances. During periods of no or low traffic, the number of running instances can be reduced, thereby freeing up unnecessary computing resources and memory space. Through the traffic distribution algorithm, cached requests are intelligently distributed to each running large model service for processing, which helps to achieve load balancing, avoid overloading of a single service, improve request processing efficiency, and then realize on-demand allocation of resources, thereby improving resource utilization.

[0024] It should be noted that the execution subject of this embodiment can be an intelligent adaptive AI large model dynamic activation and scheduling system, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a control device of an intelligent adaptive AI large model dynamic activation and scheduling system that can realize the above functions, etc. This embodiment does not specifically limit this. The following takes the intelligent adaptive AI large model dynamic activation and scheduling system as an example of the execution subject to illustrate this embodiment and the following embodiments.

[0025] Embodiment 1 Based on this, this application proposes a first embodiment of an intelligent adaptive AI large model dynamic activation and scheduling method, please refer to Figure 1 , the intelligent adaptive AI large model dynamic activation and scheduling method includes: Step S10, receiving an external request through the request router, and determining a routing direction of the external request according to the path state of the large model activator.

[0026] Through the dynamic routing mechanism, resource allocation is optimized to reduce unnecessary resource usage and ensure that large model service instances are activated when actually needed, thus avoiding resource waste.

[0027] In this embodiment, the request router is a component responsible for receiving, parsing and routing external requests. The request router decides whether to route the request to the activator or directly to the large model service according to the access status of the large model activator. The access status refers to whether the large model activator can currently accept and process the request.

[0028] In this embodiment, the large model activator is a component that manages the activation and deactivation status of the large model service. The large model activator is mainly responsible for caching and managing requests to these large model service instances when the large model service has not been started or has been scaled down to zero. The large model activator maintains real-time communication with the autoscaler through a WebSocket persistent connection and reports request indicators to support scaling decisions.

[0029] It should be noted that a service instance refers to a specific running copy of a large model service that can independently process requests. A service can have multiple instances. A service instance is the execution unit of the large model service and is responsible for the actual computing tasks.

[0030] It should be noted that external requests are requests from users or clients, which contain data and / or instructions that need to be processed by the AI ​​large model.

[0031] As an optional implementation, the request router determines the routing direction of the request based on the path status of the large model activator. When the large model service instance has not been started or the load is too high, the request is sent to the large model activator for caching.

[0032] Optionally, step S10 further includes: Step S11, if the access state of the large model activator is closed, the external request is directly routed to the large model service for processing.

[0033] In this embodiment, the big model service is a service that executes specific business logic, such as processing large model tasks such as natural language processing and image recognition. It requires more computing resources and can be dynamically expanded according to demand.

[0034] It should be noted that the access state is a state of the large model activator, indicating whether there is an available large model service instance to process the request. When the access state of the large model activator is "on", it means that there is no available large model service instance, and the request is routed to the activator; when the access state of the activator is "off", it means that there is an available large model service instance, and the request is directly routed to the large model service.

[0035] Exemplarily, when the request router detects that the access state of the large model activator is closed, the external request is directly routed to the large model service for processing to improve the response speed.

[0036] Optionally, before routing the request to the large model service, the request router performs a health check on each service instance to ensure that each service instance is in an available state. If the service instance is detected to be unhealthy, the request router removes the service instance from the routing list and routes the request to other healthy service instances.

[0037] Step S20: If the access state of the large model activator is on, the external request is routed to the large model activator, and the external request is cached by the large model activator.

[0038] When the large model service instance is unavailable or the system load is too high, the large model activator caches and manages the request to avoid request loss or service denial. By caching requests instead of continuously running service instances, resource waste is reduced and resources can be allocated on demand.

[0039] In this embodiment, caching is to temporarily store requests in memory or other fast-access storage media so that they can be processed quickly when resources are available.

[0040] As an optional implementation, when the access state of the large model activator is detected to be open, the request router routes the external request to the large model activator. After receiving the request, the large model activator caches the request until the service instance is ready and can be expanded through the autoscaler.

[0041] Optionally, set a timeout for external requests in the cache and retry after the timeout to avoid requests being lost due to long waits.

[0042] Step S30: triggering the automatic scaler to adjust the number of large model service instances based on the number of external requests cached in the large model activator.

[0043] In this embodiment, the large model service instance is the execution unit of the large model service and is responsible for the actual computing tasks.

[0044] Optionally, a threshold for triggering the scaling operation is set. When the number of requests in the cache exceeds an upper preset value, the scaling operation is triggered; when the number of requests is lower than a lower preset value, the scaling operation is triggered.

[0045] As an optional implementation, when the number of requests in the cache reaches or exceeds the expansion threshold, the large model activator sends a signal to the autoscaler to request the startup of a new large model service instance. When the number of requests is lower than the reduction threshold, a signal is sent to request the stop of the redundant service instances. After receiving the expansion request, the autoscaler adjusts the number of instances of the large model service through the Kubernetes API server according to the received signal and the preset strategy.

[0046] Step S40: Distribute the external requests cached in the large model activator to each of the large model service instances for processing through a traffic distribution algorithm.

[0047] The external requests cached in the large model activator are effectively distributed to each large model service instance for processing, ensuring that the requests can be processed in a timely manner, while optimizing the use of resources and improving the throughput and response speed of the system. By intelligently distributing requests, it is possible to avoid the situation where some service instances are overloaded while other instances are idle, thereby achieving load balancing.

[0048] In this embodiment, the traffic distribution algorithm is a method for deciding how to distribute network traffic or requests to multiple processing units, and decisions can be made based on the priority of the request, the current load of the server, the response time requirement of the request, etc.

[0049] As an optional implementation, a polling distribution method is used to distribute the requests in the cache to each large model service instance in sequence. When all service instances have been allocated a request, the distribution starts from the beginning.

[0050] As another optional implementation, a method of least connection distribution is adopted to monitor the current number of connections of each large model service instance, that is, the number of requests being processed, and allocate new requests to the service instance with the least number of connections.

[0051] As another optional implementation, a weighted round-robin distribution method is used to assign a weight to each large model service instance. A service with a higher weight receives more requests. The requests in the cache are sequentially assigned to each large model service instance according to the weight.

[0052] Optionally, the load, processing capacity, and health status of each large model service instance are monitored in real time, and the traffic distribution strategy is dynamically adjusted to adapt to load changes and optimize resource usage.

[0053] This embodiment provides an intelligent adaptive AI large model dynamic activation and scheduling method. This embodiment first activates service instances only when needed through a dynamic routing mechanism to reduce unnecessary resource usage. By monitoring the number of external requests cached in the large model activator, the automatic scaler is triggered to adjust the number of large model service instances. During periods of no or low traffic, the number of running instances can be reduced, thereby releasing unnecessary computing resources and memory space. Through the traffic distribution algorithm, the cached requests are intelligently distributed to each running large model service for processing, which helps to achieve load balancing, avoid overloading of a single service, improve request processing efficiency, and then realize on-demand allocation of resources, thereby improving resource utilization.

[0054] Based on Example 1, Example 2 of the present application proposes a method for dynamic activation and scheduling of an intelligent adaptive AI large model, referring to Figure 2 , before step S10, including: Step S50: If the number of the large model service instances is not zero, the channel state of the large model activator is marked as closed.

[0055] When large model service instances are available, ensure that external requests can be directly routed to these service instances for processing instead of entering the cache queue of the large model activator, improve request processing efficiency, reduce processing delays, and ensure that resources are effectively used.

[0056] It should be noted that a large model service instance refers to a running copy of the large model service, which can independently receive requests and process them. It can be a physical server, virtual machine, or container, depending on how the service is deployed.

[0057] As an optional implementation, when it is detected that the number of large model service instances is not zero, based on the corresponding API interface (Application Programming Interface) provided by the large model activator, the access state of the large model activator is marked as closed through an API call.

[0058] As another optional implementation, a large model activator state table is maintained in the system database, which contains the access state information of the large model activator. When a large model service instance is detected, the access state of the large model activator is set to closed through a database update operation.

[0059] Step S60: If the large model service does not receive any request within the preset time window, the number of large model service instances is reduced by the automatic scaler until it is reduced to zero.

[0060] When the large model service does not receive a request within the preset time window, the automatic scaler reduces the number of service instances until it is reduced to zero, avoiding the need to maintain unnecessary computing resources during low-demand periods, thereby reducing resource waste and achieving optimal resource utilization and cost savings.

[0061] In this embodiment, the autoscaler is the core component for achieving automatic expansion and contraction. The autoscaler decides whether to increase or decrease the number of service instances based on real-time request metrics. In the absence of requests, the autoscaler can reduce the number of service instances to zero to achieve on-demand resource allocation and further save resources. The autoscaler helps ensure that resources are used efficiently, avoiding wasting resources during low loads while providing sufficient processing power during high loads.

[0062] It should be noted that the preset time window is a pre-set time period used to monitor whether the large model service receives requests. If no request is received during this time period, the automatic scaling mechanism will be triggered. The large model service instance is a specific running entity of the large model service, which occupies a certain amount of computing resources and can process requests from clients.

[0063] As an optional implementation, use a load monitoring tool to track the request volume and response time of the large model service in real time. If the monitoring data shows that the request volume is continuously zero or extremely low within the preset time window, the autoscaler will reduce the number of service instances according to the preset rules.

[0064] As another optional implementation, a time trigger is set to regularly check the request log of the large model service within a preset time window. If there is no request record within the time period set by the time trigger, the autoscaler is triggered to reduce one or more service instances. This process is repeated until the number of instances is reduced to zero.

[0065] Step S70: If the number of the large model service instances is zero, the channel state of the large model activator is marked as open.

[0066] Ensure that the system can receive and cache new external requests when no large model service instances are running. When the number of large model service instances decreases to zero, that is, when the system has no request load at all within a certain time window, opening the large model activator's access state allows the system to respond quickly when requests come again.

[0067] It should be noted that the number of large model service instances refers to the number of large model service instances currently running. A zero instance number means that no service instance is processing requests.

[0068] As an optional implementation, when it is monitored that the number of large model service instances is zero, a control signal is sent to the large model activator to instruct the large model activator to mark the channel state as open.

[0069] As another optional implementation, the API interface of the large model activator is called to directly change the channel state to open.

[0070] This embodiment provides an intelligent adaptive AI large model dynamic activation and scheduling method. This embodiment first dynamically adjusts the number of large model service instances to allocate resources according to actual needs and avoid wasting resources at low loads. When the load suddenly increases, new service instances can be quickly started to process requests and improve the response speed of the system. By dynamically adjusting the number of instances and implementing a load balancing strategy, requests can be distributed more evenly, improving overall stability and reliability.

[0071] Based on Example 1, Example 3 of the present application proposes a method for dynamic activation and scheduling of an intelligent adaptive AI large model, referring to Figure 3 , step S30 comprises: Step S31: If the cache quantity of the external request reaches the cache threshold, a trigger signal is sent to the automatic scaler through the large model activator to instruct the automatic scaler to increase the number of large model service instances.

[0072] When the number of external request caches received by the large model service reaches the preset cache threshold, a trigger signal is sent to the automatic scaler through the large model activator, instructing it to increase the number of large model service instances, ensuring that the system can respond to the high concurrency requirements of external requests in a timely manner and maintain service stability by dynamically adjusting the number of service instances.

[0073] It should be noted that the cache threshold is the maximum number of requests that can be cached in the large model activator. If this value is exceeded, the automatic scaler needs to be triggered to expand the capacity.

[0074] As an optional implementation, based on event-driven triggering, when the cache quantity of external requests reaches or exceeds the cache threshold, an expansion event is triggered. The large model activator detects the event and sends a trigger signal to the automatic scaler.

[0075] As another optional implementation, a timer is set to periodically check the cache quantity of external requests in the large model activator. If the cache quantity reaches or exceeds the cache threshold, the large model activator immediately sends a trigger signal to the automatic scaler.

[0076] Step S32: The trigger signal is received by the automatic scaler, and the demand quantity of the large model service is calculated according to the cache quantity.

[0077] After receiving the trigger signal, it determines how many large model service instances need to be added to meet the current request load. By calculating the required number, the autoscaler can accurately adjust the number of service instances to ensure that the system can effectively handle the backlog of requests while avoiding resource waste caused by over-scaling.

[0078] It should be noted that the trigger signal is a signal sent by the large model activator to the autoscaler, indicating that the service instance needs to be increased to handle more requests. The demand quantity is based on the current load and system configuration, and calculates the number of service instances that need to be started to handle current and expected future requests.

[0079] As an optional implementation, the cache quantity is mapped to the required service instance quantity according to a preset mapping table.

[0080] For example, if each service instance can process 100 requests per second, and there are 1,000 requests waiting to be processed in the current cache, at least 10 service instances are required to meet the demand.

[0081] As another optional implementation, a machine learning model is trained based on historical request data, the number of requests in a future period is predicted by the machine learning model, and the required number of service instances is calculated based on the prediction results.

[0082] Optionally, when training a machine learning model, first collect historical request data, including information such as timestamps, quantities, and types, and ensure the time series of the data, arranging them in chronological order. Extract time series trends, periodicity, seasonality, and other features of the number of requests in the historical request data. Split the data set into a training set and a test set, and use the training set data to train the machine learning model. During the training process, the model adjusts its parameters to minimize the prediction error, and uses cross-validation to find the optimal hyperparameter combination. Use the test set data to evaluate the performance of the model, and adjust and optimize the model based on the evaluation results.

[0083] Optionally, when selecting a machine learning algorithm, you can select algorithms such as ARIMA (Autoregressive Integrated Moving Average Model) and LSTM (Long Short-Term Memory) that are suitable for analyzing time series data and predicting future trends to capture the time dependency and periodicity in the data. You can also select algorithms such as linear regression, polynomial regression, ridge regression, and lasso regression that are suitable for predicting continuous values ​​to predict the number of future requests by fitting a function whose parameters are optimized through training data.

[0084] Step S33: Based on the demand quantity, the automatic scaler interacts with the Kubernetes API server to adjust the number of large model service instances to reach the demand quantity.

[0085] It should be noted that the autoscaler can dynamically increase or decrease the number of service instances according to preset rules and conditions. It is responsible for receiving trigger signals, calculating the required number, and interacting with the Kubernetes API server to adjust the number of service instances. The Kubernetes API server is the control plane component of the Kubernetes cluster, responsible for processing API requests related to the creation, deletion, status update, etc. of large model container instances. All interactions with the Kubernetes cluster are carried out through the Kubernetes API server.

[0086] As an optional implementation, the autoscaler calculates the number of large model service instances that need to be increased or decreased based on the demand quantity. The autoscaler interacts with the Kubernetes API server and sends a request to adjust the number of large model service instances. After receiving the request from the autoscaler, the Kubernetes API server adjusts the number of replicas of the corresponding Deployment to meet the demand quantity.

[0087] This embodiment provides an intelligent adaptive AI large model dynamic activation and scheduling method. This embodiment first realizes the automatic adjustment of the number of large model service instances through the collaborative work of the large model activator, the automatic scaler and the Kubernetes API server, ensuring that the system can maintain the stability of the service under high concurrent requests while avoiding unnecessary waste of resources.

[0088] Based on Example 1, Example 4 of the present application proposes a method for dynamic activation and scheduling of an intelligent adaptive AI large model, referring to Figure 4 , step S30 further includes: Step S34, monitoring the number of requests for reasoning in the large model service as the number of concurrent requests.

[0089] Monitor the number of inference requests being processed in the large model service in real time to understand the current service load and determine whether the current service instance can handle the existing request load or whether more instances need to be added to meet the demand.

[0090] It should be noted that the number of inference requests is the number of requests sent to the large model service, requiring the model to perform calculations and return results. The number of concurrent requests is the number of requests sent to the service in the same time period, which is used to evaluate the load of the service.

[0091] As an optional implementation, the request data of the large model service is collected through a monitoring tool, key performance indicators that need to be monitored are defined, the request data is collected regularly, and analysis is performed to determine the number of concurrent requests for the service.

[0092] Step S35: Based on the number of concurrent requests and the processing capacity of the large model service instance, the automatic scaler is used to evaluate the number of the large model service instances required to process the requests as the demand number.

[0093] It should be noted that the processing capacity of a large model service instance refers to the number of inference requests that a single service instance can handle per unit time, which depends on factors such as the complexity of the model, the allocation of computing resources, and the degree of optimization.

[0094] As an optional implementation, the maximum number of concurrent processing of a single large model service instance is estimated through benchmark testing or historical data analysis as processing capacity. The required number of service instances is calculated based on the number of concurrent requests and the processing capacity of a single instance. The calculation formula is: "Required number = total number of concurrent requests / (maximum number of concurrent requests for a single large model service instance × target utilization rate)".

[0095] Optionally, calculate the processing capacity of a single large model service instance through benchmarking. Prepare input data for the benchmark based on the main application scenarios of the large model service. Deploy a single large model service instance in the target environment and ensure that its configuration is consistent with expectations. Use the benchmarking tool to send requests to the service instance. Determine the number and rate of requests based on business needs and expected load. Calculate the maximum number of concurrent requests processed by a single large model service instance through the throughput data obtained from the benchmark, and then estimate the processing capacity of the service instance under load.

[0096] Step S36: According to the demand quantity, the number of the large model service instances is adjusted through the Kubernetes API server.

[0097] After determining the required quantity, the number of large model service instances is dynamically adjusted through the Kubernetes API server to ensure that the service can cope with the current load demand, optimize resource usage, improve the service's response speed and processing capacity, and reduce resource waste when the load is low.

[0098] For example, in a Kubernetes environment, you can define a HPA (Horizontal PodAutoscaler) resource, specify scaleTargetRef (resource object) as the Deployment of the large model service, set minReplicas (minimum number of large model service container groups) and maxReplicas (maximum number of large model service container groups), and define metrics to specify the trigger conditions for scaling.

[0099] It should be noted that metrics are the metrics and target values ​​used to trigger scaling. You can set a CPU utilization target. When the actual utilization exceeds this target value, HPA will increase the number of replicas; when the utilization is lower than the target value, HPA will reduce the number of replicas.

[0100] It should be noted that HPA is an automatic scaling controller in Kubernetes, which automatically adjusts the number of Pods (large model service container groups) in Deployment or ReplicaSet according to specified resource indicators or custom indicators.

[0101] This embodiment provides an intelligent adaptive AI large model dynamic activation and scheduling method. This embodiment first monitors the number of concurrent requests in real time and responds quickly to load changes to ensure that the service remains stable during peak hours. The automatic scaler dynamically adjusts the number of services according to the current load to achieve load balancing, avoid overload or idle resources, ensure that the service has sufficient processing capacity during peak hours, and reduce request delays and response times.

[0102] Based on Example 1, Example 5 of the present application proposes a method for dynamic activation and scheduling of an intelligent adaptive AI large model, referring to Figure 5 , step S40 comprises: Step S41, monitoring the load of each of the large model service instances.

[0103] It should be noted that the load situation refers to the amount of tasks or resource usage currently being processed by the service instance, including but not limited to CPU usage, memory usage, network bandwidth, disk I / O, etc. The load situation is an important indicator for measuring the health and performance of the service instance.

[0104] For example, the Metrics Server of Kubernetes is used to collect the load data of the large model service instance.

[0105] It should be noted that Metrics Server is deployed in the Kubernetes cluster by default as a Deployment object. It is responsible for collecting monitoring data from each node in the cluster and storing it in a persistent storage inside the cluster.

[0106] Step S42: Select a target large model service instance through the traffic distribution algorithm according to the load condition and processing capacity of the large model service instance.

[0107] It should be noted that the traffic distribution algorithm is a strategy or method used to determine how to distribute network traffic or requests to multiple processing units. Processing capacity is the number of requests that a single service instance can handle per unit time.

[0108] Optionally, select a traffic distribution algorithm to intelligently select target instances based on the load and processing power of the service instances. The traffic distribution algorithm can be based on strategies such as round-robin, least connections, weighted round-robin, and take into account the load balancing of the instances.

[0109] Optionally, step S42 includes: Step A10, monitor the number of requests being processed by each of the large model service instances, evaluate the load conditions of the large model service instances, sort the large model service instances from low to high according to the load conditions, and determine the load sequence.

[0110] It should be noted that the load sequence is a list of large model service instances sorted from low to high according to load conditions.

[0111] As an optional implementation, the real-time request counts of each major model service instance are collected regularly to evaluate the CPU usage, memory usage and other load indicators of each instance. Based on the preset load threshold, the load of each major model service instance is divided into light load, medium load, heavy load and other levels, and sorted.

[0112] Step A20, analyzing the processing capabilities of each of the large model service instances, and sorting the large model service instances from strong to weak according to the processing capabilities based on the load sequence to determine a comprehensive queue.

[0113] It should be noted that the comprehensive queue is a list of large model service instances that is sorted from strong to weak according to processing capabilities based on the load sequence, which is used to more comprehensively evaluate the applicability of each instance.

[0114] As an optional implementation, the performance of each large model service instance in processing tasks is evaluated through benchmark testing. Based on the benchmark test results, a processing capability score is assigned to each instance. Based on the load sequence, the instances are sorted according to the processing capability score to form a comprehensive queue.

[0115] As another optional implementation, collect the processing request history of each major model service instance, extract features from the collected data, and count the average processing capacity and peak processing capacity of each instance in the past period of time. Use machine learning algorithms to train prediction models, evaluate the potential processing capacity of each instance through the prediction model, and sort them in combination with the load sequence to determine the comprehensive queue.

[0116] Step A30: Based on the comprehensive queue, select the large model service instance with the highest ranking as the target large model service instance.

[0117] Exemplarily, the highest ranked instance in the comprehensive queue is directly selected as the target large model service instance for processing the new request.

[0118] Optionally, other factors may be considered when selecting a target instance, including the geographical location of the instance, maintenance plan, etc.

[0119] Optionally, if the top instances in the comprehensive queue have similar loads and comparable processing capabilities, a polling or random selection method may be used to achieve load balancing.

[0120] Step S43: forward the external request to the target large model service instance for processing through the request queue agent.

[0121] It should be noted that the request queue proxy is a sidecar container in each large model service container group (Pod), which is responsible for processing external requests and forwarding them to the large model service. As the middle layer of traffic forwarding, the request queue proxy not only implements the routing and forwarding of requests, but is also responsible for counting performance indicators such as the number of requests and response time, and reporting these indicators to components such as the autoscaler. The request queue proxy also participates in the health check of the large model service to ensure the health status of the service instance. The request queue proxy receives requests by reserving specific ports and forwards the requests to the listening port of the large model service.

[0122] It should be noted that in container orchestration platforms such as Kubernetes, the sidecar container is an auxiliary container that runs in the same Pod as the main application container. The sidecar container provides additional functions for the main application, including log collection, monitoring, network proxy, etc., without interfering with the operation of the main application.

[0123] Exemplarily, the request queue agent receives external requests through a reserved specific port, and forwards the requests to the listening port of the large model service according to internal routing rules and load balancing strategies.

[0124] Optionally, the request queue agent collects and counts performance indicators such as the processing time and number of requests for each request, and regularly reports these indicators to components such as the autoscaler for resource adjustment and optimization.

[0125] Optionally, a request queue agent is used to periodically send health check requests to the large model service to detect its operating status, and based on the health check results, the routing policy is dynamically adjusted to ensure that the request is not forwarded to a faulty instance.

[0126] Optionally, step S43 includes: Step B10, receiving the external request through the request queue agent, and evaluating the processing capability of each of the large model service instances.

[0127] It should be noted that the evaluation is a quantitative analysis of the current processing capacity of the large model service instance, based on performance indicators and load conditions.

[0128] As an optional implementation, the request queue agent receives external requests through a reserved port and temporarily stores the received requests in an internal queue. The performance indicators and load conditions of each major model service instance are collected in real time, and the current processing capacity score of each instance is calculated based on the collected data.

[0129] Step B20: If the processing capacity of each of the large model service instances has reached the upper limit, or the external request cannot be processed immediately, the external request is placed in the cache area of ​​the request queue agent and sorted according to a preset priority.

[0130] It should be noted that the preset priority is a pre-set request processing order based on business requirements, request type, user importance and other factors.

[0131] As an optional implementation, when the request queue agent detects that the processing capacity of each model service instance has reached the upper limit or cannot process external requests immediately, these requests are placed in the cache area. Fixed priorities are preset for different types of requests, and the requests in the cache area are sorted according to the preset static priority rules.

[0132] Optionally, the priority of requests can be dynamically adjusted based on system status, user behavior, or request context. When a user frequently initiates requests but does not receive a response for a long time, the priority of subsequent requests can be temporarily increased.

[0133] As an optional implementation of sorting, the requests in the cache are sorted according to the arrival time, and processed in the order in which the requests arrive at the cache in a first-in-first-out manner.

[0134] Optionally, on the basis of the first-in-first-out method, a basic weight is assigned to each request based on the arrival time of the request, and the basic weight is adjusted according to dynamic factors such as the urgency of the request and the system load to obtain the final weight. The weight calculation formula is "final weight = basic weight * (1 + urgency coefficient + load adjustment coefficient)". The requests are sorted from high to low according to the final weight, and when selecting the next request to be processed, the request with the highest weight is selected from the sorted request list for processing.

[0135] As another optional implementation of sorting, sorting is performed according to priority and arrival time respectively, and then weights are assigned to the two sortings to combine into a sorting result.

[0136] Step B30, according to the result of the sorting and the load of the large model service instance, the external requests in the cache area are distributed to the large model service in sequence for processing.

[0137] Exemplarily, the large model service instances are sorted according to their load conditions to determine the load sorting to balance the load of the entire system. According to the result of the priority sorting, the requests are distributed to the corresponding large model services in the order of the load sorting.

[0138] This embodiment provides an intelligent adaptive AI large model dynamic activation and scheduling method. This embodiment first continuously monitors the load conditions of each major model service instance to understand the processing capacity and current status of each instance in real time. The traffic distribution algorithm intelligently selects the target service instance to process external requests based on the load conditions and processing capacity of the service instance, thereby improving the overall response speed and processing capacity of the system.

[0139] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the intelligent adaptive AI large model dynamic activation and scheduling method of the present application. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.

[0140] The present application provides an intelligent adaptive AI large model dynamic activation and scheduling device, which includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that at least one processor can execute the intelligent adaptive AI large model dynamic activation and scheduling method in the above-mentioned embodiment one.

[0141] Reference below Figure 6 , which shows a schematic diagram of the structure of an intelligent adaptive AI large model dynamic activation and scheduling device suitable for implementing the embodiment of the present application. The intelligent adaptive AI large model dynamic activation and scheduling device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, personal digital assistants (PDAs), tablet computers (PADs), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The intelligent adaptive AI large model dynamic activation and scheduling device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0142] like Figure 6 As shown, the intelligent adaptive AI large model dynamic activation and scheduling device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the intelligent adaptive AI large model dynamic activation and scheduling device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the intelligent adaptive AI large model dynamic activation and scheduling device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an intelligent adaptive AI large model dynamic activation and scheduling device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have alternatively.

[0143] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0144] The intelligent adaptive AI large model dynamic activation and scheduling device provided by the present application adopts the intelligent adaptive AI large model dynamic activation and scheduling method in the above embodiment, which can solve the technical problem of low resource utilization efficiency of AI large model services. Compared with the prior art, the beneficial effects of the intelligent adaptive AI large model dynamic activation and scheduling device provided by the present application are the same as the beneficial effects of the intelligent adaptive AI large model dynamic activation and scheduling method provided by the above embodiment, and the other technical features in the intelligent adaptive AI large model dynamic activation and scheduling device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0145] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0146] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0147] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the intelligent adaptive AI large model dynamic activation and scheduling method in the above-mentioned embodiment.

[0148] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequencies (RF, Radio Frequency), etc., or any suitable combination of the above.

[0149] The above-mentioned computer-readable storage medium may be included in the intelligent adaptive AI large model dynamic activation and scheduling device; or it may exist independently without being assembled into the intelligent adaptive AI large model dynamic activation and scheduling device.

[0150] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the intelligent adaptive AI large model dynamic activation and scheduling device, the intelligent adaptive AI large model dynamic activation and scheduling device can be written in one or more programming languages ​​or a combination thereof to perform computer program codes for the operation of the present application. The programming languages ​​include object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0151] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0152] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0153] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned intelligent adaptive AI large model dynamic activation and scheduling method, and can solve the technical problem of low resource utilization efficiency of AI large model services. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the intelligent adaptive AI large model dynamic activation and scheduling method provided in the above-mentioned embodiment, and will not be repeated here.

[0154] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. An intelligent adaptive AI large model dynamic activation and scheduling method, characterized in that: Applied to an AI large model service operation platform, the AI ​​large model service operation platform includes a request router, a large model activator, an automatic scaler, a Kubernetes API server, a request queue agent and a large model service, and the intelligent adaptive AI large model dynamic activation and scheduling method includes: Receiving an external request through the request router, and determining a routing direction of the external request according to a path state of the large model activator; If the access state of the large model activator is open, routing the external request to the large model activator, and caching the external request through the large model activator; Based on the number of the external requests cached in the large model activator, trigger the automatic scaler to adjust the number of large model service instances; Through the traffic distribution algorithm, the external requests cached in the large model activator are distributed to each of the large model service instances for processing.

2. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 1, characterized in that: The step of receiving the external request through the request router and determining the routing direction of the external request according to the path state of the large model activator also includes: If the access state of the large model activator is closed, the external request is directly routed to the large model service for processing.

3. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 1, characterized in that: Before the step of receiving the external request through the request router and determining the routing direction of the external request according to the path state of the large model activator, the method includes: If the number of the large model service instances is not zero, marking the access state of the large model activator as closed; If the large model service does not receive a request within a preset time window, the number of instances of the large model service is reduced by the automatic scaler until it is reduced to zero; If the number of large model service instances is zero, the channel state of the large model activator is marked as open.

4. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 1, characterized in that: The step of triggering the automatic scaler to adjust the number of large model service instances based on the number of external requests cached in the large model activator comprises: If the cache quantity of the external request reaches the cache threshold, sending a trigger signal to the automatic scaler through the large model activator to instruct the automatic scaler to increase the number of instances of the large model service; The trigger signal is received by the automatic scaler, and the demand quantity of the large model service is calculated according to the cache quantity; Based on the demand quantity, the autoscaler interacts with the Kubernetes API server to adjust the number of large model service instances to reach the demand quantity.

5. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 1, characterized in that: The step of triggering the automatic scaler to adjust the number of large model service instances based on the number of external requests cached in the large model activator further includes: Monitor the number of requests for inference in the large model service as the number of concurrent requests; Based on the number of concurrent requests and the processing capacity of the large model service instance, the number of the large model service instances required to process the request is evaluated by the automatic scaler as the demand quantity; According to the demand quantity, the number of the large model service instances is adjusted through the Kubernetes API server.

6. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 1, characterized in that: The step of distributing the external request cached in the large model activator to each of the large model service instances for processing by means of a traffic distribution algorithm comprises: Monitoring the load of each of the large model service instances; Selecting a target large model service instance according to the load and processing capacity of the large model service instance by using the traffic distribution algorithm; The external request is forwarded to the target large model service instance for processing through the request queue agent.

7. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 6, characterized in that: The step of selecting a target large model service instance according to the load and processing capacity of the large model service instance through the traffic distribution algorithm includes: Monitor the number of requests being processed by each of the large model service instances, evaluate the load of the large model service instances, sort the large model service instances from low to high according to the load, and determine the load sequence; Analyze the processing capabilities of each of the large model service instances, and on the basis of the load sequence, sort the large model service instances from strong to weak according to the processing capabilities to determine a comprehensive queue; Based on the comprehensive queue, the large model service instance with the highest ranking is selected as the target large model service instance.

8. The intelligent adaptive AI large model dynamic activation and scheduling method according to claim 6, characterized in that: The step of forwarding the external request to the target large model service instance for processing through the request queue proxy also includes: Receiving the external request through the request queue agent and evaluating the processing capability of each of the large model service instances; If the processing capacity of each of the large model service instances has reached the upper limit, or the external request cannot be processed immediately, the external request is placed in the buffer area of ​​the request queue agent and sorted according to a preset priority; According to the result of the sorting and the load condition of the large model service instance, the external requests in the cache area are distributed to the large model service in sequence for processing.

9. An intelligent adaptive AI large model dynamic activation and scheduling device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent adaptive AI large model dynamic activation and scheduling method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the intelligent adaptive AI large model dynamic activation and scheduling method as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Remote procedure calling system and method

    CN117493046A

  • Data processing method and apparatus, electronic device, and storage medium

    WO2024213026A1

Cited By

  • Hierarchical multi-modal tracking method and system based on large model cognitive driving

    CN121811326A