Flow limiting method and device for AI large model calling, equipment and medium
By identifying client and application identifiers in the gateway server and dynamically building a distributed current limiter with ring buffers, the problem of insufficient flexibility in calling traffic control strategies for existing AI large-scale models is solved, adaptive traffic management is realized, and service capabilities in high concurrency scenarios are improved.
Patent Information
- Application Number
- CN202510830879.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-08
AI Technical Summary
The existing AI large-scale AI model calls traffic control strategies with poor flexibility and cannot efficiently and accurately respond to users' dynamic access needs, resulting in traffic control problems in high concurrency scenarios.
The query key name is built through the gateway server identification client and application identification, combined with the ring buffer maintained by the model server, a distributed current limiter is dynamically built, and adaptive traffic restrictions are performed based on actual user access.
It realizes adaptive and dynamic traffic control for users of different calling levels, improves the service capabilities and resource allocation efficiency of AI models in high concurrency scenarios, and meets the convenient and precise traffic restrictions.
Smart Images

Figure CN120455379A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, and medium for limiting traffic for calling a large AI model. Background Art
[0002] With the in-depth application of large AI (Artificial Intelligence) models in fields such as cloud computing and intelligent customer service, the number of API (Application Programming Interface) calls has shown explosive growth, and the problem of traffic control in high-concurrency scenarios has become increasingly severe.
[0003] To cope with the short, high-load interface call pressure of large AI models, traffic control can be implemented based on the accessing user. For example, a single user can only access the large AI model a fixed number of times within a preset time period. If a user frequently sends service call requests to the large AI model within a short period of time, these requests will be directly throttled and no longer sent to the large AI model.
[0004] In the process of implementing the present invention, the inventors discovered that the existing traffic restriction processing mechanism for AI large model calls, which uses users as the flow limiting object, has a relatively fixed flow control strategy and poor flexibility, and cannot efficiently and accurately respond to users' dynamically changing real-time access needs for AI large models. Summary of the Invention
[0005] Embodiments of the present invention provide a method, apparatus, device, and medium for traffic restriction for AI large model calls, which can adaptively and dynamically restrict user traffic based on actual user access conditions in AI large model call scenarios.
[0006] According to one aspect of an embodiment of the present invention, a method for limiting traffic called by a large AI model is provided, which is executed by a gateway server, and the method includes:
[0007] Identify the client ID and application ID in the received target call request for the AI big model, and construct a target query key name based on the client ID and application ID;
[0008] If the target distributed rate limiter matching the target query key name is not cached locally, the target call level of the target call request is obtained;
[0009] Query the data cache status of the ring buffer maintained by the model server in real time, and obtain the remaining traffic value that matches the target call level under the traffic restriction segment to which the target call request belongs;
[0010] The ring buffer is constructed using time units divided based on standard duration. The time units correspond to request queues, which store call requests of various call levels to be executed under the time units. Each time segment of the system time divided by the standard duration is mapped to the ring buffer.
[0011] After building a target distributed rate limiter that matches the target query key name based on the rate limit segment and the remaining rate value, the target call request is sent to the model server.
[0012] According to another aspect of an embodiment of the present invention, there is also provided a flow restriction device for calling an AI large model, which is configured in a gateway server and includes:
[0013] The request identification module is used to identify the client ID and application ID in the received target call request for the AI large model, and construct the target query key name based on the client ID and application ID;
[0014] A rate limiting check module is used to obtain the target call level of the target call request if the target distributed rate limiter matching the target query key name is not cached locally;
[0015] The traffic calculation module is used to query the data cache status of the ring buffer maintained by the model server in real time, and obtain the remaining traffic value that matches the target call level under the traffic restriction section to which the target call request belongs;
[0016] The ring buffer is constructed using time units divided based on standard duration. The time units correspond to request queues, which store call requests of various call levels to be executed under the time units. Each time segment of the system time divided by the standard duration is mapped to the ring buffer.
[0017] The flow limiting execution module is used to build a target distributed flow limiter that matches the target query key name according to the flow limit segment and the remaining flow value, and then send the target call request to the model server.
[0018] According to another aspect of an embodiment of the present invention, an electronic device is provided, the electronic device comprising:
[0019] at least one processor; and
[0020] a memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the traffic limitation method called by an AI large model described in any embodiment of the present invention.
[0022] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a traffic limiting method for calling an AI large model as described in any embodiment of the present invention when executed.
[0023] According to another aspect of an embodiment of the present invention, a computer program product is provided, comprising computer instructions, which implement the steps of the method according to any embodiment of the present invention when executed by a processor.
[0024] According to the technical solution of an embodiment of the present invention, when a gateway server determines that a new distributed rate limiter needs to be constructed for a user based on a received target call request, it does not use a fixed rate limiting strategy to construct the distributed rate limiter. Instead, it combines the target call level to which the target call request belongs and the data cache status of the ring buffer maintained in real time by the model server to obtain the remaining traffic value that matches the target call level within the traffic restriction segment to which the target call request belongs. After constructing a target distributed rate limiter that matches the target query key name based on the traffic restriction segment and the remaining traffic value, the target call request is sent to the model server. This new rate limiting method can achieve the effect of adaptively and dynamically limiting traffic for users of different call levels based on actual user access in AI large model call scenarios, meeting people's growing demand for convenient and precise rate limiting. This solution breaks through the limitations of the extensive control of traditional rate limiting technology. Through multi-level refined management, while ensuring the service quality of high-priority business, it significantly improves the overall service capabilities of AI large models in high-concurrency scenarios, providing reliable technical support for large-scale AI service deployment.
[0025] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 This is a flow chart of a method for limiting traffic for calling a large AI model according to the first embodiment of the present invention;
[0028] Figure 2This is a flowchart of another method for limiting traffic for calling a large AI model according to the second embodiment of the present invention;
[0029] Figure 3 This is a flowchart of another method for limiting traffic for calling a large AI model according to the third embodiment of the present invention;
[0030] Figure 4 This is a structural diagram of a flow limiting device for calling an AI large model provided according to a fourth embodiment of the present invention;
[0031] Figure 5 It is a structural diagram of an electronic device that implements a traffic limitation method called by an AI large model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] To facilitate explanation of the implementation of each embodiment of the present invention, the specific application scenarios to which the embodiments of the present invention are adapted are first described. Each embodiment of the present invention can be applied to a hardware architecture consisting of a terminal, a gateway server, and a model server.
[0035] The terminal is installed with client software. When using the client software, the user can trigger the generation of an AI big model call request. For example, the user uses the AI assistant function provided by the client software to ask questions in natural language and request the AI big model to provide natural language processing services. The above operation will trigger the AI big model call request to be sent to the gateway server.
[0036] The gateway server is mainly used to perform flow control on multiple AI model call requests sent by multiple users to the model server through the client software, so as to avoid service anomalies caused by the model server receiving a large number of AI call requests in a short period of time. For example, a flow limiter is set with the user as the control dimension, and only a certain user is allowed to send 10 AI large model call requests within 1 minute. Accordingly, when the gateway server determines that an AI call request does not need to be flow-limited, it can directly forward this AI call request to the model server. When the gateway server determines that the AI call request needs to be executed, it can directly discard the AI call request and no longer send it to the model server. Furthermore, the gateway server can also directly feedback a request failure response to the sender of the AI model call request that is flow-limited.
[0037] The model server can be understood as a backend server that corresponds one-to-one with the client software. When the AI model is independently developed by the client software's service provider, it can be directly configured on the model server. When the AI model is developed by a third-party model service provider, the model server can include an access interface to the AI model configured on the third-party server. When the model server receives an AI model call request that has been filtered by the gateway server's traffic, it can call the AI model service locally or through a third-party call.
[0038] In each embodiment of the present invention, a new traffic limiting method implemented by a gateway server is proposed. Based on the real-time access of multiple users to the AI large model and the specific user call level, the traffic limiting strategy for different types of users at different access times can be determined to meet the dynamic traffic limiting requirements.
[0039] Example 1
[0040] Figure 1A flowchart of a traffic limiting method for AI large model calls provided in Example 1 of the present invention. This embodiment is applicable to the case where the gateway server performs dynamic traffic limiting on users as the traffic limiting object. The method can be executed by a traffic limiting device for AI large model calls. The device can be implemented in the form of hardware and / or software and can generally be configured in an electronic device, such as a gateway server, and used in conjunction with a client that initiates a call request for the AI large model and a model server for providing AI large model services.
[0041] Correspondingly, such as Figure 1 As shown, the method includes:
[0042] S110. Identify the client identifier and application identifier in the received target call request for the AI big model, and construct a target query key name based on the client identifier and application identifier.
[0043] Among them, the AI big model can be understood as a pre-trained large language model (LLM) based on a deep neural network architecture (e.g., Transformer) with a large number of model parameters (e.g., more than 1 billion). The AI big model can be a generative AI system, that is, after training with massive data, it can realize natural language understanding, generation, and reasoning to obtain new content for multi-source input data such as text, images, or audio and video. Typically, the AI big model can be an autoregressive language model that supports setting the context size, etc.
[0044] It is understandable that when a user calls the AI big model service within the client software, it will trigger the generation of a call request for the AI big model, that is, a target call request. Furthermore, the target call request will include the identity identification information of the sender of the call request. This identity identification information can include two dimensions: one is the client identification at the device or user level, and the other is the application identification at the application level.
[0045] Accordingly, when the gateway server processes AI model call requests, it first extracts two key identifiers from the received target call request: the client identifier and the application identifier. The client identifier uniquely identifies the source of the request and is typically in the form of a user ID, device ID, or API key, such as "user_001" (for user 001) or "device_abcd001" (for device abcd001). The application identifier identifies the specific application or service initiating the request. This identifier typically takes the form of an AppID (Application Identifier) or an application ID. Its importance lies in distinguishing requests from different applications within the same client (e.g., the same user account). For example, a developer account may have multiple AI applications, each with its own unique identifier. These two identifiers together constitute the complete source information for the request. The system then combines these two identifiers according to specific rules to generate the target query key. This key is typically constructed using the format "client identifier: application identifier." This composite key is crucial, serving as the primary basis for system queries and matching rate limiting policies. Through this combination, precise and differentiated traffic control strategies can be implemented for each client + application combination.
[0046] S120: If the target distributed rate limiter matching the target query key name is not cached locally, obtain the target call level of the target call request.
[0047] In this embodiment, the gateway server performs traffic control on a user-by-user basis, distinguishing different users through a combination of client ID and application ID. In other words, a distributed rate limiter is independently constructed for each user, limiting the number of calls to the AI model by that user within the time limiter's validity period.
[0048] That is, the gateway server stores multiple distributed rate limiters, each corresponding to a set query key name. Each distributed rate limiter includes a rate limit time window, a rate limit value, and a flow counter.
[0049] The current limiting time window is used to describe the effective time period of the distributed current limiter, which can also be called the flow restriction section. A current limiting time window corresponds to a variable or fixed current limiting duration (for example, 1 minute or 2 minutes). When the current system time is within the current limiting time window, the distributed current limiter is effective. When the current system time falls outside the current limiting time window, the distributed current limiter is invalid and can be directly deleted from the gateway server.
[0050] The limit value can be understood as the number of calls to the user corresponding to the query key within the limit time window. In other words, it limits the number of calls to the user's AI model within the limit time window. This traffic counter is used to count the number of calls to the AI model by the user within the limit time window in real time.
[0051] Furthermore, the flow limiting principle based on the distributed flow limiter is as follows: whenever the gateway server receives a target call request, it first obtains the user identity identification information (i.e., the target query key name) that matches the target call request. If there is no target distributed flow limiter that matches the target query key name currently stored, it means that the most recent flow limiting operation for the user has expired. At this time, a new distributed flow limiter can be set for the user to perform flow limiting control in a new flow limiting time window. If there is a target distributed flow limiter that matches the target query key name currently stored, the target call request can be directly sent to the model server according to the target distributed flow limiter, or the flow limiting operation can be directly performed.
[0052] As mentioned previously, a distributed rate limiter contains three pieces of information: a rate limit window, a rate limit value, and a flow counter. The rate limit window is directly determined based on the current system time and a preset fixed rate limit duration. The flow counter is a standard counter used to count call counts and can be flexibly configured. Therefore, as long as the rate limit value is accurately set, the distributed rate limiter can be fully configured.
[0053] In each embodiment of the present invention, instead of using a fixed current limit value to set the distributed current limiter, a creative implementation method is proposed to dynamically set different current limit values for different users based on the real-time access situation of the AI large model.
[0054] After deployment in the AI model server cluster, there is often a standard traffic value within a standard time period (for example, 1 minute, or 10 minutes, etc.). The standard traffic can be understood as the standard interface call throughput, that is, the maximum number of interface calls that can be stably processed within the standard time period, for example, 10,000 times per minute. Furthermore, when setting a distributed flow limiter for a certain user, the flow limit value of the user within the flow limit time window can be dynamically determined based on the real-time estimated total traffic value (number of interface calls) available within the flow limit time window and the real-time estimated total number of users who need to make interface calls within the flow limit time window.
[0055] Considering that call requests for large AI models often have different call levels, the call level can be understood as the user level of the call request sender. The user level can be divided according to a preset level division method, for example, divided into paying users and ordinary users according to whether they pay, or divided into corporate users and individual users according to user type, or divided into high-frequency users or ordinary users according to application activity, etc.
[0056] Furthermore, after obtaining the standard traffic value within the standard duration, the standard traffic value can be further divided into two parts according to the call level to determine the traffic value corresponding to each call level within the standard duration. For example, assuming that the call levels are divided into call level 1 and call level 2, and call level 2 has a higher priority than call level 1, the 10,000 calls of the AI large model in 1 minute can be divided into: users of call level 1 can call it 7,000 times, while users of call level 2 can call it 3,000 times.
[0057] Based on this, to accurately determine the rate limit value in the distributed rate limiter created for the target query key, the target call level corresponding to the target call request can be first obtained. Optionally, the gateway server can cache the mapping relationship between different client identifiers and call levels. After obtaining the client identifier in the target call request, the target call level corresponding to the target call request can be identified by querying this mapping relationship.
[0058] S130. Query the data cache status of the ring buffer maintained in real time by the model server to obtain the remaining flow value that matches the target call level under the flow restriction section to which the target call request belongs.
[0059] Among them, the ring buffer is constructed using time units divided based on standard time lengths. The time units correspond to request queues. The request queues store call requests of various call levels that need to be executed under the time units. The system time is mapped to the ring buffer in each time segment divided by the standard time length.
[0060] In this embodiment, the ring buffer is maintained in real time by the model server, describing traffic distribution and remaining traffic in different time periods. When the gateway server needs to create a new distributed rate limiter, it can query the model server for the data cache status of the ring buffer to more accurately set the rate limiter value.
[0061] In the embodiment of the present invention, the ring buffer can be understood as a circular queue architecture, and the time span of the ring buffer corresponds to the standard duration. The standard duration can be understood as the regeneration time of the aforementioned standard traffic value. For example, every 1 minute, the AI large model will regain the ability to respond to 10,000 calls. The standard duration is divided into multiple time units. For example, after dividing the standard duration of 1 minute into 60 1-second time units, a ring buffer with the same time span can be divided into 60 time units accordingly.
[0062] The mapping relationship between time segments and ring buffers can be understood as follows: after dividing the continuous time axis into equal and continuous time segments according to the standard duration (such as 1 minute), each time segment can be mapped to the same ring buffer in the order of time flow. In a specific example, the system time can be divided into: time segment 1 (April 18, 2025, 16:15:01-April 18, 2025, 16:15:59), time segment 2 (April 18, 2025, 16:16:01-April 18, 2025, 16:16:59), etc., along the order of time extension. Furthermore, only one ring buffer cycle can be used to establish a mapping relationship with each time segment. That is, each time unit in the ring buffer can correspond to a specific time segment in each time segment. For example, the first time unit in the ring buffer corresponds to the time segment (April 18, 2025, 16:15:00-April 18, 2025, 16:15:01) in the aforementioned time segment 1.
[0063] Each time unit in the ring buffer corresponds to a request queue, which stores the call requests of each call level required by the AI model within that time unit. With this setup, for example, if one minute is divided into 60 time units, the ring buffer can be used to query the request queue corresponding to each upcoming second.
[0064] In this embodiment, the large model server pre-arranges the call requests that the AI large model needs to process in each time segment of each time period based on the circular buffer and the standard traffic value. The specific arrangement method is that if it is determined that the call request 1 needs to be processed in a future time segment A, the time unit 1 that matches the time segment A is first located in the circular buffer, and the call request 1 is stored in the request queue 1 corresponding to the time segment 1. Then, when the system time elapses to the time segment A, all call requests can be obtained from the request queue 1 and provided to the AI large model for model processing. In this embodiment, the specific encoding method of each call request by the large model server is not limited.
[0065] As previously mentioned, in order to set a new distributed rate limiter for a target call request, it is necessary to dynamically set the rate limit value corresponding to this new distributed rate limiter based on the current access status of each user to the AI large model. Furthermore, it is necessary to first determine the flow restriction section corresponding to the distributed rate limiter, that is, the effective time period of the distributed rate limiter. This flow restriction section can be determined by the time the gateway server requests the target call request and the preset rate restriction duration.
[0066] Assume that the target call request is requested at 15:29:00 on April 25, 2025, and the throttling duration is 1 minute. The flow limit range of the distributed throttler that matches the target call request is: (15:29:00 on April 25, 2025 - 15:29:59 on April 25, 2025).
[0067] After obtaining the traffic restriction segment, the remaining traffic matching the target call level within the traffic restriction segment to which the target call request belongs can be estimated by combining the mapping relationship between the ring buffer and each time segment, the standard traffic value regenerated by the AI large model at standard intervals, and the call request cache status in each request queue matching the ring buffer. In other words, within the traffic restriction segment, the total number of AI large model calls that can be allocated to respond to call requests of the target call level can be estimated.
[0068] S140. After constructing a target distributed flow limiter that matches the target query key name based on the flow limit segment and the remaining flow value, the target call request is sent to the model server.
[0069] Among them, after obtaining the remaining traffic value that matches the target call level under the traffic restriction segment to which the target call request belongs, if the total number of users that match the target call level under the traffic restriction segment is obtained, through simple division calculation, the target traffic value of a single user allocated to the target call level under the traffic restriction segment can be obtained, that is, the target flow limit value.
[0070] The total number of users matching the target call level in the traffic restriction section can be data-fitted or model-predicted based on historical actual user access situations, and this embodiment does not impose any restrictions on this.
[0071] Finally, based on the traffic limit segment and the calculated target limit value, a target distributed limiter matching the target query key can be constructed. After the target distributed limiter is constructed, the gateway server can limit the call requests matching the target query key within the traffic limit segment.
[0072] Based on the above embodiments, if the target distributed rate limiter is not successfully established due to various reasons, the target call request can be directly abandoned and sent to the model server. Instead, the target call request can be fed back to the sender of the target call request, and the sender of the target call request can perform local rate limiting on the client. For example, local rate limiting on the client can be performed using a token bucket. This two-level rate limiting mechanism, namely gateway server and client-side rate limiting, can further improve the effectiveness and accuracy of rate limiting.
[0073] According to the technical solution of an embodiment of the present invention, when a gateway server determines that a new distributed rate limiter needs to be constructed for a user based on a received target call request, it does not use a fixed rate limiting strategy to construct the distributed rate limiter. Instead, it combines the target call level to which the target call request belongs and the data cache status of the ring buffer maintained in real time by the model server to obtain the remaining traffic value that matches the target call level within the traffic restriction segment to which the target call request belongs. After constructing a target distributed rate limiter that matches the target query key name based on the traffic restriction segment and the remaining traffic value, the target call request is sent to the model server. This new rate limiting method can achieve the effect of adaptively and dynamically limiting traffic for users of different call levels based on actual user access in AI large model call scenarios, meeting people's growing demand for convenient and precise rate limiting. This solution breaks through the limitations of the extensive control of traditional rate limiting technology. Through multi-level refined management, while ensuring the service quality of high-priority business, it significantly improves the overall service capabilities of AI large models in high-concurrency scenarios, providing reliable technical support for large-scale AI service deployment.
[0074] Example 2
[0075] Figure 2 This is a flowchart of another method for limiting the flow of large AI model calls, provided in Example 2 of the present invention. This example optimizes the above examples. This example specifically details the operation of "querying the data cache status of the ring buffer maintained in real time by the model server to obtain the remaining flow value that matches the target call level within the flow restriction section to which the target call request belongs."
[0076] Correspondingly, such as Figure 2 As shown, the method may specifically include:
[0077] S210. Identify the client identifier and application identifier in the received target call request for the AI big model, and construct a target query key name based on the client identifier and application identifier.
[0078] S220: If the target distributed rate limiter matching the target query key name is not cached locally, obtain the target call level of the target call request.
[0079] S230: Obtain the flow restriction section to which the target call request belongs according to the request time point of the target call request and the preset flow restriction duration.
[0080] Specifically, when it is determined based on the target query key name that the target distributed flow limiter corresponding to the target query key name is not currently stored, a flow limiting segment that matches the flow limiting duration can be constructed with the request time point of the target call request as the starting point, and flow limiting processing of the target flow limiting value can be performed within the flow limiting segment.
[0081] S240. According to the request time point, obtain the target time segment to which the target call request belongs, and query the consumed traffic value of the AI big model for the target call level in the target time segment from the model server.
[0082] Among them, the consumed traffic value of the target call level can be understood as: the total amount of resource quota (that is, the number of calls to the AI large model) used by requests of a certain call level (such as high, medium and low) in a specific time period. This value will be updated in real time to reflect the actual usage of call requests of this call level in the current time period. For example, in the minute from 10:15:00 to 10:15:59, the high-level call request may have occupied 3000 of the 5000 call quota of its corresponding level, so its consumed traffic value is 3000 times.
[0083] In this embodiment, whenever a target call request is received, the time segment to which it belongs is first determined based on the timestamp of the target call request. For example, a target call request with a high call level arriving at 14:30:23 can be classified into the time segment of 14:30:00-14:30:59. The gateway server then queries the model server for the consumed traffic value for all high call levels within that time segment, that is, to count how many high call level call requests have been processed in that minute. For example, the query determines that 3000 high call level call quotas have been consumed.
[0084] S250. Obtain the target traffic value assigned by the AI large model to the target call level under the standard duration, and calculate the initial remaining traffic value based on the target traffic value and the consumed traffic value.
[0085] In this embodiment, the total traffic quota pre-allocated for a certain type of call level (such as high level) within a standard time length (such as 1 minute) is first obtained. For example, the high call level may be set to allow a maximum of 5,000 calls per minute, and this "5,000 times" is the target traffic value. Next, the traffic value consumed by the call level in the current time period will be queried, for example, the query determines that 3,000 times have been used. Finally, through a simple calculation (target traffic value 5,000 times - consumed traffic value 3,000 times), the initial remaining traffic value is 2,000 times.
[0086] Furthermore, this calculation is performed dynamically and in real time. With each new call request, the latest consumed traffic value is retrieved to ensure the accuracy of the remaining traffic value. For example, if the model server processes 500 high-level requests in the same time period after the above example, the consumed traffic value will be 3500 in the next calculation, and the remaining traffic value will be 1500.
[0087] S260 , identifying each future time unit in the ring buffer according to the request time point, and obtaining a total value of call requests matching the target call level in each request queue matching each future time unit.
[0088] Here, future time units can be understood as the time slots in the ring buffer that will be processed sequentially after the request time. For example, if the request time is 10:15:23 (corresponding to unit 24 in the buffer), then units 25-60 are future time units, each representing a fixed time segment (such as 1 second). In other words, starting from the time unit in the ring buffer where the request time falls, all time units obtained by traversing backward to the end of the ring buffer are future time units.
[0089] In this embodiment, the position of a request in the ring buffer can be determined based on its arrival time, and then the request queues corresponding to all subsequent future time units can be scanned. In these request queues, the number of all pending requests with the same call level (e.g., a higher level) as the current target call request is counted to obtain the total call request value.
[0090] S270 , determining whether the current limiting duration is less than the future duration determined by each future time unit; if so, executing S280 ; otherwise, executing S290 .
[0091] S280. Calculate the remaining traffic value that matches the target call level in the traffic restriction section to which the target call request belongs based on the initial remaining traffic value, the total value of the call requests, and the proportion of the traffic restriction duration in the future duration, and execute S2120.
[0092] If the preset throttling duration is shorter than the total duration covered by the future time unit, a proportional conversion is required: First, the total available remaining traffic value is obtained by subtracting the total number of call requests from the initial remaining traffic value. Then, the total available remaining traffic value corresponding to the total duration covered by the future time unit is converted based on the proportion of the throttling duration to the total duration covered by the future time unit, resulting in a remaining traffic value that matches the target call level within the traffic restriction segment to which the target call request belongs.
[0093] In a specific example, suppose a limit of 5,000 calls per minute is set for the high call level. When a high-level user initiates a request at 10:15:30, the system performs the following calculations: First, the time period (10:15:00-10:15:59) to which 10:15:30 belongs is determined. Then, the number of high-level calls consumed by the model server during this time period (e.g., 3,000) is counted. Then, based on the ring buffer, the number of high-level requests queued in the second half minute from 10:15:30 to 10:15:59 (e.g., 1,000) is counted. This results in a total remaining traffic of 1,000 (5,000-3,000-1,000). If the preset flow limiting time is only 15 seconds (covering only 50% of the second half minute), the traffic that can actually be allocated to all premium users needs to be converted into 500 times (1000 times × 15 / 30) based on the time ratio.
[0094] S290: Calculate a first surplus duration obtained by subtracting the future duration from the current limiting duration, and execute S2100.
[0095] Generally speaking, when the throttling duration exceeds the future time range covered by the ring buffer, an additional time difference is generated. This difference is called the first surplus time. Specifically, the future time range of the ring buffer is the period from the current time point to the end of the buffer cycle (for example, in a 60-second buffer, if the current time is 40 seconds, the future time range is 20 seconds). If the throttling duration is 30 seconds at this time, the 10-second difference between the two (30 seconds throttling duration - 20 seconds future time range) is the first surplus time.
[0096] S2100: Calculate the target number of complete standard durations included in the first surplus duration, and the second surplus duration obtained by subtracting the target number of standard durations from the first surplus duration.
[0097] S2110. Based on the target traffic value, the initial remaining traffic value, the total value of call requests, the target number value, and the proportion of the second surplus time in the standard time, calculate the remaining traffic value that matches the target call level under the traffic restriction section to which the target call request belongs, and execute S2120.
[0098] Generally speaking, when calculating the first surplus duration, the number of complete standard duration cycles (i.e., the target number value) is first determined. For example, if the first surplus duration is 75 seconds and the standard duration is 60 seconds, then it includes 1 complete standard duration (target number value = 1). Next, the second surplus duration is calculated, which is the total time corresponding to these complete cycles subtracted from the first surplus duration (75 seconds - 1 × 60 seconds = 15 seconds). This 15 seconds is the second surplus duration, which represents the remaining time period that is less than a complete standard duration. In flow control, this subdivision method can more accurately handle the quota calculation problem of time periods that exceed the current buffer period.
[0099] S2120. After building a target distributed flow limiter that matches the target query key name based on the flow limit segment and the remaining flow value, the target call request is sent to the model server.
[0100] In this embodiment, it is necessary to integrate the three core parameters of target traffic value (total quota of the call level), initial remaining traffic value (remaining quota of the current period), and total value of call requests (number of queued requests of the same call level) to establish a basic traffic control framework. When it comes to time periods across cycles, the target quantity value (number of complete standard duration cycles) will be introduced for calculation. For example, if there is 1 complete cycle, all quotas of the cycle will be included in the calculation range. At the same time, special consideration will be given to the impact of the second surplus time (the remaining time that is less than the complete cycle), and the corresponding quota ratio will be converted by calculating its proportion in the standard time (such as 15 seconds / 60 seconds = 25%).
[0101] Ultimately, these parameters are combined for a weighted calculation: the full quota for the full cycle is taken into account, the quota for the second most abundant period is proportionally converted, and the number of queued requests is deducted to obtain the precise remaining traffic value. This calculation method ensures that traffic control maintains mathematical accuracy and temporal continuity, regardless of whether it is a full cycle or scattered periods.
[0102] In a specific example, the target traffic value set for high-level users is 5,000 requests per minute. When a high-level user initiates a request at 10:15:30, it is detected that 3,000 high-level traffic requests have been consumed in the time period (10:15:00-10:15:59) at the current time point, and 1,000 high-level requests have been queued for the remaining time (10:15:30-10:15:59). If the throttling duration is 120 seconds, the available traffic for the remaining period (10:15:30-10:15:59) is calculated as (5000-3000-1000) = 1000 times. The quota for the full period (10:16:00-10:16:59) is then calculated as 5000 times per minute. Finally, the second most abundant period (10:17:00-10:17:29) is converted based on the time proportion to obtain a quota of 2500 times (5000 times x 30 / 60). After comprehensive calculations, the total remaining traffic for all premium users during the next 120-second throttling period is 1000 + 5000 + 2500 = 8500 times, which accurately reflects the actual available quota across multiple time periods.
[0103] In this embodiment, a special case can be considered for the case where the target call request belongs to a flow-restricted section that spans a time section. Due to the circular storage of the ring buffer, it is possible that the model server will schedule the future time units corresponding to the next time section while scheduling the call request based on the future time units corresponding to the time section of the current system time. For example, if the model server schedules a call request processing policy for the next 20 seconds, then when the current system time is 10:17:50, a new request queue corresponding to the first 10 seconds of a minute will be encoded.
[0104] When this happens, in order to further improve the accuracy of the calculation, when the traffic restriction segment to which the target call request belongs spans a time segment, when calculating the total value of the call requests, the call requests in all request queues corresponding to all time units in the ring buffer can be obtained as the total value of the call requests.
[0105] The technical solution of the embodiment of the present invention provides a new method for traffic restriction. Based on this method, it is possible to achieve the effect of adaptive and dynamic traffic restriction for users of different call levels based on the actual user access situation in the AI large model call scenario, meeting people's growing demand for convenient and precise traffic restriction. This solution breaks through the limitations of the extensive control of traditional flow restriction technology. Through multi-level refined management, while ensuring the quality of high-priority business services, it significantly improves the overall service capabilities of the AI large model in high-concurrency scenarios, providing reliable technical support for large-scale AI service deployment.
[0106] Furthermore, based on the above embodiments, obtaining the target traffic value allocated by the AI large model to the target call level under the standard duration may also include:
[0107] Obtain the standard traffic value of the AI large model under standard duration;
[0108] Based on the historical number of request users for each call level in multiple historical time periods, calculate the average number of users corresponding to each call level under standard duration;
[0109] The equivalent traffic volume is obtained by weighting the average number of users corresponding to each call level under the standard duration and the level weight corresponding to each call level.
[0110] The target flow value is calculated based on the standard flow value, equivalent flow equivalent and the level weight corresponding to the target call level.
[0111] Generally speaking, the standard traffic value of a large AI model within a standard duration (such as 1 minute) is divided into user quotas at different call levels. Taking a total standard traffic value of 10,000 calls per minute as an example, it can be allocated in different levels according to preset rules: the high-level user group can obtain a call quota of 5,000 times, the intermediate user group is allocated 3,000 times, and the low-level user group is allocated the remaining 2,000 times. The specific allocation method can be dynamically determined based on the historical access history of users at each call level to the large AI model.
[0112] Specifically, the average number of users for each call level can be calculated by analyzing historical data. The specific approach is to collect the historical actual number of request users for each call level (such as high, medium, and low) in multiple standard time periods in the past (such as every minute), and then average these historical values. For example, the total number of high-level user requests in the last 100 1-minute periods can be counted. Assume that there are 48, 52, 50, ..., and so on, and then add up these total requests and divide by 100. The average number of users for high-level users is approximately 50 / minute. The same method is also applicable to the calculation of the average number of users for medium and low-level users. These calculated average numbers of users will serve as the basic reference value for predicting future traffic distribution.
[0113] This averaging method smooths out short-term fluctuations and reflects the stable user base over a longer period. These averages are typically updated regularly, for example, hourly, to ensure forecast accuracy. This allows for a more reasonable estimate of the resource requirements for each call level in the future, enabling more accurate traffic allocation decisions.
[0114] Generally speaking, the equivalent traffic equivalent can be calculated by combining the average number of users for each call level and the preset weights. Specifically, each call level has two key parameters: the average number of users calculated from historical data (e.g., 50 users / minute for high-level users), and the level weight set based on the importance of the service (e.g., weight 3 for high-level users, weight 2 for medium-level users, and weight 1 for low-level users). By multiplying these two parameters and summing them, the total equivalent traffic equivalent can be obtained. For example, assume the average number of high-level users is 50 (weight 3), the average number of medium-level users is 100 (weight 2), and the average number of low-level users is 200 (weight 1). The calculation process is: 50 × 3 = 150 equivalents for high-level users, 100 × 2 = 200 equivalents for medium-level users, and 200 × 1 = 200 equivalents for low-level users. The total equivalent traffic equivalent = 150 + 200 + 200 = 550. This comprehensive indicator reflects the weighted demand for system resources by different call levels, providing a quantitative basis for subsequent traffic allocation. Through this weighted calculation, the resource usage of users at different call levels can be more reasonably balanced.
[0115] Furthermore, the target flow value for each call level can be calculated based on three core parameters: standard flow value (total system capacity), equivalent flow equivalent (weighted user demand) and call level weight (priority coefficient). In the specific calculation, QPM is the total number of calls per minute (i.e., standard flow value), and the target flow value corresponding to the target call level is Among them C i is the average number of users of the i-th calling level, W i is the weight value corresponding to the i-th call level, W is the weight corresponding to the target call level, W∈[W i , i=1~n], n is the total number of pre-divided call levels.
[0116] The core idea of the calculation formula is to allocate the total system capacity (QPM) according to the weighted user demand ratio of each call level. When calculating specifically, first sum the product of the average number of users of all call levels and the weight (i.e., the equivalent traffic equivalent). ), and then divide the product of the weight (W) of the target call level to be calculated and QPM by the total equivalent, and finally obtain the target flow value that should be allocated to the target call level.
[0117] Example 3
[0118] Figure 3 This is a flowchart of another method for limiting traffic flow for a large AI model, provided in Example 3 of the present invention. This example optimizes the above examples. This example specifically refines the operation of "constructing a target distributed rate limiter that matches the target query key name based on the rate limit segment and the remaining rate value."
[0119] Correspondingly, such as Figure 3 As shown, the method may specifically include:
[0120] S310. Identify the client identifier and application identifier in the received target call request for the AI big model, and construct a target query key name based on the client identifier and application identifier.
[0121] S320: If the target distributed rate limiter matching the target query key name is not cached locally, obtain the target call level of the target call request.
[0122] S330, querying the data cache status of the ring buffer maintained in real time by the model server, and obtaining the remaining flow value matching the target call level under the flow restriction section to which the target call request belongs;
[0123] Among them, the ring buffer is constructed using time units divided based on standard time lengths. The time units correspond to request queues. The request queues store call requests of various call levels that need to be executed under the time units. The system time is mapped to the ring buffer in each time segment divided by the standard time length.
[0124] S340: Construct a flow-limiting time window according to the flow-limiting section.
[0125] Specifically, the flow restriction section can be directly determined as the flow restriction time window.
[0126] S350. Calculate the target flow limit value according to the remaining flow value, the flow limit duration corresponding to the flow limit section, and the average number of users corresponding to the target call level under the standard duration.
[0127] The remaining traffic value is the total number of available calls corresponding to the target call level within the traffic restriction section. After obtaining the average number of users corresponding to the target call level under the standard duration, the converted number of users under the traffic restriction section can be calculated based on the data ratio between the traffic restriction duration and the standard duration.
[0128] For example, if the throttling duration is 2 minutes and the average number of users corresponding to the target call level in 1 minute is 10,000, then the converted number of users in the throttling section is 20,000. For another example, if the throttling duration is 30 seconds and the average number of users corresponding to the target call level in 1 minute is 10,000, then the converted number of users in the throttling section is 5,000.
[0129] Finally, the integer value of the quotient obtained by dividing the remaining traffic value by the converted number of users in the traffic limit section can be used as the target traffic limit value.
[0130] S360: Construct a target distributed current limiter that matches the target query key name based on the current limit time window, the target current limit value, and a traffic counter with an initial count value of 1.
[0131] Since there is currently no target distributed rate limiter, the target call request should be sent directly to the model server without being rate limited. Therefore, a traffic counter with an initial count value of 1 can be constructed to count the number of times this target call request is sent.
[0132] In an embodiment of the present invention, the system achieves precise control of traffic by constructing a target distributed rate limiter. The rate limiter is composed of three core elements working together: first, setting a rate limiting time window (such as 14:30:00-14:30:59) to clearly define the boundaries of the control period, secondly configuring a target rate limiting value (such as 10 times) as a trigger threshold, and finally initializing a traffic counter to count the number of requests in real time. The system ensures that the rate limiter acts accurately on a specific object by combining the client identifier and the application identifier to generate a unique key name (such as "user_001:app_ai"). When the counter reaches the preset threshold, new requests are automatically intercepted until the rate limiting time window expires.
[0133] S370: Send the target call request to the model server.
[0134] The embodiment of the present invention provides a method for fine-grained traffic control of AI large models based on a ring buffer. First, a unique query key is constructed through the client and application identifiers to achieve accurate identification of requests. When a request arrives, the traffic restriction segment to which it belongs is determined based on the time point and the flow restriction duration, and the ring buffer is queried to obtain real-time traffic data. Through a dynamic calculation mechanism, the remaining traffic value, flow restriction duration and user scale are comprehensively considered, and a segmented algorithm is used to accurately calculate the target flow restriction value: the remaining traffic in the current period is directly calculated, and when crossing cycles, it is dynamically adjusted based on the complete cycle standard value and the surplus period ratio value. Based on these parameters, a distributed flow limiter containing precise time windows, traffic thresholds and traffic counters is constructed to ensure accurate control in a cluster environment. The entire solution achieves efficient and accurate traffic monitoring through the circular queue design of the ring buffer, which can not only ensure the quality of high-priority services, but also effectively improve the overall utilization of system resources, providing a reliable traffic control solution for AI large model services.
[0135] Optionally, based on the above embodiments, after constructing the target query key name according to the client identifier and the application identifier, the following steps may be further included:
[0136] If the local cache has a target distributed rate limiter that matches the target query key name, then the count value of the traffic counter in the target distributed rate limiter is obtained and incremented by one to obtain an updated count value;
[0137] Check whether the update count value is greater than the target current limit value in the target distributed current limiter;
[0138] If not, the target call request is sent to the model server; if so, the target call request is directly throttled.
[0139] Generally speaking, when the system detects that there is a distributed rate limiter matching the target query key (such as "user_001:app_ai") in the local cache, the flow control mechanism of the rate limiter will be activated immediately. In a specific example, suppose that the rate limiter of an AI application user "user_001:app_ai" is set to process a maximum of 10 requests per minute. When the 8th request arrives, the corresponding distributed rate limiter will be found, and the flow counter (for example, the current value is 7) will be read from the distributed rate limiter and safely updated to 8 through atomic operations. The updated flow counter value will be immediately compared with the target flow limit value in real time: if the new value (8) is less than or equal to the target flow limit value (10), the request will be immediately released; but if the counter has reached or exceeded the target flow limit value (for example, the 11th request causes the flow counter to change from 10 to 11), the flow control mechanism will be triggered immediately, for example, the target call request will be directly discarded, and the prompt message "Server busy, please try again later" will be returned to the sender of the target call request.
[0140] Optionally, based on the above embodiments, this solution further includes:
[0141] Real-time time detection of the current limiting time window of each distributed current limiter in the local cache;
[0142] Based on the detection results, delete the failed distributed current limiters.
[0143] Generally speaking, the gateway server acts like a patrol sentry, continuously checking the timeliness of each distributed rate limiter in its local cache. By comparing the current system time to see if it exceeds the rate limiter window in each distributed rate limiter in real time, it can accurately identify expired distributed rate limiters. For example, when the system time reaches 14:01:00, the distributed rate limiter with a rate limiter window of (14:00:00-14:01:00) will be automatically marked as invalid.
[0144] A cleanup mechanism can be immediately activated for these failed distributed rate limiters. By scanning all distributed rate limiters stored on the gateway server, objects whose rate limit time windows have expired can be precisely removed, ensuring that only currently valid control objects remain. This dynamic maintenance mechanism not only optimizes resource utilization but also prepares for the next round of traffic control.
[0145] Example 4
[0146] Figure 4 This is a structural diagram of a flow restriction device called by an AI large model provided in the fourth embodiment of the present invention. Figure 4 As shown, the device includes:
[0147] The request identification module 410 is used to identify the client identifier and application identifier in the received target call request for the AI large model, and construct a target query key name based on the client identifier and application identifier;
[0148] a current limiting checking module 420 for obtaining a target call level of a target call request if a target distributed current limiter matching the target query key name is not cached locally;
[0149] The traffic calculation module 430 is used to query the data cache status of the ring buffer maintained in real time by the model server, and obtain the remaining traffic value matching the target call level under the traffic restriction section to which the target call request belongs;
[0150] The ring buffer is constructed using time units divided based on standard duration. The time units correspond to request queues, which store call requests of various call levels to be executed under the time units. Each time segment of the system time divided by the standard duration is mapped to the ring buffer.
[0151] The flow limiting execution module 440 is used to build a target distributed flow limiter that matches the target query key name according to the flow limiting segment and the remaining flow value, and then send the target call request to the model server.
[0152] According to the technical solution of an embodiment of the present invention, when a gateway server determines that a new distributed rate limiter needs to be constructed for a user based on a received target call request, it does not use a fixed rate limiting strategy to construct the distributed rate limiter. Instead, it combines the target call level to which the target call request belongs and the data cache status of the ring buffer maintained in real time by the model server to obtain the remaining traffic value that matches the target call level within the traffic restriction segment to which the target call request belongs. After constructing a target distributed rate limiter that matches the target query key name based on the traffic restriction segment and the remaining traffic value, the target call request is sent to the model server. This new rate limiting method can achieve the effect of adaptively and dynamically limiting traffic for users of different call levels based on actual user access in AI large model call scenarios, meeting people's growing demand for convenient and precise rate limiting. This solution breaks through the limitations of the extensive control of traditional rate limiting technology. Through multi-level refined management, while ensuring the service quality of high-priority business, it significantly improves the overall service capabilities of AI large models in high-concurrency scenarios, providing reliable technical support for large-scale AI service deployment.
[0153] Based on the above embodiments, the flow calculation module 430 may specifically include:
[0154] A time segment positioning unit is used to obtain the flow restriction segment to which the target call request belongs based on the request time point of the target call request and the preset flow restriction duration;
[0155] The real-time traffic query unit is used to obtain the target time segment to which the target call request belongs based on the request time point, and query the consumed traffic value of the AI large model for the target call level in the target time segment from the model server;
[0156] The initial remaining traffic calculation unit is used to obtain the target traffic value allocated by the AI large model to the target call level under the standard duration, and calculate the initial remaining traffic value based on the target traffic value and the consumed traffic value;
[0157] a dynamic traffic calculation unit, configured to identify each future time unit in the ring buffer according to the request time point, and obtain a sum of call requests matching the target call level in each request queue matching each future time unit;
[0158] The remaining traffic value calculation unit is used to calculate the remaining traffic value that matches the target call level under the traffic restriction segment to which the target call request belongs based on the initial remaining traffic value, the total value of the call requests, and the proportion of the flow restriction duration in the future duration if the flow restriction duration is less than the future duration determined by each future time unit.
[0159] Based on the above embodiments, the flow calculation module 430 may further include:
[0160] a cross-cycle length calculation unit, configured to calculate a first surplus time obtained by subtracting the future time from the current limiting time if the current limiting time is greater than or equal to the future time determined by each future time unit;
[0161] The time period decomposition calculation unit is used to calculate the target number of complete standard durations included in the first surplus duration, and the second surplus duration obtained by subtracting the target number of standard durations from the first surplus duration;
[0162] The multi-cycle traffic comprehensive calculation unit is used to calculate the remaining traffic value that matches the target call level in the traffic restriction section to which the target call request belongs based on the target traffic value, the initial remaining traffic value, the total value of call requests, the target number value, and the proportion of the second surplus time in the standard time.
[0163] Based on the above embodiments, the initial residual flow calculation unit is specifically used to:
[0164] Obtain the standard traffic value of the AI large model under standard duration;
[0165] Based on the historical number of request users for each call level in multiple historical time periods, calculate the average number of users corresponding to each call level under standard duration;
[0166] The equivalent traffic volume is obtained by weighting the average number of users corresponding to each call level under the standard duration and the level weight corresponding to each call level.
[0167] The target flow value is calculated based on the standard flow value, equivalent flow equivalent and the level weight corresponding to the target call level.
[0168] Based on the above embodiments, the current limiting execution module 440 may further include:
[0169] A time window generating unit is used to construct a flow limiting time window according to a flow limiting section;
[0170] A dynamic flow limit value calculation unit, configured to calculate a target flow limit value based on the remaining flow value, the flow limit duration corresponding to the flow limit section, and the average number of users corresponding to the target call level under the standard duration;
[0171] The distributed flow limiter assembly unit is used to construct a target distributed flow limiter that matches the target query key name according to the flow limit time window, the target flow limit value, and a flow counter with an initial count value of 1.
[0172] Furthermore, based on the above embodiments, the following may be further included:
[0173] A traffic counter dynamic update module is used to obtain the count value of the traffic counter in the target distributed flow limiter and add one to obtain an updated count value if there is a target distributed flow limiter matching the target query key name in the local cache;
[0174] An intelligent current limiting decision module is used to detect whether the update count value is greater than the target current limiting value in the target distributed current limiter;
[0175] The current limiting decision module is used to send the target call request to the model server if no; if yes, directly limit the target call request.
[0176] Based on the above embodiments, the current limiting execution module 440 may further include:
[0177] The current limiting timeliness detection unit is used to perform timeliness detection on the current limiting time windows of each distributed current limiter cached locally in real time;
[0178] The resource recovery unit is used to delete failed distributed current limiters according to the detection results.
[0179] The flow limiting device for AI large model calls provided in an embodiment of the present invention can execute a flow limiting method for AI large model calls provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0180] In the technical solutions of the embodiments of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0181] Example 5
[0182] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0183] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0184] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0185] The processor 11 can be various general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, for example, executing a traffic limiting method called by an AI large model as described in any of the embodiments of the present invention, that is:
[0186] Identify the client ID and application ID in the received target call request for the AI big model, and construct a target query key name based on the client ID and application ID;
[0187] If the target distributed rate limiter matching the target query key name is not cached locally, the target call level of the target call request is obtained;
[0188] Query the data cache status of the ring buffer maintained by the model server in real time, and obtain the remaining traffic value that matches the target call level under the traffic restriction segment to which the target call request belongs;
[0189] The ring buffer is constructed using time units divided based on standard duration. The time units correspond to request queues, which store call requests of various call levels to be executed under the time units. Each time segment of the system time divided by the standard duration is mapped to the ring buffer.
[0190] After building a target distributed rate limiter that matches the target query key name based on the rate limit segment and the remaining rate value, the target call request is sent to the model server.
[0191] In some embodiments, a traffic restriction for a large AI model call may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the traffic restriction for a large AI model call described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured in any other appropriate manner (for example, by means of firmware) to execute a traffic restriction method for a large AI model call as described in any one of the embodiments of the present invention.
[0192] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0194] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0196] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0197] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0198] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0199] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for limiting the flow of AI large model calls, characterized in that: Executed by a gateway server, the method includes: Identify the client ID and application ID in the received target call request for the AI big model, and construct a target query key name based on the client ID and application ID; If the target distributed rate limiter matching the target query key name is not cached locally, the target call level of the target call request is obtained; Query the data cache status of the ring buffer maintained in real time by the model server to obtain the remaining traffic value that matches the target call level under the traffic restriction section to which the target call request belongs; The ring buffer is constructed using time units divided based on standard duration. The time units correspond to request queues, which store call requests of various call levels to be executed under the time units. Each time segment of the system time divided by the standard duration is mapped to the ring buffer. After building a target distributed rate limiter that matches the target query key name based on the rate limit segment and the remaining rate value, the target call request is sent to the model server.
2. The method according to claim 1, characterized in that Query the data cache status of the ring buffer maintained in real time by the model server, and obtain the remaining traffic value that matches the target call level under the traffic restriction segment to which the target call request belongs, including: According to the request time point of the target call request and the preset flow limit duration, obtain the flow limit section to which the target call request belongs; Based on the request time point, obtain the target time segment to which the target call request belongs, and query the consumed traffic value of the AI large model for the target call level in the target time segment from the model server; Obtain the target traffic value assigned by the AI big model to the target call level under the standard duration, and calculate the initial remaining traffic value based on the target traffic value and the consumed traffic value; Identifying each future time unit in the ring buffer according to the request time point, and obtaining a sum of call requests matching the target call level in each request queue matching each future time unit; If the flow limit duration is less than the future duration determined by each future time unit, the remaining flow value that matches the target call level under the flow limit segment to which the target call request belongs is calculated based on the initial remaining flow value, the total value of call requests, and the proportion of the flow limit duration in the future duration.
3. The method according to claim 2, characterized in that After obtaining the sum of call requests matching the target call level, it also includes: If the flow-limiting duration is greater than or equal to the future duration determined by each future time unit, the first surplus duration is obtained by subtracting the future duration from the flow-limiting duration. Calculate the target number of complete standard durations included in the first surplus duration, and the second surplus duration obtained by subtracting the target number of standard durations from the first surplus duration; Based on the target traffic value, the initial remaining traffic value, the total value of call requests, the target number value, and the proportion of the second surplus time in the standard time, the remaining traffic value that matches the target call level in the traffic restriction segment to which the target call request belongs is calculated.
4. The method according to claim 2 or 3, characterized in that Obtain the target traffic value assigned by the AI large model to the target call level under standard duration, including: Obtain the standard traffic value of the AI large model under standard duration; Based on the historical number of request users for each call level in multiple historical time periods, calculate the average number of users corresponding to each call level under standard duration; The equivalent traffic volume is obtained by weighting the average number of users corresponding to each call level under the standard duration and the level weight corresponding to each call level. The target flow value is calculated based on the standard flow value, equivalent flow equivalent and the level weight corresponding to the target call level.
5. The method according to claim 1, wherein Based on the traffic limit segment and the remaining traffic value, a target distributed rate limiter is constructed that matches the target query key name, including: According to the flow restriction section, a flow restriction time window is constructed; Calculate the target traffic limit value based on the remaining traffic value, the traffic limit duration corresponding to the traffic limit section, and the average number of users corresponding to the target call level under the standard duration; According to the flow limiting time window, the target flow limiting value, and the traffic counter with an initial count value of 1, a target distributed flow limiter that matches the target query key name is constructed.
6. The method according to claim 5, characterized in that After constructing the target query key name based on the client ID and application ID, it also includes: If the local cache has a target distributed rate limiter that matches the target query key name, then the count value of the traffic counter in the target distributed rate limiter is obtained and incremented by one to obtain an updated count value; Check whether the update count value is greater than the target current limit value in the target distributed current limiter; If not, the target call request is sent to the model server; if so, the target call request is directly throttled.
7. The method according to claim 5, characterized in that The method further comprises: Real-time time detection of the current limiting time window of each distributed current limiter in the local cache; Based on the detection results, delete the failed distributed current limiters.
8. A flow limiting device for calling a large AI model, characterized in that: Configured in the gateway server, the device includes: The request identification module is used to identify the client ID and application ID in the received target call request for the AI large model, and construct the target query key name based on the client ID and application ID; A rate limiting check module is used to obtain the target call level of the target call request if the target distributed rate limiter matching the target query key name is not cached locally; The traffic calculation module is used to query the data cache status of the ring buffer maintained by the model server in real time, and obtain the remaining traffic value that matches the target call level under the traffic restriction section to which the target call request belongs; The ring buffer is constructed using time units divided based on standard duration. The time units correspond to request queues, which store call requests of various call levels to be executed under the time units. Each time segment of the system time divided by the standard duration is mapped to the ring buffer. The flow limiting execution module is used to build a target distributed flow limiter that matches the target query key name according to the flow limit segment and the remaining flow value, and then send the target call request to the model server.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a traffic limitation method for AI large model calls according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a traffic limiting method for AI large model calling according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Flow control method, device and equipment for large model service and storage medium
CN121037308A