Model token uniform measurement method and system
Patent Information
- Application Number
- CN202611080956.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-29
AI Technical Summary
若直接基于各模型对应的模型原生Token数量进行计量,则容易导致不同模型之间的Token计量标准不一致,从而降低不同模型之间Token计量结果的可比性
[0016]本发明的模型Token统一计量方法及系统的有益效果是:本发明通过获取目标模型对应的模型原生Token数量,并结合对应的Token转换参数确定统一计量Token数量,从而能够对不同模型对应的Token消耗进行统一量纲转换,能够解决由于不同模型采用不同Tokenization方式而导致的Token计量标准不统一的问题,提高了不同模型之间Token计量的一致性与可比性。
Smart Images

Figure CN122840255A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence model reasoning technology, and more specifically, to a unified measurement method and system for model tokens. Background Technology
[0002] With the development of artificial intelligence model inference technology, applications such as text generation, code generation, question-answering interaction, and content creation based on large language models are gradually increasing. Different artificial intelligence models typically employ different tokenization methods, meaning that different models may differ in their token segmentation rules, encoding granularity, and encoding results for the same input content.
[0003] In this scenario, the same input content may have different numbers of native tokens depending on the model. For example, the same text content may be segmented into different numbers of tokens depending on the model. If the tokens are directly measured based on the number of native tokens for each model, it can easily lead to inconsistent token measurement standards between different models, thereby reducing the comparability of token measurement results between different models. Summary of the Invention
[0004] The problem addressed by this invention is: how to uniformly measure the number of native tokens for different models, so as to improve the consistency and comparability of token measurement between different models.
[0005] To address the aforementioned issues, this invention provides a unified measurement method and system for model tokens.
[0006] In a first aspect, the present invention provides a unified measurement method for model tokens, including: In response to a model inference request, determine the target model corresponding to the model inference request; The inference process corresponding to the model inference request is executed through the target model, and the number of native model tokens corresponding to the inference process is determined. Obtain the Token conversion parameters corresponding to the target model; Based on the original number of tokens in the model and the token conversion parameters, the corresponding unified measurement token number is determined.
[0007] Optionally, after determining the corresponding unified measurement token quantity, the model token unified measurement method further includes: Based on the unified metering token quantity and the preset unit token price, the billing result corresponding to the model inference request is determined. Based on the billing results, deduct fees and output the inference results corresponding to the model inference request.
[0008] Optionally, the target model is selected from a preset model library, which includes at least one model using a tokenization method.
[0009] Optionally, determining the target model corresponding to the model inference request in response to the model inference request includes: In response to a model inference request, authentication processing is performed on the user who initiated the model inference request; Upon successful authentication, rate limiting detection is performed on the model inference request. Upon successful rate limiting detection, the user's billing policy information and remaining token amount information are obtained. In response to the remaining token amount information satisfying the preset amount conditions, the target model is determined based on the model inference request and the billing strategy information.
[0010] Optionally, the preset limit conditions include: The remaining token amount is greater than or equal to the estimated token consumption amount corresponding to the model inference request.
[0011] Optionally, the authentication process for the user initiating the model inference request in response to the model inference request includes: In response to a model inference request, user identifier, scene identifier, and user input text are extracted based on the model inference request; Based on the user identifier, the authentication process is performed on the user who initiates the model inference request.
[0012] Optionally, determining the target model based on the model inference request and the billing policy information includes: Based on the scene identifier, determine the target model type corresponding to the model inference request; Based on the billing policy information, determine the model invocation permissions corresponding to the user; Based on the model access permissions and the load information of each candidate model corresponding to the target model type, the target model is determined from each of the candidate models corresponding to the target model type.
[0013] Optionally, obtaining the Token conversion parameters corresponding to the target model includes: Obtain the preset token conversion parameters corresponding to the target model from the preset parameter library, and use the preset token conversion parameters as the token conversion parameters; Alternatively, the preset Token conversion parameters can be dynamically adjusted based on at least one of the model resource consumption corresponding to the target model, the application scenario information corresponding to the model inference request, and the billing strategy information to determine the Token conversion parameters. The Token conversion parameters include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing power cost factor.
[0014] Optionally, the model token unified measurement method further includes: Obtain the resource consumption data corresponding to the reasoning process; The resource consumption data is associated with the number of unified metering tokens to generate a token cost attribution record; wherein the token cost attribution record is used to dynamically adjust the token conversion parameters.
[0015] Secondly, this invention provides a unified measurement system for model tokens, comprising: The application layer is used to generate model inference requests in response to user input; The Token billing engine layer communicates with the application layer; the Token billing engine layer is used to respond to the model inference request, determine the target model corresponding to the model inference request, and obtain the Token conversion parameters corresponding to the target model; A large model layer is communicatively connected to the Token billing engine layer; the large model layer includes at least one target model instance, and the large model layer is used to execute the inference process corresponding to the model inference request through the target model, and determine the number of native model tokens corresponding to the inference process; A computing resource layer, which is communicatively connected to the large model layer; the computing resource layer includes heterogeneous computing resources for deploying the target model instance; The Token billing engine layer is also used to determine the corresponding unified metering token quantity based on the original token quantity of the model and the token conversion parameters.
[0016] The beneficial effects of the model token unified measurement method and system of the present invention are as follows: The present invention obtains the number of native tokens corresponding to the target model and determines the unified measurement token number in combination with the corresponding token conversion parameters, thereby enabling unified dimensional conversion of token consumption for different models. This solves the problem of inconsistent token measurement standards caused by different models using different tokenization methods, and improves the consistency and comparability of token measurement between different models. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a unified measurement method for model tokens in an embodiment of the present invention. Figure 2 This is a flowchart illustrating a unified measurement method for model tokens in an embodiment of the present invention. Figure 3 This is a structural block diagram of a unified measurement system for model tokens in an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.
[0020] Combination Figure 1 As shown, this embodiment of the invention provides a unified measurement method for model tokens, including: Step 100: In response to the model inference request, determine the target model corresponding to the model inference request.
[0021] The method in this embodiment can be applied to artificial intelligence model inference scenarios (such as text generation, code generation, question-and-answer interaction, content creation, etc. based on large language models). It is used to uniformly measure the native tokens generated by different models (such as heterogeneous models) during the inference process, so as to realize the unified dimensional conversion of token consumption between different models and improve the consistency and comparability of token measurement between different models.
[0022] Specifically, in step 100, when a model inference request is received, such as a model inference request received from user input, the target model corresponding to the model inference request is determined. The model inference request may include user input content and request information representing the corresponding inference requirement; the target model is an artificial intelligence model used to perform the corresponding inference task, such as a text generation model, code generation model, question-and-answer interaction model, or content creation model. Different target models may adopt different tokenization methods; that is, different target models may have different token segmentation rules, encoding granularity, and encoding results for the same input content.
[0023] Step 200: Execute the inference process corresponding to the model inference request through the target model, and determine the number of native tokens of the model corresponding to the inference process.
[0024] Specifically, after determining the target model, in step 200, the inference process corresponding to the model inference request is executed through the target model to generate the corresponding inference result. Simultaneously, during the inference process, the number of native model tokens consumed by the target model during the inference process can be counted. The number of native model tokens represents the number of tokens generated by the target model after encoding the input and / or output content based on its own tokenization method.
[0025] Step 300: Obtain the Token conversion parameters corresponding to the target model.
[0026] Specifically, since different models typically employ different tokenization methods, the number of native tokens for the same input content may differ across models. For example, different models may use different segmentation granularities, encoding rules, and token segmentation methods for the same text content, resulting in significant differences in the number of native tokens. In this case, directly measuring based on the number of native tokens can easily lead to inconsistent measurement standards between different models. Therefore, after determining the number of native tokens, in step 300, the token conversion parameters corresponding to the target model can be further obtained to convert the number of native tokens into a unified measurement token number. These token conversion parameters characterize the differences in token measurement between different models.
[0027] Step 400: Determine the corresponding unified measurement token quantity based on the model's native token quantity and token conversion parameters.
[0028] Specifically, in step 400, a unified measurement token quantity is determined based on the model's native token quantity and token conversion parameters. This unified measurement token quantity can be used to perform a standardized dimensional conversion of token consumption across different models, enabling the conversion of native tokens generated by different models under different tokenization methods into token quantities under a unified measurement standard. Thus, even if different models use different tokenization methods for the same content, unified measurement token quantities can achieve consistent measurement of token consumption across different models.
[0029] Thus, the method in this embodiment obtains the number of native tokens corresponding to the target model and determines the number of tokens to be uniformly measured by combining the corresponding token conversion parameters. This enables a uniform dimensional conversion of token consumption for different models, which solves the problem of inconsistent token measurement standards caused by different tokenization methods used by different models, and improves the consistency and comparability of token measurement between different models.
[0030] Optionally, combined Figure 2 As shown, after determining the corresponding unified measurement token quantity, the model token unified measurement method also includes: Step 500: Based on the unified metering token quantity and the preset unit token price, determine the billing result corresponding to the model inference request; Step 600: Perform the deduction process based on the billing results and output the inference results corresponding to the model inference request.
[0031] Specifically, after determining the quantity of uniform measurement tokens, in step 500, the method of this embodiment can further determine the billing result corresponding to the model inference request based on the quantity of uniform measurement tokens and the preset unit token price. For example, the billing result corresponding to the model inference request = quantity of uniform measurement tokens * preset unit token price. The preset unit token price can be used to characterize the price parameter corresponding to a unit of uniform measurement token.
[0032] After determining the billing result, in step 600, the deduction process is performed based on the billing result, and the inference result corresponding to the model inference request is output. For example, the corresponding amount can be deducted from the user's account balance, token limit, or prepaid fees based on the billing result, and the inference result generated by the target model is returned to the corresponding user terminal.
[0033] Thus, the method in this embodiment not only achieves unified token measurement across different models, but also determines the corresponding billing result based on the unified measurement token quantity, thereby improving the consistency of billing standards across different models.
[0034] For example, as the demand for large language model inference grows, operators can build a unified computing resource pool based on data center GPU resources, edge node NPU resources, etc., and provide artificial intelligence model inference services to users based on a unified service platform. Since model inference services typically exhibit peak and trough characteristics, some computing resources may be idle during low-load periods, and resource utilization needs improvement. Therefore, the aforementioned unified service platform can be used to uniformly manage model inference requests, token measurement, billing processing, and resource scheduling to improve the efficiency of computing resource utilization. The method in this embodiment can be applied to scenarios of artificial intelligence model inference services provided by operators. For example, operators can deploy multiple artificial intelligence models based on computing resources such as data center GPU resources and edge node NPU resources, and provide model inference services to users through a unified service platform. When a user initiates a model inference request through the operator's corresponding upper-layer application (such as an intelligent customer service application, resume generation application, code assistant application, etc.), the corresponding inference process can first be executed through the target model, and the corresponding number of native tokens for the model can be determined.
[0035] Since different target models may employ different tokenization methods, the number of native tokens for the same input content may differ. Therefore, this paper further obtains the token transformation parameters corresponding to the target models and determines a unified token quantity based on the model's native token quantity and the token transformation parameters. For example, for different target models using different tokenization methods, the corresponding token transformation parameters can be used to convert the number of native tokens for different models into a unified token quantity under a unified measurement standard.
[0036] After determining the quantity of unified metering tokens, the corresponding billing result can be determined based on the quantity of unified metering tokens and the preset unit token price, and the deduction process can be performed based on the billing result. For example, the corresponding fee can be deducted from the user's account balance, token limit account, or prepaid account, and the corresponding reasoning result can be returned to the user.
[0037] Thus, the method in this embodiment can not only achieve unified measurement of token consumption among different models, but also achieve billing processing under a unified measurement standard among different models in the context of operator artificial intelligence services, thereby improving the consistency and comparability of billing standards among different models.
[0038] Optionally, the target model is selected from a preset model library, which includes at least one model using the tokenization method.
[0039] Specifically, the target model determined in the method of this embodiment can be selected from a preset model library. The preset model library includes at least one model employing a tokenization method; that is, the preset model library may include only one model for performing inference tasks, or it may include multiple models for performing different inference tasks or adapting to different inference needs. When the preset model library includes multiple models, at least some of the models may employ different tokenization methods. For example, different models may have different tokenization rules, encoding granularity, or encoding results for the same input content. In this case, the multiple models can constitute a heterogeneous model system.
[0040] In this embodiment, the target model can be determined from a preset model library based on the inference requirements corresponding to the model inference request. Since the target model has a corresponding tokenization method, after determining the number of native tokens for the model, the number of native tokens can be further converted into a unified measurement token number by combining the token conversion parameters corresponding to the target model.
[0041] Optionally, in response to a model inference request, determining the target model corresponding to the model inference request includes: In response to a model inference request, authenticate the user who initiated the request.
[0042] Specifically, upon receiving a model inference request, the user initiating the request can be authenticated first. This can be done by verifying the user's identity, API access permissions, or account status based on the user identifier corresponding to the model inference request, to determine whether the current user has the necessary permissions to call the corresponding model inference service.
[0043] For example, the authentication process includes: obtaining the corresponding user account information based on the user identifier corresponding to the model inference request; verifying at least one of the following based on the user account information: user identity information, interface access permission information, and account status information; and determining that the current user has the permission to call the corresponding model inference service if the verification is successful. Specifically, firstly, the corresponding user account information is obtained based on the user identifier corresponding to the model inference request. The user identifier may include a user account, mobile phone number, user ID, application interface identifier, or other information used to uniquely identify the user; the user account information may include user identity information, interface access permission information, account status information, and other information related to calling the model inference service. Subsequently, at least one of the following can be verified based on the user account information: user identity information; interface access permission information; and account status information. The user identity information can be used to characterize the legitimacy of the current user's identity; the interface access permission information can be used to characterize whether the current user has the permission to call the corresponding model inference service; and the account status information can be used to characterize whether the current user's account is in a normal state. If the verification is successful, it can be determined that the current user has the permission to call the corresponding model inference service, and the subsequent model inference process can be allowed to continue. Conversely, if verification fails, the current model inference request can be rejected to prevent unauthorized users from invoking the model inference service. Thus, by authenticating users who initiate model inference requests, the security and reliability of the model inference service invocation process can be improved.
[0044] In response to successful authentication, rate limiting is performed on model inference requests.
[0045] After authentication is successful, rate limiting can be further applied to model inference requests. This can be done by considering factors such as the current user's historical request frequency, the number of requests per unit time, or the number of token calls to determine if the current model inference request meets preset rate limiting conditions, thus preventing excessive system resource consumption due to a large number of requests in a short period.
[0046] For example, rate limiting detection includes: obtaining at least one of the following information for the current user within a preset time range: request frequency information, request count information, token call count information, and resource usage information; determining whether the current model inference request meets the preset rate limiting conditions based on at least one of the following information; and allowing the current model inference request to continue executing subsequent processing procedures if the preset rate limiting conditions are met. Specifically, firstly, at least one of the following information for the current user within a preset time range is obtained: request frequency information, request count information, token call count information, and resource usage information. The request frequency information can be used to characterize the frequency at which the current user initiates model inference requests per unit time; the request count information can be used to characterize the number of model inference requests corresponding to the current user within the preset time range; the token call count information can be used to characterize the token call status corresponding to the current user; and the resource usage information can be used to characterize the system resource usage status corresponding to the model inference request of the current user. Subsequently, it can be determined whether the current model inference request meets the preset rate limiting conditions based on at least one of the following information: request frequency information, request count information, token call count information, and resource usage information. For example, preset rate limiting conditions include: the number of model inference requests made by the current user within a preset time range is less than or equal to a preset request count threshold; and / or, the cumulative number of tokens invoked by the current user within a preset time range is less than or equal to a preset token invoke threshold; and / or, the system resources occupied by the current user's corresponding model inference request are less than or equal to a preset resource usage threshold. If the preset rate limiting conditions are met, the current model inference request can continue to execute subsequent processing steps. Conversely, if the preset rate limiting conditions are not met, the current model inference request can be restricted from continuing to execute, thereby reducing the pressure on system resources caused by a large number of model inference requests in a short period of time. In this way, by rate limiting detection of model inference requests, it is possible to avoid excessive system resource usage caused by a concentrated influx of model inference requests in a short period of time, which is beneficial to improving the stability of the model inference service operation.
[0047] Upon successful rate limiting detection, obtain the user's billing policy information and remaining token amount. Specifically, after the rate limiting detection passes, the user's billing policy information and remaining token amount information can be further obtained. The billing policy information can be used to represent the service policy corresponding to the current user, such as different users corresponding to different token amounts, different calling permissions, or different model service ranges, etc.; the remaining token amount information can be used to represent the current user's available token amount.
[0048] In response to the remaining token amount information meeting the preset amount conditions, the target model is determined based on the model inference request and billing policy information.
[0049] Specifically, it determines whether the remaining token amount meets preset limits, such as whether the current remaining token amount is greater than or equal to the estimated token consumption amount corresponding to the current model inference request. When the remaining token amount meets the preset limits, the corresponding target model can be further determined based on the model inference request and billing policy information. Specifically, the target model suitable for executing the corresponding inference task can be determined according to the inference requirements corresponding to the model inference request and the billing policy information corresponding to the current user. For example, users corresponding to different billing policies can call different types of models, or models with different performance levels.
[0050] Thus, the method in this embodiment determines the target model through authentication processing, rate limiting detection, token quota verification, and information based on billing policies. This ensures the stability of the model inference service while enabling reasonable scheduling of model inference services for different users. Furthermore, by performing authentication processing, rate limiting detection, and token quota verification on model inference requests, this method avoids unauthorized calls, a large number of requests in a short period of time, and abnormal inference requests due to insufficient quota, thereby improving the security and stability of the model inference service operation.
[0051] Optionally, the preset credit limit conditions include: The remaining token amount is greater than or equal to the estimated token consumption amount corresponding to the model inference request.
[0052] Specifically, after obtaining the remaining token amount for the current user, the estimated token consumption for the current model inference request can be further determined. The estimated token consumption amount can be used to characterize the number of tokens expected to be consumed by the current model inference request.
[0053] For example, the estimated token consumption amount for the current model inference request can be determined based on at least one of the following: the length of the input content corresponding to the model inference request, the target model type, historical token consumption, and the inference task type. Generally, the longer the input content or the finer the token encoding granularity of the target model, the larger the estimated token consumption amount may be. Subsequently, the remaining token amount for the current user can be compared with the estimated token consumption amount. When the remaining token amount is greater than or equal to the estimated token consumption amount, it can be determined that the current remaining token amount meets the preset limit condition, and the subsequent model inference process can be allowed to continue. Conversely, when the remaining token amount is less than the estimated token consumption amount, the current model inference request can be restricted from continuing.
[0054] Thus, by verifying the remaining token amount and the estimated token consumption amount before model inference, the method in this embodiment can avoid continuing the model inference process when the token amount is insufficient, which helps to improve the rationality of the model inference service management process.
[0055] Optionally, in response to a model inference request, the authentication process for the user initiating the model inference request includes: In response to the model inference request, extract the user identifier, scene identifier, and user input text based on the model inference request; Based on the user identifier, the user who initiates the model inference request is authenticated.
[0056] Specifically, during authentication, the model inference request can be parsed first to extract the user identifier, scenario identifier, and user input text carried in the request. The user identifier uniquely identifies the user initiating the model inference request, such as a user ID, phone number, account identifier, or API call identifier. The scenario identifier represents the application scenario corresponding to the current model inference request, such as a text generation scenario, code generation scenario, question-and-answer interaction scenario, or content creation scenario. The user input text serves as the input content for the target model to perform the inference process.
[0057] After extracting the user identifier, the corresponding user account information can be queried based on the user identifier, and authentication processing can be performed on the user initiating the model inference request based on the user account information. For example, the validity of the user's identity, the validity of the interface call permissions, and the normality of the account status can be verified based on the user account information. If the authentication is successful, subsequent processing steps such as rate limiting detection, quota verification, and target model determination can continue; if the authentication fails, the processing steps corresponding to the current model inference request can be terminated.
[0058] Thus, the method in this embodiment extracts the user identifier, scene identifier, and user input text from the model inference request, and performs authentication processing on the user based on the user identifier. On the one hand, it can confirm the legitimate source of the model inference request, and on the other hand, it can provide basic data for subsequent determination of the target model based on the scene identifier and execution of model inference based on the user input text.
[0059] Optionally, based on model inference requests and billing policy information, the target model is determined, including: Based on the scene identifier, determine the target model type corresponding to the model inference request; Based on billing policy information, determine the model access permissions corresponding to the user; Based on model access permissions and the load information of each candidate model corresponding to the target model type, the target model is determined from each candidate model corresponding to the target model type.
[0060] Specifically, when determining the target model, the target model type corresponding to the model inference request can be determined first based on the scene identifier extracted from the model inference request. The scene identifier can be used to characterize the application scenario corresponding to the current model inference request. For example, when the scene identifier corresponds to a text generation scenario, the corresponding target model type can be determined to be a text generation model; when the scene identifier corresponds to a code generation scenario, the corresponding target model type can be determined to be a code generation model.
[0061] After determining the target model type, model invocation permissions corresponding to the user can be further determined based on billing policy information. This billing policy information can be used to characterize the service policy corresponding to the current user. For example, users corresponding to different billing policies may have different model invocation scopes, different model invocation priorities, or different model service levels. For instance, users corresponding to different billing policies may have different model invocation priorities, where users corresponding to higher-level billing policies may have priority in invoking high-performance model instances.
[0062] Subsequently, based on model access permissions and the load information of each candidate model corresponding to the target model type, the target model can be determined from among the candidate models corresponding to the target model type. The load information of the candidate models can be used to characterize the current resource consumption of the corresponding candidate model, such as GPU utilization, CPU utilization, video memory usage, power consumption, or the number of current requests. In some embodiments, among multiple candidate models that satisfy the current user's model access permissions, the candidate model with the lowest current load can be preferentially selected as the target model to execute the corresponding model inference request.
[0063] Thus, the method in this embodiment determines the target model by combining the scenario identifier, billing policy information, and load information corresponding to the candidate model. This enables the reasonable scheduling of model inference requests while meeting the model call requirements of different users, thereby improving the resource scheduling efficiency of the model inference service.
[0064] Optionally, obtaining the Token conversion parameters corresponding to the target model includes: Retrieve the preset token conversion parameters corresponding to the target model from the preset parameter library, and use the preset token conversion parameters as token conversion parameters; Alternatively, based on at least one of the following: model resource consumption corresponding to the target model, application scenario information corresponding to the model inference request, and billing strategy information, the preset Token conversion parameters can be dynamically adjusted to determine the Token conversion parameters. The Token conversion parameters include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing power cost factor.
[0065] Specifically, the token conversion parameters include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing power cost factor. The model type factor can be used to characterize differences in model capabilities, parameter scale, or inference resource consumption between different models; the scenario factor can be used to characterize differences in value weights corresponding to different application scenarios, such as different scenario factors for resume creation, code generation, and entertainment / casual conversation scenarios; the package type factor can be used to characterize differences in metering corresponding to different billing strategies, such as different package type factors for monthly packages, annual packages, pay-as-you-go, and tiered billing. In some embodiments, different package types correspond to different unit token prices or different token conversion rules, such as the unit token price for an annual package being lower than the unit token price for a monthly package, in order to meet the differentiated model inference service needs of different users and improve the flexibility of model inference service billing methods; the real-time computing power cost factor can be used to characterize the real-time resource consumption of the current target model, such as dynamic indicators like GPU utilization, memory usage, and power consumption. For example, different target models may have different model type factors due to differences in model structure, parameter size, tokenization method, and inference resource consumption. For instance, different models such as DeepSeek V4, Qwen, and Llama can correspond to different model type factors; different application scenarios such as resume creation, code generation, and entertainment interaction can correspond to different scenario factors; different billing strategies such as monthly plans, annual plans, pay-as-you-go, and tiered billing can correspond to different plan type factors; and different resource consumption such as GPU utilization, memory usage, and power consumption can correspond to different real-time computing cost factors. Therefore, the token conversion parameters in this embodiment can adopt at least one of the model type factor, scenario factor, plan type factor, and real-time computing cost factor. This allows for dynamic adjustment and unified measurement of token consumption based on at least one of different models, application scenarios, billing strategies, and resource consumption, improving the flexibility and rationality of token measurement results.
[0066] In determining the token transformation parameters, one can obtain the preset token transformation parameters corresponding to the target model from a preset parameter library and use these preset token transformation parameters as the token transformation parameters. The preset parameter library can pre-store preset token transformation parameters corresponding to different models to characterize the differences in token measurement between different models.
[0067] Alternatively, in determining the token conversion parameters, in addition to obtaining the preset token conversion parameters corresponding to the target model from the preset parameter library, the preset token conversion parameters can be dynamically adjusted based on at least one of the following: model resource consumption of the target model, application scenario information corresponding to the model inference request, and billing strategy information, to determine the corresponding token conversion parameters. Here, model resource consumption can be used to characterize the resource usage of the target model during the inference process, such as GPU utilization, CPU utilization, video memory usage, power consumption, or inference duration; application scenario information can be used to characterize the application scenario corresponding to the current model inference request, such as resume creation, code generation, question-and-answer interaction, or entertainment interaction; billing strategy information can be used to characterize the service strategy corresponding to the current user, such as monthly packages, annual packages, pay-as-you-go, or tiered billing. For example, when the resource consumption of the target model is high, the corresponding real-time computing cost factor can be increased; conversely, when the resource consumption of the target model is low, the corresponding real-time computing cost factor can be decreased. This can reduce the corresponding token metering cost when the computing resource load is low, guiding some model inference tasks to execute during low-load periods, which is beneficial to improving the utilization rate of computing resources. For example, under different application scenarios or different billing strategies, the corresponding scenario factors or package type factors can also change dynamically. In the above cases, the Token conversion parameters in the method of this embodiment can not only reflect the differences in Token measurement between different models, but also be dynamically adjusted in combination with at least one of the model resource consumption, application scenario information, and billing strategy information.
[0068] Thus, the method in this embodiment determines the token conversion parameters by combining preset token conversion parameters and a dynamic adjustment mechanism, thereby improving the accuracy, flexibility, and dynamic adaptability of the unified token measurement results across different models.
[0069] For example, if the token conversion parameters include model type factor, scenario factor, package type factor, and real-time computing power cost factor, then the uniformly measured token quantity can be determined based on the model's native token quantity, model type factor, scenario factor, and real-time computing power cost factor, for example: The unified metering token quantity = the model's native token quantity * model factor * scenario factor * computing power factor * package factor; the billing result = the unified metering token quantity * the preset unit token price.
[0070] Optionally, the unified measurement method for model tokens also includes: Obtain resource consumption data corresponding to the reasoning process; Resource consumption data is associated with the number of unified metering tokens to generate token cost attribution records; these records are used to dynamically adjust token conversion parameters.
[0071] Specifically, during the inference process of the target model, resource consumption data for the corresponding inference process can be obtained. This resource consumption data can be used to characterize the system resource consumption of the current inference process. For example, resource consumption data may include GPU utilization, CPU utilization, video memory usage, power consumption, inference duration, or other data used to characterize resource consumption.
[0072] Subsequently, resource consumption data can be associated with the corresponding number of unified metering tokens to generate corresponding token cost attribution records. These token cost attribution records can be used to characterize the resource consumption corresponding to a given number of unified metering tokens. For example, the number of unified metering tokens corresponding to a model inference request can be associated with resource consumption data such as GPU utilization, memory usage, and power consumption for that model inference request to generate corresponding token cost attribution records.
[0073] Token cost attribution records can be used for subsequent token conversion parameter optimization and cost accounting. For example, during user request processing, the monitoring module of the computing resource pool can asynchronously collect resource consumption data such as GPU utilization, CPU utilization, memory usage, and power consumption of each computing node in the background, and associate the resource consumption data of the corresponding inference request with the unified metering token quantity to generate the corresponding token cost attribution record. Subsequently, based on the token cost attribution record, the token conversion parameters of the corresponding target model can be updated so that the updated token conversion parameters can more accurately reflect the actual resource consumption of the corresponding target model. At the same time, the token cost attribution record can also be used to statistically analyze the resource consumption costs corresponding to different model inference services, providing data support for subsequent cost accounting, resource scheduling optimization, and billing strategy adjustments for model inference services. The process of collecting resource consumption data and generating token cost attribution records can be executed asynchronously in the background to avoid blocking the output process of the inference results corresponding to user requests.
[0074] Thus, the method in this embodiment obtains resource consumption data corresponding to the inference process and generates corresponding Token cost attribution records, thereby enabling the association between the Token consumption of different models and the actual resource consumption. This allows for subsequent optimization and adjustment of Token conversion parameters based on the actual resource consumption, and provides data support for cost accounting corresponding to model inference services. This is beneficial for improving the accuracy and rationality of unified Token measurement results among different models.
[0075] Combination Figure 3 As shown, another embodiment of the present invention provides a unified measurement system for model tokens, including: The application layer is used to generate model inference requests in response to user input; The Token billing engine layer communicates with the application layer. In response to model inference requests, the Token billing engine layer determines the target model corresponding to the model inference request and obtains the Token conversion parameters corresponding to the target model. The large model layer communicates with the Token billing engine layer. The large model layer includes at least one target model instance. The large model layer is used to execute the inference process corresponding to the model inference request through the target model and determine the number of native tokens of the model corresponding to the inference process. The computing resource layer communicates with the large model layer; the computing resource layer includes heterogeneous computing resources for deploying target model instances. The Token billing engine layer is also used to determine the corresponding unified metering token quantity based on the model's native token quantity and token conversion parameters.
[0076] The unified model token measurement system in this embodiment can be applied to artificial intelligence model inference service scenarios to uniformly measure the native model tokens generated by different models during inference. For example, it can be applied to artificial intelligence application scenarios such as text generation, code generation, question-answering interaction, and content creation.
[0077] The unified metering system for model tokens comprises a computing resource layer, a large model layer, a token billing engine layer, and an application layer. These layers can interact via standardized interfaces to enable collaborative transmission of model inference requests, token metering information, and resource consumption data.
[0078] The application layer is used to generate model inference requests in response to user input. For example, the application layer includes at least one upper-layer application, such as an intelligent customer service application, a resume generation application, a code assistant application, or other artificial intelligence applications. Users can input model inference requests through the corresponding application, such as inputting text content, code requirements, image generation requirements, or question-and-answer requests. The application layer is used to generate inference requests in response to user input, send the inference requests to the Token Billing Engine layer, and receive the inference results returned by the Token Billing Engine layer.
[0079] The Token billing engine layer communicates with the application layer. In response to a model inference request, the Token billing engine layer determines the target model corresponding to the request and obtains the Token conversion parameters for that target model. For example, the Token billing engine layer may include an API gateway module. The API gateway module can receive model inference requests from the application layer and perform processing such as request parsing, user authentication, rate limiting detection, and request routing. For instance, the API gateway module can parse the model inference request based on the user identifier, scenario identifier, and user input text, and route the request to the corresponding target model based on the parsing result. In some embodiments, the Token billing engine layer can obtain preset Token conversion parameters corresponding to the target model from a preset parameter library and use these preset parameters as the Token conversion parameters. These Token conversion parameters characterize the differences in Token measurement between different models. For example, the Token conversion parameters may include at least one of a model type factor, a scenario factor, a package type factor, and a real-time computing power cost factor.
[0080] The large model layer communicates with the token billing engine layer. The large model layer includes at least one target model instance. It is used to execute the inference process corresponding to the model inference request through the target model and determine the number of native model tokens corresponding to the inference process. For example, different target model instances in the large model layer can use different tokenization methods. For instance, different model instances such as DeepSeek, Qwen, and Llama can be deployed in the large model layer. Because different models can use different tokenization methods, the number of native model tokens corresponding to the same input content may differ across different models.
[0081] The computing resource layer communicates with the large model layer. The computing resource layer includes heterogeneous computing resources for deploying target model instances. For example, heterogeneous computing resources may include data center GPU resources, edge node NPU resources, or other computing resources used to perform AI model inference tasks. During model inference, the target model instance in the large model layer can execute the inference process based on the corresponding model inference request and generate the corresponding inference result. Simultaneously, the number of native model tokens consumed in the corresponding inference process can be counted and returned to the token billing engine layer. Subsequently, the token billing engine layer can determine the corresponding unified metering token quantity based on the number of native model tokens and the corresponding token conversion parameters. For example, for different target models using different tokenization methods, the number of native model tokens corresponding to different models can be converted into a unified metering token quantity under a unified metering standard through the corresponding token conversion parameters.
[0082] In some embodiments, the Token billing engine layer can also dynamically adjust the corresponding Token conversion parameters based on the model resource consumption of the target model, the application scenario information corresponding to the model inference request, and the billing strategy information. For example, when the resource utilization rate of the corresponding computing node is high, the corresponding real-time computing power cost factor can be increased; conversely, when the resource utilization rate of the corresponding computing node is low, the corresponding real-time computing power cost factor can be decreased, so as to improve the adaptability of the unified Token metering result to the actual resource consumption.
[0083] Thus, this embodiment, by setting up an application layer, a token billing engine layer, a large model layer, and a computing resource layer, and combining token conversion parameters to uniformly measure the number of native tokens for different models, can solve the problem of inconsistent token measurement standards caused by different models using different tokenization methods, thereby improving the consistency and comparability of token measurement results between different models. On the other hand, through the collaborative cooperation between the application layer, the token billing engine layer, the large model layer, and the computing resource layer, the unified model token measurement system can achieve coordinated processing between model inference, unified token measurement, and resource scheduling, thereby improving the scalability of the overall system architecture.
[0084] For example, the computing resource layer includes heterogeneous computing resources and a monitoring unit. Heterogeneous computing resources may include data center GPU resources, edge node NPU resources, or other computing resources used to perform artificial intelligence model inference tasks. The monitoring unit is used to collect resource consumption metrics for each computing node in the heterogeneous computing resources. For example, resource consumption metrics may include GPU utilization, CPU utilization, memory usage, power consumption, inference time, or other data characterizing resource consumption. For example, heterogeneous computing resources can form a unified computing resource pool for different target model instances to share and utilize.
[0085] The large model layer comprises at least one heterogeneous model instance deployed on the computing resource layer. Different heterogeneous model instances can employ different tokenization methods; that is, different models may have different token segmentation rules, encoding granularity, and encoding results for the same input content. For example, the large model layer can deploy different model instances such as DeepSeek, Qwen, and Llama. The large model layer is used to receive inference requests from the token billing engine layer, invoke the corresponding heterogeneous model instance to execute the inference process, generate inference results, and determine the number of native model tokens consumed in the inference process.
[0086] The Token billing engine layer communicates with the large model layer, monitoring unit, and application layer. It receives inference requests from the application layer and routes them to the corresponding large model layer. Subsequently, the Token billing engine layer obtains the number of native model tokens returned by the large model layer and resource consumption metrics collected by the monitoring unit. Based on the number of native model tokens, resource consumption metrics, and token conversion parameters, it determines the unified metering token quantity. The token conversion parameters characterize the differences in token metering between different models. For example, the token conversion parameters may include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing power cost factor. The model type factor can be used to characterize differences in model capabilities, parameter scale, or inference resource consumption between different models; the scenario factor can be used to characterize differences in value weights corresponding to different application scenarios; the package type factor can be used to characterize metering differences corresponding to different billing strategies; and the real-time computing power cost factor can be used to characterize the real-time resource consumption of the current model instance. After determining the number of unified metering tokens, the token billing engine layer can also determine the corresponding billing result based on the number of unified metering tokens, and forward the inference result returned by the large model layer to the application layer. For example, the inference result can be encapsulated as network response data and returned to the corresponding application layer.
[0087] The application layer includes at least one upper-layer application. Upper-layer applications may include intelligent customer service applications, resume generation applications, code assistant applications, or other artificial intelligence applications. The application layer is used to generate inference requests in response to user input, send the inference requests to the Token Billing Engine layer, and receive the inference results returned by the Token Billing Engine layer.
[0088] Thus, this embodiment, by setting up a computing resource layer, a large model layer, a token billing engine layer, and an application layer, and combining token conversion parameters to uniformly measure the number of native tokens for different models, can solve the problem of inconsistent token measurement standards caused by different models using different tokenization methods, thereby improving the consistency and comparability of token measurement results between different models. Furthermore, through the collaborative cooperation between the application layer, the token billing engine layer, the large model layer, and the computing resource layer, the unified model token measurement system can achieve collaborative processing between model inference, unified token measurement, resource monitoring, and billing processing, thereby improving the scalability of the overall system architecture.
[0089] Optionally, the Token billing engine layer includes an API gateway module. The API gateway module can receive model inference requests from the application layer and perform processing such as request parsing, user authentication, rate limiting detection, request routing, and result return. For example, the API gateway module can parse model inference requests based on user identifiers, scene identifiers, and user input text, and route the model inference requests to the corresponding target model based on the parsing results. In this way, by providing a unified model call interface for different upper-layer applications through the API gateway, the complexity of different applications accessing the model inference service can be reduced, which is beneficial to improving the universality and scalability of the model inference service.
[0090] For example, combined Figure 3As shown, the unified model token metering system in this embodiment may include an application layer, a token billing engine layer, a large model layer, and a computing resource layer. These layers can interact with each other through standardized interfaces to achieve coordinated transmission of model inference requests, model inference results, token metering data, and resource consumption data. The application layer, located at the upper layer of the system, provides users with different types of artificial intelligence application services. For example, the application layer may include image generation applications, video generation applications, music generation applications, customer service assistant applications, code assistant applications, personal assistant applications, and resume creation applications. Users can input model inference requests through the corresponding applications, such as inputting text content, code requirements, image generation requirements, or question-and-answer requests. The token billing engine layer, located between the application layer and the large model layer, is used to implement functions such as model inference request management, unified token metering, and billing processing. For example, the token billing engine layer may include an API gateway and a billing platform. The API gateway can be used to receive model inference requests from the application layer and perform routing, authentication, and rate limiting processing. For example, an API gateway can parse model inference requests based on user identifiers, scenario identifiers, and user input, and route the requests to the corresponding target models according to user access permissions, target model load, and application scenarios. A billing platform can be used for unified metering and billing of token consumption corresponding to model inference requests. For example, the billing platform can support different billing strategies such as monthly plans, annual plans, pay-as-you-go, and tiered billing. Different billing strategies can correspond to different token amounts, different model access permissions, or different billing rules. A large model layer can be used to execute model inference tasks. For example, a large model layer can include multiple heterogeneous models, such as DeepSeek V4, Qwen, and Llama models. Different models can use different tokenization methods; that is, different models may have different token segmentation rules, encoding granularity, and encoding results for the same input content. Therefore, the number of native tokens corresponding to the same input content may differ across models. In this case, the token billing engine layer can obtain the number of native tokens generated during the inference process of the corresponding target model and determine the unified metering token quantity based on the corresponding token conversion parameters. The token conversion parameters may include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing power cost factor, to achieve unified measurement of token consumption across different models. The computing power resource layer may be located at the bottom layer of the system and is used to provide the computing power resources required for model inference to the large model layer. For example, the computing power resource layer may include a computing power resource pool and a real-time computing power monitoring module.The computing resource pool can include data center GPUs, edge node NPUs, computing centers, or other heterogeneous computing resources to deploy different model instances and execute model inference tasks. The real-time computing power monitoring module can be used to collect resource consumption indicators for each computing node in real time, such as GPU utilization, CPU utilization, memory usage, and power consumption.
[0091] In some embodiments, the unified measurement system for model tokens also includes a data layer, and the large model layer can also communicate with the data layer. The data layer can be used to provide inference-related data to the corresponding model. For example, the data layer may include multimodal token corpora, customer service data, user data, resume templates, press releases, or other model training and inference-related data.
[0092] In some embodiments, the Token billing engine layer can also obtain resource consumption indicators collected by the real-time computing power monitoring module and dynamically adjust the Token conversion parameters based on the corresponding resource consumption. For example, when the resource utilization rate of the corresponding computing node is high, the corresponding real-time computing power cost factor can be increased; conversely, when the resource utilization rate of the corresponding computing node is low, the corresponding real-time computing power cost factor can be decreased, so as to achieve dynamic measurement and dynamic adjustment for different model inference services.
[0093] Thus, this embodiment sets up an application layer, a token billing engine layer, a large model layer, and a computing resource layer, and combines the token conversion parameters corresponding to different models to achieve a unified measurement of the number of native tokens in the model. This solves the problem of inconsistent token measurement standards caused by different models using different tokenization methods, and improves the consistency and comparability of token measurement results between different models.
[0094] For example, the following describes the unified measurement method of model tokens in this application by taking a user initiating a model inference request through a resume creation application as an example, combined with the unified measurement system of model tokens.
[0095] Users can initiate model inference requests through the corresponding application layer. For example, university students can upload their resume text through a resume creation application and request an AI model to generate resume optimization suggestions. The resume creation application can act as a higher-level application instance within the application layer and initiate a corresponding API call request to the Token billing engine layer. This model inference request can carry a user identifier, a scenario identifier, and user input text. The user identifier can include a mobile phone number, user ID, or account identifier; the scenario identifier can be used to represent the application scenario corresponding to the current model inference request, such as a resume creation scenario; and the user input text can be used as input content for the target model to perform the inference process.
[0096] When the API gateway receives a model inference request from the application layer, it can parse the request to extract the user identifier, scene identifier, and user input text. Then, based on the user identifier, it can obtain the corresponding user account information and verify at least one of the following based on the user account information: user identity information, interface access permission information, and account status information, to determine whether the current user has the authority to call the corresponding model inference service. For example, it can verify whether the current user is a legitimate registered user and whether the current API key is valid. When authentication fails, it can return a corresponding authentication error response and terminate the processing flow corresponding to the current model inference request.
[0097] After authentication, the system can further check whether the current model inference request meets preset rate limiting conditions based on factors such as the current user's historical request frequency, the number of requests per unit time, or the number of token calls. For example, it can determine whether the current model inference request exceeds a preset request threshold based on the number of model inference requests made by the current user within a preset time range. In some embodiments, different billing strategies correspond to different rate limiting rules; for example, monthly plan users can make 60 calls per minute, while trial users can make 10 calls per minute. When the preset rate limiting conditions are not met, a rate limiting error response can be returned, and the processing flow corresponding to the current model inference request can be terminated.
[0098] After authentication and rate limiting checks are passed, the API gateway can also call the unified billing platform to obtain the billing policy information and remaining token amount information for the current user. The billing policy information may include different billing policies such as monthly plans, annual plans, pay-as-you-go, or tiered billing. Based on the billing policy information for the current user, the unified billing platform can determine the corresponding remaining token amount and the estimated token consumption amount for the current model inference request. When the remaining token amount is less than the estimated token consumption amount, an insufficient balance error response can be returned, prompting the user to recharge or adjust the corresponding billing policy.
[0099] After the quota verification is passed, the target model type corresponding to the current model inference request can be determined based on the scenario identifier. Then, combined with the current user's model invocation permissions and the load information of the corresponding candidate models, the target model is determined from the candidate models. For example, in a resume creation scenario, the target model type that excels in text generation can be prioritized. Simultaneously, among multiple candidate models that satisfy the current user's model invocation permissions, the candidate model with lower current GPU utilization or fewer current requests can be prioritized as the target model. In this embodiment, the current model inference request can be routed to the Qwen model instance.
[0100] Subsequently, after receiving the corresponding user input text, the target model instance in the large model layer can execute the corresponding model inference process to generate the corresponding inference result. For example, it can generate an optimized resume text. During the inference process, the target model can also count the number of native model tokens consumed by the corresponding inference process. Since different models can use different tokenization methods, the number of native model tokens for the same input content may differ between different models.
[0101] After obtaining the number of native tokens for the model, preset token conversion parameters corresponding to the target model can be retrieved from the preset parameter library and used as token conversion parameters. These token conversion parameters can include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing cost factor. For example, different models such as DeepSeek V4, Qwen, and Llama can correspond to different model type factors; different application scenarios such as resume creation and code generation can correspond to different scenario factors; different billing strategies can correspond to different package type factors; and resource consumption such as GPU utilization, memory usage, and power consumption can correspond to different real-time computing cost factors.
[0102] In some embodiments, the preset token conversion parameters can be dynamically adjusted based on the model resource consumption of the target model, the application scenario information corresponding to the model inference request, and the billing policy information to determine the corresponding token conversion parameters. For example, when the GPU utilization of the target model is high, the corresponding real-time computing cost factor can be increased; conversely, when the GPU utilization of the target model is low, the corresponding real-time computing cost factor can be decreased.
[0103] After determining the token conversion parameters, the corresponding unified metering token quantity can be determined based on the model's original token quantity and the token conversion parameters. The corresponding billing result is then determined based on the unified metering token quantity and the preset unit token price. Subsequently, the corresponding token amount or monetary value can be deducted from the user's account based on the billing result, while simultaneously recording the corresponding consumption details.
[0104] After completing model inference and billing processing, the API gateway can also encapsulate the inference results generated by the target model into corresponding network response data and return it to the corresponding application layer. For example, the optimized resume text can be returned to the resume creation application so that users can view the corresponding generated results.
[0105] In some embodiments, during the model inference request processing, the monitoring module in the computing resource pool can also asynchronously collect resource consumption data such as GPU utilization, CPU utilization, video memory usage, and power consumption of the corresponding computing node in the background, and associate the corresponding resource consumption data with the number of unified metering tokens corresponding to the current model inference request to generate a corresponding token cost attribution record. The token cost attribution record can be used for subsequent optimization and adjustment of token conversion parameters and for cost accounting. The above resource consumption data collection and token cost attribution process can be executed asynchronously in the background to avoid blocking the output process of the inference result corresponding to the current model inference request.
[0106] Thus, this embodiment combines the number of native tokens in the model with the corresponding token conversion parameters to uniformly measure the number of native tokens for different models. This solves the problem of inconsistent token measurement standards caused by different tokenization methods used by different models, improving the consistency and comparability of token measurement results across different models. Furthermore, this embodiment can dynamically adjust the token conversion parameters based on model resource consumption, application scenario information, and billing strategy information, thereby enhancing the flexibility and rationality of the unified token measurement results for model inference services.
[0107] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A unified measurement method for model tokens, characterized in that, include: In response to a model inference request, determine the target model corresponding to the model inference request; The inference process corresponding to the model inference request is executed through the target model, and the number of native model tokens corresponding to the inference process is determined. Obtain the Token conversion parameters corresponding to the target model; Based on the original number of tokens in the model and the token conversion parameters, the corresponding unified measurement token number is determined.
2. The model token unified measurement method as described in claim 1, characterized in that, After determining the corresponding unified measurement token quantity, the model token unified measurement method further includes: Based on the unified metering token quantity and the preset unit token price, the billing result corresponding to the model inference request is determined. Based on the billing results, deduct fees and output the inference results corresponding to the model inference request.
3. The model token unified measurement method as described in claim 1, characterized in that, The target model is selected from a preset model library, which includes at least one model using the tokenization method.
4. The model token unified measurement method as described in any one of claims 1-3, characterized in that, The step of determining the target model corresponding to the model inference request in response to the model inference request includes: In response to a model inference request, authentication processing is performed on the user who initiated the model inference request; Upon successful authentication, rate limiting detection is performed on the model inference request. Upon successful rate limiting detection, the user's billing policy information and remaining token amount information are obtained. In response to the remaining token amount information satisfying the preset amount conditions, the target model is determined based on the model inference request and the billing strategy information.
5. The unified measurement method for model tokens as described in claim 4, characterized in that, The preset limit conditions include: The remaining token amount is greater than or equal to the estimated token consumption amount corresponding to the model inference request.
6. The unified measurement method for model tokens as described in claim 4, characterized in that, The authentication process for the user initiating the model inference request in response to the model inference request includes: In response to a model inference request, user identifier, scene identifier, and user input text are extracted based on the model inference request; Based on the user identifier, the authentication process is performed on the user who initiates the model inference request.
7. The model token unified measurement method as described in claim 6, characterized in that, The process of determining the target model based on the model inference request and the billing policy information includes: Based on the scene identifier, determine the target model type corresponding to the model inference request; Based on the billing policy information, determine the model invocation permissions corresponding to the user; Based on the model access permissions and the load information of each candidate model corresponding to the target model type, the target model is determined from each of the candidate models corresponding to the target model type.
8. The model token unified measurement method as described in any one of claims 1-3, characterized in that, The step of obtaining the Token conversion parameters corresponding to the target model includes: Obtain the preset token conversion parameters corresponding to the target model from the preset parameter library, and use the preset token conversion parameters as the token conversion parameters; Alternatively, the preset Token conversion parameters can be dynamically adjusted based on at least one of the model resource consumption corresponding to the target model, the application scenario information corresponding to the model inference request, and the billing strategy information to determine the Token conversion parameters. The Token conversion parameters include at least one of the following: model type factor, scenario factor, package type factor, and real-time computing power cost factor.
9. The model token unified measurement method as described in any one of claims 1-3, characterized in that, The unified token measurement method for the model also includes: Obtain the resource consumption data corresponding to the reasoning process; The resource consumption data is associated with the number of unified metering tokens to generate a token cost attribution record; wherein the token cost attribution record is used to dynamically adjust the token conversion parameters.
10. A unified measurement system for model tokens, characterized in that, include: The application layer is used to generate model inference requests in response to user input; The Token billing engine layer communicates with the application layer. The Token billing engine layer is used to respond to the model inference request, determine the target model corresponding to the model inference request, and obtain the Token conversion parameters corresponding to the target model; A large model layer is communicatively connected to the Token billing engine layer; the large model layer includes at least one target model instance, and the large model layer is used to execute the inference process corresponding to the model inference request through the target model, and determine the number of native model tokens corresponding to the inference process; A computing resource layer, which is communicatively connected to the large model layer; the computing resource layer includes heterogeneous computing resources for deploying the target model instance; The Token billing engine layer is also used to determine the corresponding unified metering token quantity based on the original token quantity of the model and the token conversion parameters.