Gateway-based model charging method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610771921.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]ARM云手机平台能够同时承载大量用户的OpenClaw实例并发运行,每个实例均可独立发起模型调用,大模型服务的真实成本持续按令牌消耗发生,且不同云手机设备实例、不同模型、不同服务提供方、不同部署接入点之间存在显著成本差异,按次数控制或套餐控制的方式进行计费,会导致成本不可控,难以统一治理,可扩展性较差
通过在每个计算节点集群配置一个第一网关,对该计算节点集群中任一智能代理发送的大模型接入请求,估算所需令牌数量,并进行鉴权,在鉴权通过的情况下,再把该接入请求发送给第二网关,第二网关能够对所有计算节点集群的第一网关处理结果统一出口到模型厂商,由第二网关路由到实际调用模型服务的部署节点,向提供该模型服务的厂商发送调用请求,再在调用请求的返回结果中统计实际所需令牌数量,完成计费。从而能够在模型流量入口层针对多实例并发场景构建统一计费基础设施,实现针对不同云手机设备实例、不同模型服务提供方以及不同模型的实际令牌消耗进行精细化计量、鉴权、扣费与归集,提高大模型计费的灵活性和可扩展性。
Smart Images

Figure CN122845314A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of cloud phones, intelligent agents, large models, and deep learning, and specifically to a gateway-based large model billing method, device, electronic device, and storage medium. Background Technology
[0002] On the ARM cloud phone platform, users calling large model capabilities through the intelligent agent framework deployed on the cloud phone has become a core business scenario.
[0003] The ARM cloud phone platform can simultaneously support a large number of OpenClaw instances running concurrently. Each instance can independently initiate model calls. The actual cost of large model services is continuously incurred based on token consumption, and there are significant cost differences between different cloud phone device instances, different models, different service providers, and different deployment access points. Billing based on the number of calls or package pricing leads to uncontrollable costs, difficulty in unified governance, and poor scalability. Therefore, a more refined and flexible method is needed to achieve unified billing management for diverse model calls. Summary of the Invention
[0004] This disclosure aims to at least partially address one of the technical problems in the related art.
[0005] Therefore, the purpose of this disclosure is to propose a gateway-based large-scale billing method, apparatus, electronic device and storage medium to achieve refined billing based on token consumption and improve the flexibility and scalability of large-scale billing.
[0006] According to the first aspect of this disclosure, a gateway-based large-model billing method is provided, comprising:
[0007] Using the first gateway of the computing node cluster, an access request sent by any intelligent agent unit in the computing node cluster is processed to determine the request object and authentication context, wherein the request object contains the first identifier of the target calling large model; Based on the first identifier and the authentication context, estimate the number of first tokens required for the call; The access request is authenticated based on the request object, the authentication context, and the number of the first tokens to obtain the authentication result; If the authentication result is successful, the request object is sent to the second gateway, which determines the target model service to be called and sends a call request to the vendor providing the model service. The second gateway is associated with all computing node clusters. Billing is completed based on the number of second tokens in the result returned by the call request.
[0008] According to a second aspect of this disclosure, a gateway-based large-scale billing apparatus is provided, comprising: The processing module is used to process access requests sent by any intelligent agent unit in the computing node cluster using the first gateway of the computing node cluster, and to determine the request object and authentication context, wherein the request object contains a first identifier of the target calling large model. The prediction module is used to estimate the number of first tokens required for the call based on the first identifier and the authentication context; The authentication module is used to authenticate the access request based on the request object, the authentication context, and the number of the first tokens, and obtain the authentication result. The invocation module is used to send the request object to the second gateway when the authentication result is passed. The second gateway determines the target model service to be invoked and sends an invocation request to the vendor providing the model service. The second gateway is associated with all computing node clusters. The billing module is used to complete the billing based on the number of second tokens in the result returned by the call request.
[0009] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the gateway-based large-model billing method as described in the first aspect.
[0010] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the gateway-based large-model billing method as described in the first aspect.
[0011] According to a fifth aspect of this disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the gateway-based large-model billing method as described in the first aspect.
[0012] The gateway-based large-model billing method, apparatus, electronic device, and storage medium disclosed herein have the following beneficial effects: By configuring a first gateway in each compute node cluster, the number of tokens required for a large model access request sent by any intelligent agent in the cluster is estimated and authenticated. If authentication is successful, the access request is then forwarded to a second gateway. The second gateway can uniformly output the processing results of the first gateways in all compute node clusters to the model vendor. The second gateway then routes the request to the deployment node that actually calls the model service, sends a call request to the vendor providing the model service, and calculates the actual number of tokens required in the return result of the call request to complete the billing. This enables the construction of a unified billing infrastructure at the model traffic entry layer for multi-instance concurrent scenarios, achieving fine-grained metering, authentication, billing, and aggregation of actual token consumption for different cloud mobile device instances, different model service providers, and different models, improving the flexibility and scalability of large model billing.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, which are provided for a better understanding of the present invention and are not intended to limit the scope of this disclosure, wherein: Figure 1 This is a flowchart illustrating a gateway-based large-model billing method according to an embodiment of this disclosure; Figure 2 This is a flowchart illustrating a gateway-based large-model billing method according to another embodiment of this disclosure; Figure 3 This is a flowchart illustrating a gateway-based large-model billing method according to another embodiment of this disclosure; Figure 4 This is a schematic diagram of a gateway-based large-scale billing device according to an embodiment of the present disclosure; Figure 5 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0016] This disclosure relates to the fields of artificial intelligence technologies such as cloud phones, intelligent agents, large models, and deep learning.
[0017] Artificial Intelligence (AI) is a new technological science that studies, develops, and applies theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence.
[0018] Large-scale artificial intelligence models (or simply "large models") refer to a class of artificial intelligence models with a large number of parameters built from artificial neural networks. They are typically pre-trained on massive datasets using self-supervised or semi-supervised learning, and then their performance and capabilities are further optimized through fine-tuning based on instructions and human alignment. Large models are characterized by a large number of parameters, large training data sets, and large computational resources, and possess the ability to solve general tasks, follow human instructions, and perform complex reasoning.
[0019] Cloudphone, or simply cloud, is a cloud service that uses edge-cloud integrated virtualization technology and digital capabilities such as cloud network, security, and AI to flexibly adapt to users' personalized needs, free up the hardware resources of the phone itself, and load massive cloud applications on demand. It is a cloud service that has an operating system and virtual phone functions, and can run complex calculations and large amounts of data in the cloud.
[0020] An intelligent agent is a proxy capable of perceiving its environment and taking actions to achieve specific goals. It can be software, hardware, or a system, possessing autonomy, adaptability, and interactivity. Intelligent agents perceive changes in the environment (e.g., through sensors or data input), make judgments and decisions based on their learned knowledge and algorithms, and then execute actions to influence the environment or achieve predetermined goals. Intelligent agents are widely used in the field of artificial intelligence, commonly found in automated systems, robots, virtual assistants, and game characters. Their core strength lies in their ability to learn autonomously and continuously evolve to better complete tasks and adapt to complex environments.
[0021] Deep learning learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly aids in interpreting data such as text, images, and sound. The ultimate goal of deep learning is to enable machines to possess analytical and learning capabilities similar to humans, allowing them to recognize data such as text, images, and sound.
[0022] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0023] The following description, with reference to the accompanying drawings, outlines a gateway-based large-model billing method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure.
[0024] It should be noted that the execution entity of the gateway-based large model billing method in this embodiment is a gateway-based large model billing device. This device can be implemented by software and / or hardware. The device can be configured in an electronic device, which may include, but is not limited to, a terminal, a server, etc.
[0025] Figure 1 This is a flowchart illustrating a gateway-based large-model billing method according to an embodiment of this disclosure.
[0026] like Figure 1 As shown, this gateway-based large-model billing method includes: S101: Using the first gateway of the computing node cluster, process the access request sent by any intelligent agent unit in the computing node cluster, and determine the request object and authentication context.
[0027] Among them, a computing node cluster refers to a computer room or other independent network domain.
[0028] This disclosure embodiment can be applied to multi-datacenter scenarios such as enterprise intranets, where networks, Domain Name System (DNS), gateways, and data are isolated between multiple datacenters, and each datacenter corresponds to an independent cloud mobile platform (or a specific region of the cloud platform). Each cloud mobile platform (i.e., a computing node cluster) can create and manage multiple intelligent agent units.
[0029] It should be noted that in this disclosure, the available vendors and / or available models can be configured independently for different computing node clusters.
[0030] The first gateway, also known as the model gateway, is used to perform identity authentication, access control, model routing, access rate limiting, token metering, cost calculation, and log auditing.
[0031] Among them, the intelligent agent unit refers to the OpenClaw instance configured in the cloud phone that can independently initiate model calls and process tasks.
[0032] In this embodiment of the disclosure, multiple intelligent agent units created within the same computing node cluster are isolated from each other. The intelligent agent units can read the identity token (JSON Web Token, JWT) of the ARM architecture cloud phone, read symmetric encrypted files, parse user IDs or cloud phone tenant information, and generate authentication signatures.
[0033] In this embodiment, any intelligent agent unit running on an ARM cloud phone within the computing node cluster can access the same preset domain name (e.g., redclaw-llm) through a terminal access plugin. The access plugin is responsible for token parsing, encrypted signing, and request context encapsulation, enabling the OpenClaw instance of the intelligent agent unit on the cloud phone side to complete standardized access without directly storing external model service access credentials. Then, the unified domain name is resolved to the corresponding intranet access address of the computing node cluster through DNS domain name redirection, thereby achieving controlled traffic redirection and proximity access for model requests in the enterprise intranet environment. This allows model traffic to preferentially enter the model gateway node in the local data center for processing, so that the first gateway of the computing node cluster can process access requests sent by any intelligent agent unit in the computing node cluster.
[0034] It should be noted that in this disclosure, the access request can be an OpenAI-compatible protocol request, which can seamlessly switch between different model service providers without being bound to a specific model vendor. Flexible scheduling and dynamic replacement of models can be fully achieved through the gateway layer.
[0035] The request object may include, but is not limited to, model identifier, message body or input body, whether it is streaming, tool definition, structured output constraints, user identifier, and extended metadata.
[0036] In this embodiment of the disclosure, the first gateway can uniformly receive access requests sent by the intelligent agent unit through the protocol adaptation layer. The access request can include at least the request form of chat / completions, responses, and vector embeddings. The first gateway can form a standardized request object for the access request.
[0037] It should be noted that in this disclosure, the first gateway can parse the actual hit model after generating a standardized request object. The request object contains the first identifier of the target calling large model. The first identifier can be a unified model name exposed externally. For example, the unified model name can be a general term for multiple models from the same vendor, or it can be a type of logical model, etc., so that the terminal or business system can only perceive the unified model name, and the specific deployment nodes can intelligently route according to cost or performance.
[0038] In this embodiment, the first gateway does not directly rely on the private fields of any external model service. Instead, it extracts common fields from the compatible protocol in a unified manner and merges context information from different levels into a unified authentication context. This authentication context includes at least the user identifier, session identifier, terminal device identifier, terminal code, user terminal association identifier, and access token in the request header, the real source address information injected by the proxy layer, and information such as user, metadata, and extra request headers in the request body. This ensures that the terminal plugin header, proxy layer forwarding header, and gateway internal injection header are processed consistently in the same standard object, avoiding inconsistencies in authentication and billing standards caused by multiple entry points, multiple proxies, and multiple terminal forms.
[0039] In this disclosure, after a standardized request object is formed, the request header, metadata, and business context can be normalized, and token pre-estimation and business authentication can be performed to transform the access request into a model call request that can be uniformly governed.
[0040] S102: Based on the first identifier and authentication context, estimate the number of first tokens required for the call.
[0041] In this embodiment of the disclosure, input token pre-estimation can be performed before model invocation, and the actual token usage can be extracted after model invocation. The two can be used for pre-verification and post-settlement, respectively.
[0042] It should be noted that in this disclosure, the estimation of the first token quantity before the call can adopt a dual-channel mechanism. Depending on whether the upstream external model service corresponding to the first identifier supports the official token counting interface, different methods can be selected to estimate the first token quantity, thereby improving the accuracy of the token estimation before the call and providing a more reliable data foundation for subsequent authentication.
[0043] S103: Authentication is performed on the access request based on the request object, authentication context, and number of first tokens to obtain the authentication result.
[0044] In this embodiment of the disclosure, the business authentication service can be invoked to complete the verification of balance, permissions and device legality based on the request object, authentication context and the number of first tokens. For example, if a user has a balance of 'a' of the number of available tokens for the model of the first identifier and the number of first tokens is 'b', the balance verification passes if 'a' is greater than 'b', thereby obtaining the authentication result of the access request.
[0045] S104: If the authentication result is successful, the request object is sent to the second gateway, which determines the model service to be called and sends a call request to the vendor providing the model service.
[0046] The second gateway is the unified exit gateway.
[0047] In this disclosure, the second gateway is associated with all compute node clusters. The processing results of access requests from the first gateways in multiple compute node clusters can be aggregated to the second gateway. The second gateway then routes the request to the actual external model service based on the first identifier in the request object, and initiates a call request based on information such as the input body, whether it is a streaming response, and output constraints in the request object.
[0048] It should be noted that the second gateway can include whitelist control, access auditing, cost scheduling, fault rollback, and external egress governance functions. It can be used to uniformly manage the access behavior of the public network model at the enterprise security boundary, thereby converging the attack surface, reducing exposure, and greatly reducing the risk of the enterprise's internal network being penetrated.
[0049] In this publicly disclosed public network model service environment, the second gateway can simultaneously connect to multiple model service providers and perform unified scheduling across different providers, models, and service capabilities. When a model response is returned, it can be transmitted to the client in a lossless manner. The second gateway further aggregates token usage, request identifiers, and latency information to form unified billing and auditing records, thereby enabling flexible switching and unified orchestration in multi-vendor, multi-model scenarios.
[0050] It should be noted that if the model call fails after the call request is sent, exception classification, error logging, and failure response logging can be performed, and a failure result can be returned.
[0051] S105: Complete billing based on the number of second tokens in the request response.
[0052] The second token quantity is the number of tokens actually required for the call.
[0053] In this embodiment, the actual token statistics after the call prioritize extracting fields such as input token prompt_tokens, output token completion_tokens, and total tokens from the upstream returned results to obtain the second token count. This second token count can then be reported to the unified billing interface to complete the billing process.
[0054] In this embodiment of the disclosure, when the request object indicates a streaming return, the streaming response processing unit can perform lossless pass-through without changing the semantics of the original data block output, and simultaneously monitor client disconnection, cancellation and abnormal situations, obtain token usage information at the end of the streaming process or perform supplementary aggregation in the log to determine the second token quantity.
[0055] It should be noted that after determining the number of the second tokens, the number of the second tokens can be compared and calibrated with the estimated number of the first tokens.
[0056] In this embodiment, a first gateway is configured in each compute node cluster. This gateway estimates the required number of tokens for any large model access request sent by any intelligent agent in the cluster and performs authentication. If authentication is successful, the access request is then sent to a second gateway. The second gateway can uniformly output the processing results of the first gateways across all compute node clusters to the model vendor. The second gateway then routes the request to the deployment node that actually calls the model service, sends a call request to the vendor providing the model service, and calculates the actual number of tokens required in the return result of the call request to complete the billing. This enables the construction of a unified billing infrastructure at the model traffic entry layer for multi-instance concurrent scenarios. It allows for refined metering, authentication, billing, and aggregation of actual token consumption for different cloud mobile device instances, different model service providers, and different models, improving the flexibility and scalability of large model billing.
[0057] It should be noted that when the first gateway proxies traffic for compatible protocols, it does not destructively rewrite critical business semantic fields in the request and response bodies. It maintains at least the transparent or equivalent transmission of message structure, tool calls, response format, structured output constraints, streaming data block order, request identifier, token usage fields, and the model response body. The first gateway only supplements observation fields, authentication fields, and streaming token usage injection fields when necessary, without altering the core protocol semantics between the client and the upstream model. This achieves a transparent gateway capability that allows for fine-grained billing without affecting existing Software Development Kits (SDKs), plugins, and compatible client access methods.
[0058] In this disclosure, since the terminal only accesses the unified domain name of the enterprise intranet and the access point in the data center, and does not directly access the public network model service address, nor directly expose the vendor key, routing policy and billing logic, it can reduce security risks such as key leakage, terminal bypassing the gateway, unauthorized direct connection to the public network model interface and uncontrolled cost.
[0059] The gateway-based large model billing method disclosed herein can unify the model call entry point through the terminal access plugin on the OpenClaw instance side deployed on ARM cloud phones, complete traffic takeover on the network side through controlled DNS redirection, intranet address mapping and proxy access, and achieve a complete closed loop on the data center model gateway node and the central unified exit gateway side through standardized request access, unified authentication, unified routing, usage aggregation, billing reporting and failure compensation.
[0060] Figure 2 This is a flowchart illustrating a gateway-based large-model billing method proposed in another embodiment of this disclosure.
[0061] like Figure 2 As shown, this gateway-based large-model billing method includes: S201: Using the first gateway of the compute node cluster, process the access request sent by any intelligent agent unit in the compute node cluster, and determine the request object and authentication context.
[0062] Optionally, the access request can be standardized first to determine the request object. The access request can include at least a dialog, response, and vector-embedded request format. Then, the target fields of the request header, metadata, and business context in the access request are normalized to obtain the authentication context.
[0063] The target field is a common field in the compatibility protocol and is consistent across different external model services.
[0064] In this embodiment of the disclosure, the first gateway can uniformly receive access requests sent by the intelligent agent unit through the protocol adaptation layer. The access request can include at least the request form of chat / completions, responses, and vector embeddings. The first gateway can form a standardized request object for the access request.
[0065] In this embodiment, the first gateway does not directly rely on the private fields of any external model service. Instead, it extracts common fields from the compatible protocol in a unified manner and merges context information from different levels into a unified authentication context. This authentication context includes at least the user identifier, session identifier, terminal device identifier, terminal code, user terminal association identifier, and access token in the request header, the real source address information injected by the proxy layer, and information such as user, metadata, and extra request headers in the request body. This ensures that the terminal plugin header, proxy layer forwarding header, and gateway internal injection header are processed consistently in the same standard object, avoiding inconsistencies in authentication and billing standards caused by multiple entry points, multiple proxies, and multiple terminal forms.
[0066] This disclosure improves the performance and efficiency of subsequent authentication, routing, and billing processes by processing access requests sent from different computing node clusters and different proxy units into unified standard data, thereby unifying observability and facilitating refined auditing.
[0067] Optionally, it can receive access requests sent by any intelligent agent unit through accessing a preset domain name, perform targeted resolution on the preset domain name based on the second identifier of the computing node cluster, determine the access address of the first gateway, and then send the access request to the first gateway based on the access address.
[0068] In this disclosure, a preset domain name is resolved to the gateway of each computing node cluster by internal network DNS hijacking. In the enterprise intranet environment, controlled traffic redirection and proximity access for model requests are realized, so that model traffic can be preferentially processed by the model gateway node of this computing node cluster, thereby achieving forced traffic control.
[0069] S202: If the model service corresponding to the first identifier meets the preset conditions, the token counting interface associated with the first identifier is called according to the authentication context to determine the number of first tokens.
[0070] In this embodiment of the disclosure, the preset condition is that the upstream external model service corresponding to the first identifier supports the official token counting interface. When the official token counting interface is supported, that is, the official token counting interface can be directly called to obtain a high-precision input token result and determine the number of the first tokens.
[0071] Alternatively, if the model service corresponding to the first identifier does not meet the preset conditions, the token segmenter of the first gateway is invoked according to the authentication context to determine the number of first tokens.
[0072] In this embodiment of the disclosure, the failure to meet the preset conditions means that the upstream external model service corresponding to the first identifier does not support the official counting interface, or only provides compatible but not completely equivalent proxy capabilities. When the preset conditions are not met, the system can fall back to the local tokenizer of the first gateway to perform token estimation on the authentication context and determine the number of first tokens.
[0073] It should be noted that in this disclosure, model-level calibration coefficients can be configured for different models to correct systematic deviations between different tokenizer implementations.
[0074] S203: Authentication is performed on the access request based on the request object, authentication context, and number of first tokens to obtain the authentication result.
[0075] S204: If the authentication result is successful, the request object is sent to the second gateway, which determines the model service to be called and sends a call request to the vendor providing the model service.
[0076] S205: Complete billing based on the number of second tokens in the request response.
[0077] For a detailed description of steps S203 to S205 above, please refer to other embodiments of this disclosure, which will not be repeated here.
[0078] In this embodiment, by adopting a dual-channel mechanism to perform token estimation before the call, the accuracy of the token estimation result can be improved, providing a more reliable basis for the gateway to intercept overdraft risks in advance.
[0079] Figure 3 This is a flowchart illustrating a gateway-based large-model billing method proposed in another embodiment of this disclosure.
[0080] like Figure 3 As shown, this gateway-based large-model billing method includes: S301: Using the first gateway of the compute node cluster, process the access request sent by any intelligent agent unit in the compute node cluster, and determine the request object and authentication context.
[0081] S302: Estimate the number of first tokens required for the call based on the first identifier and authentication context.
[0082] For a detailed description of steps S301 to S302 above, please refer to other embodiments of this disclosure, which will not be repeated here.
[0083] S303: Based on the mapping relationship in the second gateway at the current time and the first identifier in the request object, determine the second identifier and the deployment node of the model corresponding to the second identifier.
[0084] In this embodiment of the disclosure, the second gateway can maintain a three-layer mapping relationship of "external model alias (i.e., first identifier) - internal model identifier (i.e., second identifier) - actual deployment node". The terminal or business system only perceives the first identifier, while the second gateway can map the request to the real model second identifier and inference access point of the specific external model service according to the routing configuration.
[0085] Optionally, the process of determining the mapping relationship may include at least one of the following: for each first identifier corresponding to multiple candidate second identifiers, determine the target second identifier corresponding to the first identifier at the current time based on at least one of the model cost, current performance, and region to which each candidate second identifier belongs; update the target second identifier corresponding to the first identifier based on the status data of the deployment node corresponding to each candidate second identifier; and determine the target second identifier corresponding to the first identifier based on the authentication context.
[0086] In this embodiment of the disclosure, the first identifier can be mapped to equivalent implementations of different model service providers to achieve dynamic routing based on cost priority, performance priority, or region priority.
[0087] For example, when the goal is to reduce expenses, the mapping relationship will direct requests to the equivalent model service provider with a lower unit price or caching capabilities; or when pursuing low latency or high throughput, the mapping relationship will lock the deployment node with the fastest response speed and the lowest load; or it can be based on the principle of network proximity to map requests to the actual deployment node that is geographically closest to the caller to reduce network transmission time.
[0088] The status data may include whether an anomaly has occurred, the type of anomaly, and the real-time concurrency or quota usage.
[0089] In this embodiment of the disclosure, routing control may also support fault rollback, load balancing, deployment-based switching, and retrying based on anomaly type. For example, when an actual deployment node on the mapping path experiences an anomaly (such as timeout or triggering rate limiting), the mapping relationship can automatically switch to a pre-configured backup model or service address.
[0090] In this embodiment of the disclosure, the scope of the mapping relationship is often limited by the identity information carried in the request. Therefore, based on the authentication context, it can support unified configuration according to dimensions such as model group, deployment instance, service address, access key, retry policy, and maximum number of retries to determine the target second identifier corresponding to the first identifier. For example, if a business line or user's token quota is about to be exhausted, the mapping relationship may automatically redirect it to a low-cost model or trigger an interception.
[0091] In this embodiment of the disclosure, attribution during the billing and auditing phases is based on the actual model identifier that was hit, rather than solely on the model name requested by the client.
[0092] In this embodiment, by using multi-dimensional governance strategies and dynamic decision-making rules, the mapping relationship between the unified model name exposed to the outside world and the actual model service called can be determined, thereby achieving more comprehensive and reliable intelligent dynamic routing.
[0093] S304: Determine the model service to be invoked based on the deployment node.
[0094] In this embodiment, by utilizing the mapping relationship maintained within the gateway, the true second identifier of the specific external model service mapped to the first identifier of the currently requested model and the deployment node are determined. This allows the system to identify the target model service being called, enabling it to initiate requests only to a unified model name without needing to be aware of the underlying model vendor, specific version, or physical deployment location. When it is necessary to change the model vendor, upgrade the model version, or migrate the deployment node, only the mapping configuration needs to be modified at the gateway layer, greatly reducing the system's coupling and providing high flexibility, scalability, and security controllability even in complex production environments.
[0095] S305: If the authentication result is successful, the request object is sent to the second gateway, which determines the model service to be called and sends a call request to the vendor providing the model service.
[0096] S306: Complete billing based on the number of second tokens in the request response.
[0097] Optionally, the number of second tokens can be reported to the billing interface first. Then, if the billing interface fails, a compensation log can be generated based on the model service, authentication context, and the number of second tokens. Subsequently, an exception retry can be performed based on the compensation log at preset time intervals.
[0098] In this embodiment of the disclosure, if the billing interface fails, the model service, authentication context, and number of second tokens can be further written into the compensation log. Then, when the preset time interval is reached, an independent compensation task will perform an asynchronous retry to complete the deduction. Thus, a complete accounting loop can be formed through pre-call authentication, post-call billing, and failure compensation to ensure eventual accounting consistency.
[0099] For a detailed description of steps S305 and S306 above, please refer to other embodiments of this disclosure, which will not be repeated here.
[0100] The gateway-based large-model billing method proposed in this disclosure, compared to existing solutions that only control whether something can be used, can control which OpenClaw instance on which cloud phone is using it, what model is being used, how many tokens are being used, how much should be deducted, and how to recharge after a failure. Furthermore, existing solutions implement membership on / off at the product layer, while this invention builds a unified billing infrastructure at the model traffic entry layer for concurrent multi-instance scenarios on ARM cloud phones.
[0101] Figure 4 This is a schematic diagram of the structure of a gateway-based large-scale billing device proposed in one embodiment of the present disclosure.
[0102] like Figure 4 As shown, the gateway-based large-model billing device 40 includes: Processing module 401 is used to process access requests sent by any intelligent agent unit in the computing node cluster using the first gateway of the computing node cluster, and determine the request object and authentication context, wherein the request object contains the first identifier of the target calling large model. Prediction module 402 is used to estimate the number of first tokens required for the call based on the first identifier and authentication context; The authentication module 403 is used to authenticate the access request based on the request object, authentication context and the number of first tokens, and obtain the authentication result. The calling module 404 is used to send the request object to the second gateway if the authentication result is successful. The second gateway determines the target model service to be called and sends a call request to the vendor providing the model service. The second gateway is associated with all computing node clusters. The billing module 405 is used to complete the billing based on the number of second tokens in the result returned by the call request.
[0103] In some possible embodiments, the prediction module 402 may specifically be used for: Standardize access requests and determine the request object, wherein the access request includes at least dialogue, response and vector embedding request forms; The target fields of the request header, metadata, and business context in the access request are normalized to obtain the authentication context.
[0104] In some possible embodiments, the prediction module 402 may specifically be used for: Receive access requests sent by any intelligent agent unit through accessing a preset domain name; Based on the second identifier of the computing node cluster, the preset domain name is resolved to determine the access address of the first gateway; Based on the access address, the access request is sent to the first gateway.
[0105] In some possible embodiments, the prediction module 402 may specifically be used for: If the model service corresponding to the first identifier meets the preset conditions, the token counting interface associated with the first identifier is called according to the authentication context to determine the number of first tokens; or, If the model service corresponding to the first identifier does not meet the preset conditions, the token segmenter of the first gateway is invoked according to the authentication context to determine the number of first tokens.
[0106] In some possible embodiments, module 404 is invoked, which can specifically be used for: Based on the mapping relationship in the second gateway at the current moment and the first identifier in the request object, determine the second identifier and the deployment node of the model corresponding to the second identifier; Based on the deployment node, determine the model service to be invoked by the target.
[0107] In some possible embodiments, the process of determining the mapping relationship includes at least one of the following: For each first identifier corresponding to multiple candidate second identifiers, the target second identifier corresponding to the first identifier at the current time is determined based on the model cost, current time performance, and at least one of the regions to which each candidate second identifier belongs. Based on the status data of the deployment node corresponding to each candidate second identifier, update the target second identifier corresponding to the first identifier; Based on the authentication context, determine the target second identifier corresponding to the first identifier.
[0108] In some possible embodiments, the billing module 405 may specifically be used for: Report the number of the second token to the billing interface; In the event of a billing interface failure, a compensation log is generated based on the model service, authentication context, and the number of second tokens. Based on a preset time interval, exceptions are retried using compensation logs.
[0109] It should be noted that the foregoing explanation of the gateway-based large model billing method also applies to the gateway-based large model billing device of this embodiment, and will not be repeated here.
[0110] In this embodiment, a first gateway is configured in each compute node cluster. This gateway estimates the required number of tokens for any large model access request sent by any intelligent agent in the cluster and performs authentication. If authentication is successful, the access request is then sent to a second gateway. The second gateway can uniformly output the processing results of the first gateways across all compute node clusters to the model vendor. The second gateway then routes the request to the deployment node that actually calls the model service, sends a call request to the vendor providing the model service, and calculates the actual number of tokens required in the return result of the call request to complete the billing. This enables the construction of a unified billing infrastructure at the model traffic entry layer for multi-instance concurrent scenarios. It allows for refined metering, authentication, billing, and aggregation of actual token consumption for different cloud mobile device instances, different model service providers, and different models, improving the flexibility and scalability of large model billing.
[0111] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0112] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0113] like Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0114] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0115] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the gateway-based big model billing method. For example, in some embodiments, the gateway-based big model billing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the gateway-based big model billing method described above can be performed. Alternatively, in other embodiments, computing unit 501 may be configured to perform a gateway-based large-model billing method by any other suitable means (e.g., by means of firmware).
[0116] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0117] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0120] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0121] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0122] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0123] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. In the description of this disclosure, the words "if" and "suppose" as used may be interpreted as "when," "when," "in response to determination," or "in the circumstances."
[0124] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A gateway-based large-scale billing method, characterized in that, include: Using the first gateway of the computing node cluster, an access request sent by any intelligent agent unit in the computing node cluster is processed to determine the request object and authentication context, wherein the request object contains the first identifier of the target calling large model; Based on the first identifier and the authentication context, estimate the number of first tokens required for the call; The access request is authenticated based on the request object, the authentication context, and the number of the first tokens to obtain the authentication result; If the authentication result is successful, the request object is sent to the second gateway, which determines the target model service to be called and sends a call request to the vendor providing the model service. The second gateway is associated with all computing node clusters. Billing is completed based on the number of second tokens in the result returned by the call request.
2. The method as described in claim 1, characterized in that, The first gateway of the computing node cluster processes access requests sent by any intelligent agent unit in the computing node cluster, and determines the request object and authentication context, including: The access request is standardized to determine the request object, wherein the access request includes at least dialogue, response and vector-embedded request forms; The authentication context is obtained by normalizing the target fields of the request header, metadata, and business context in the access request.
3. The method as described in claim 2, characterized in that, Before the first gateway of the computing node cluster processes the access request sent by any intelligent agent unit in the computing node cluster and determines the request object and authentication context, the method further includes: Receive an access request sent by any of the intelligent agent units through accessing a preset domain name; Based on the second identifier of the computing node cluster, the preset domain name is resolved in a targeted manner to determine the access address of the first gateway; Based on the access address, the access request is sent to the first gateway.
4. The method as described in claim 1, characterized in that, The step of estimating the number of first tokens required for the call based on the first identifier and the authentication context includes: If the model service corresponding to the first identifier meets preset conditions, the token counting interface associated with the first identifier is called according to the authentication context to determine the number of the first tokens; or, If the model service corresponding to the first identifier does not meet the preset conditions, the token segmenter of the first gateway is invoked according to the authentication context to determine the number of the first tokens.
5. The method as described in claim 1, characterized in that, The step of sending the request object to the second gateway, whereby the second gateway determines the model service to be invoked, includes: Based on the mapping relationship in the second gateway at the current moment and the first identifier in the request object, determine the second identifier and the deployment node of the model corresponding to the second identifier; Based on the deployment node, determine the model service to be invoked by the target.
6. The method as described in claim 5, characterized in that, The process of determining the mapping relationship includes at least one of the following: For each first identifier corresponding to multiple candidate second identifiers, the target second identifier corresponding to the first identifier at the current time is determined based on the model cost, current time performance, and at least one of the regions to which each candidate second identifier belongs. Based on the status data of the deployment node corresponding to each candidate second identifier, update the target second identifier corresponding to the first identifier; Based on the authentication context, a target second identifier corresponding to the first identifier is determined.
7. The method as described in claim 1, characterized in that, The step of completing billing based on the number of second tokens in the result returned by the call request includes: Report the number of the second token to the billing interface; In the event of a failure of the billing interface, a compensation log is generated based on the model service, the authentication context, and the number of the second tokens. Based on a preset time interval, exception retries are performed based on the compensation log.
8. A gateway-based large-scale billing device, characterized in that, include: The processing module is used to process access requests sent by any intelligent agent unit in the computing node cluster using the first gateway of the computing node cluster, and to determine the request object and authentication context, wherein the request object contains a first identifier of the target calling large model. The prediction module is used to estimate the number of first tokens required for the call based on the first identifier and the authentication context; The authentication module is used to authenticate the access request based on the request object, the authentication context, and the number of the first tokens, and obtain the authentication result. The invocation module is used to send the request object to the second gateway when the authentication result is passed. The second gateway determines the target model service to be invoked and sends an invocation request to the vendor providing the model service. The second gateway is associated with all computing node clusters. The billing module is used to complete the billing based on the number of second tokens in the result returned by the call request.
9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the gateway-based large-model billing method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to execute the gateway-based large-model billing method according to any one of claims 1-7.
11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the gateway-based large-model billing method according to any one of claims 1-7.