A standardized pricing and transaction system, method, and apparatus for AI inference services.
By constructing an AI service unit computing module, a hybrid transaction execution module, and a regulatory interface protocol, the problems of inconsistent value measurement and price distortion in the AI inference service market have been solved. This has enabled standardized pricing and transactions across modalities and platforms, improving market resource allocation efficiency and risk management capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-31
AI Technical Summary
The existing AI inference service market suffers from problems such as the inability to uniformly measure value across modalities and platforms, price distortion, uneven resource allocation, and risk exposure. The lack of standardized pricing and transaction systems leads to waste of market resources and difficulties in risk management.
The AI service unit computing module is constructed for quality assessment and unified measurement. Combined with the hybrid transaction execution module, the computing node marginal pricing module, and the institutional interface protocol, standardized pricing and trading across modalities and platforms are realized. Transaction orders are processed through a central price limit order ledger and an automated market maker mechanism. A dynamic risk value model is used for settlement. Institutional interface protocols are established for the technical abstraction layer, the market incentive layer, and the regulatory adaptation layer.
It enables unified measurement of AI service value across modalities and platforms, reduces price information distortion, improves market liquidity and trading efficiency, enhances market robustness and operability, and provides real-time scarcity signal transmission and compliance audit capabilities.
Smart Images

Figure CN122492276A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of pricing and trading of artificial intelligence computing power services, design of computing markets, microstructure of financial markets and decentralized finance, specifically to a standardized pricing and trading system, method and apparatus for AI inference services. Background Technology
[0002] Artificial intelligence (AI) inference services have become a key production factor in the digital economy, with applications rapidly expanding from early text generation and understanding to a wide range of fields including image generation, video synthesis, speech recognition and synthesis, and multimodal content understanding. As of the date of this patent application, the global AI inference market has reached an annual transaction volume of tens of billions of US dollars, with enterprise clients consuming trillions of inference service calls daily. However, this rapid market growth is accompanied by significant structural deficiencies in AI inference services at the infrastructure level, including transaction mechanisms, pricing transparency, and cross-modal value measurement, hindering the effective allocation of resources and the further healthy development of the industry.
[0003] The current AI inference service market generally adopts a closed model with independent pricing and metering across platforms. Different service providers set prices based on their own cost structures, brand positioning, and competitive strategies, and the billing unit also varies depending on the modality: text services are often priced per thousand or per million tokens, image generation services are often billed on a tiered basis based on a single image or output resolution, video generation services are billed based on a combination of duration, frame rate, or resolution, and voice services are charged based on audio duration or number of characters. This heterogeneity in billing units makes it difficult for users to make intuitive cost comparisons and value judgments between services on different platforms and in different modalities. A more prominent problem is that even within the same modality, due to the lack of a unified quality assessment and adjustment framework, the price difference between services with equivalent technical capabilities can be tens or even hundreds of times, resulting in a severe distortion of market price signals.
[0004] Further analysis reveals that the shortcomings of existing technologies can be summarized into three interrelated levels. At the resource allocation level, the average utilization rate of GPUs in data centers and computing nodes globally has long been low, with many newly built computing facilities achieving utilization rates of less than 30%. The idleness and waste of computing resources are extremely prominent, stemming from the lack of real-time price signals and scheduling mechanisms to guide inference requests dynamically to the optimal nodes. At the price formation mechanism level, the pricing of AI inference services fails to systematically incorporate multiple factors such as quality, latency, credit risk, and request scale into a unified standardized framework. This leads to information asymmetry for users when selecting services, significant errors in AI cost prediction for enterprise clients, and a lack of effective risk management tools. At the market completeness level, the existing market is limited to spot trading on various platforms, lacking the institutional arrangements and infrastructure necessary for mature financial markets, such as forward contracts, derivatives hedging, cross-platform arbitrage, and automated clearing and settlement. This exposes both the supply and demand sides of computing services to unpredictable market volatility risks.
[0005] The rapid emergence and commercialization of multimodal AI models have exacerbated the severity of these problems. The technical paths and resource consumption characteristics of text, image, video, voice, and multimodal services differ significantly, and current technologies lack a unified benchmark and pricing protocol for measuring the value of AI services across modalities. When users need to manage their spending on AI services across multiple modalities, they can only rely on fragmented bills and experience-based judgments, lacking a unified and comparable standardized unit of measurement. Simultaneously, regulatory agencies and industry self-regulatory organizations lack a verifiable technical interface for implementing data residency compliance, tiered asset supervision, and privacy protection audits.
[0006] In summary, the AI inference service market urgently needs a systematic solution that can achieve standardized measurement, transparent pricing, efficient trading, and automated clearing across modalities and platforms to address key issues in the current market such as resource misallocation, price distortion, excessive risk exposure, and fragmented multimodal pricing. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a standardized pricing and transaction system, method, and apparatus for AI inference services, thereby resolving the problems described in the background.
[0008] To achieve the above objectives, the present invention provides the following technical solution: Firstly, a standardized pricing and transaction system for AI inference services includes: The AI service unit computing module is used to collect raw performance scores from third-party benchmarking systems in the corresponding domains for AI inference services of different modalities, and generate a unified quality index through a quality assessment model adapted to the modality. Based on the original billing price of each service and the aforementioned quality index, combined with the mode-related congestion factor... Calculate standardized AI service unit ASU values to establish a cross-modal, cross-platform comparable price measurement benchmark; The hybrid trading execution module is used to process trading orders for the assets associated with the AI service unit in parallel based on the Central Limit Order Book (CLOB) and the Automated Market Maker (AMM) mechanism; wherein, the CLOB generates quotes based on the market maker's optimal control model, and the AMM mechanism maintains liquidity and calculates impermanent loss based on a constant product formula; The C-LMP module for computing node marginal pricing is used to solve scheduling optimization problems with capacity and bandwidth constraints for AI inference computing networks with the goal of maximizing social net surplus. It solves the problem iteratively using the Alternating Direction Multiplier Method (ADMM), and after convergence, it directly decomposes the system marginal cost, capacity congestion premium, and bandwidth premium from the dual variables to form the node price and broadcast it to the system. The Institutional Interface Protocol (IIP) module, as middleware, provides a technical abstraction layer, a market incentive layer, and a regulatory adaptation layer. It is used to map heterogeneous billing rules and congestion factors of different modalities into standardized AI service units, execute capacity pre-sales and secondary royalty allocation, and implement data residency and compliance audits. The derivatives and clearing module is used to calculate the trading margin based on the real-time price and position data output by the hybrid trading execution module, through a dynamic risk value model, and to execute T+0 settlement through a smart contract.
[0009] Preferably, the process by which the AI service unit calculation module generates the quality index includes: Determine the modality type of the AI inference service to be evaluated, and select at least one third-party benchmark evaluation indicator for the corresponding capability dimension. The capability dimension includes, but is not limited to, language understanding, code generation, mathematical reasoning, and image quality assessment. The modality type includes text, image, video, voice, and multimodal. The raw performance scores of the AI service model on the capability dimension are collected from the selected third-party benchmarking system; The raw performance scores collected are standardized to eliminate the dimensional differences between different indicators; The standardized scores are then synthesized into a comprehensive quality index with values ranging from [0,1] using principal component analysis or weighted aggregation methods.
[0010] Preferably, the third-party benchmark evaluation indicators include, but are not limited to: For text modalities, at least benchmark metrics in the fields of language understanding, code generation, or mathematical reasoning should be included; For image generation modalities, at least benchmark evaluation metrics in the fields of image fidelity or image-text matching should be included; For video generation modalities, at least benchmark evaluation metrics in the fields of video fidelity or temporal consistency should be included; For multimodal models, at least benchmark metrics in the areas of cross-modal understanding and generation capabilities should be included.
[0011] Preferably, the AI service unit calculation module calculates the standardized AI service unit value. The method is as follows: ;in, The original billing price for the AI inference service is uniformly converted into a comparable price based on equivalent token units through a preset modality conversion factor. The modality conversion factor is determined by using the unit output price of a specified reference platform at a base time as an anchor point and the market price ratio method to inversely calculate the equivalent token quantity corresponding to each modality unit output. The quality index is mentioned above. This refers to the mass elasticity parameter; The delay factor is... This refers to the delay elasticity parameter; The credit premium is measured on a logarithmic scale. This is a parameter related to credit premium sensitivity. The congestion factor characterizes the platform's real-time load level. This refers to the congestion resilience parameter.
[0012] Preferably, the scheduling optimization problem solved by the computing node marginal pricing module is based on the amount of inference requests allocated to each computing node. The decision variables are: demand utility and total cost at each node. The objective function is to maximize the difference between demand utility and total cost at each node. Constraints include node capacity limits. and bandwidth limit ; The node price The dual variables obtained from solving this problem are directly derived, and their decomposition is: ,in For the system's marginal cost, For capacity congestion premium, This is a bandwidth premium.
[0013] Preferably, the derivatives and clearing module calculates the dynamic minimum margin. The formula is: ,in, Value at Risk (VaR) This is the position ratio buffer factor. The value of holdings calculated based on historical simulation methods within a specific time window The maximum possible loss at a 99% confidence level. Real-time median price Secondly, a standardized pricing and transaction method for AI inference services includes the following steps: S1. Data Acquisition and Quality Index Generation: For AI inference services of different modalities, raw performance scores are collected from third-party benchmark evaluation systems in the corresponding fields, covering multiple dimensions including but not limited to language understanding, code generation, mathematical reasoning, and image quality assessment. The collected raw performance scores are standardized, and a comprehensive quality index with a value range of [0,1] is generated through principal component analysis or weighted aggregation methods. ; S2. Parameter Estimation: Obtain the original billing prices of each AI inference service and convert them into comparable prices based on equivalent token units using a preset modality conversion factor; based on the hedonic regression model. Quality index Delay coefficient Credit premium and congestion factor Estimate the elasticity parameters. For comparable prices, > 0 represents the market clearing constant. for ; S3. ASU Calculation: Using the estimated parameters, through the formula... Calculate the standardized AI service unit value for each service; S4. Exchange Rate Generation and Arbitrage Detection: Generate cross-platform exchange rates based on the AI service unit value. ,in and They represent the first The and the first The standardized service unit value of an AI inference service platform, and generates arbitrage signals when the exchange rate deviates from a preset threshold; S5. Transaction Execution: Receive transaction orders for the assets associated with the AI service unit and match transactions based on a hybrid mode of central limit order ledger quotation mechanism and constant product automatic market maker mechanism; S6. Node Price Generation: For AI inference computing networks, with the goal of maximizing social net surplus, a scheduling optimization problem with capacity and bandwidth constraints is solved, and the marginal price of nodes is obtained by iteratively solving the problem using the alternating direction multiplier method. The obtained marginal price of nodes is decomposed into system marginal cost, capacity congestion premium and bandwidth premium, and broadcast to network nodes.
[0014] Preferably, the generation of a uniform quality index The steps further include: Determine the modality type of the AI inference service to be evaluated, including text, image, video, voice, and multimodal. Based on the modality type, raw performance scores are collected from the corresponding third-party benchmark evaluation system on at least one capability dimension, including language understanding, code generation, mathematical reasoning, image fidelity, image-text matching, or video fidelity. The raw performance scores collected are standardized. The standardized scores are then synthesized into a comprehensive quality index with values ranging from [0,1] using principal component analysis or weighted aggregation methods. .
[0015] Preferably, the congestion factor This represents the real-time load level of the corresponding platform within the sampling time window, and is taken as the ratio of the platform's current active requests to its maximum concurrent processing capacity, with a minimum value of 1. As the platform load increases, the congestion factor increases, as indicated by the formula... This reduces the standardized AI service unit value of the platform, directing inference requests to nodes with lower loads.
[0016] Thirdly, an AI inference service system interface protocol device includes: Technical abstraction layer, containing a transformation operator It is used to map heterogeneous billing rules, platform congestion status, quality assessment results and service priority information of AI inference services of different modalities into a unified AI service unit value. The market incentive layer is used to execute capacity pre-sales in the form of AI service unit forward contracts, embed credit premium factors into cross-platform exchange rate calculations, and automatically allocate secondary market transfer royalties through smart contracts. The regulatory adaptation layer is used to classify AI service unit-related assets into utility and transaction categories, drive data residency routing based on data tags, and provide regulatory authorities with compliance audit proofs for cross-modal services through zero-knowledge proof technology.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention constructs a modality-adaptive quality assessment model and a unified measurement framework, incorporating AI inference services of different modalities such as text, images, videos, voice, and multimodal services into the same comparable price measurement system. It introduces an equivalent token conversion factor and a modality-related congestion factor, enabling the originally heterogeneous billing units to be uniformly converted into standardized AI service unit values. This effectively solves the technical problem that cross-modal and cross-platform service values cannot be directly compared and uniformly priced.
[0018] 2. This invention uses a hedonic regression model to estimate the parameters of the quality index, delay coefficient, credit premium, and congestion factor. It also synthesizes multi-dimensional third-party benchmark evaluation scores into a comprehensive quality index through principal component analysis or weighted aggregation methods. This achieves objective quantification and cross-platform standardization of AI service quality and reduces price information distortion caused by differences in service quality.
[0019] 3. This invention, by deploying a central price-limiting order ledger and an automated market maker mechanism in parallel, and combining market maker quotations based on an optimal control model with a constant product liquidity maintenance algorithm, forms a hybrid trading execution mechanism suitable for AI service unit-related assets. It can provide continuous liquidity while realizing price discovery, and generate cross-platform arbitrage signals through exchange rate monitoring, thereby promoting market prices to converge toward equilibrium levels.
[0020] 4. This invention employs the alternating direction multiplier method to solve the scheduling optimization problem with the goal of maximizing social net surplus, and directly decomposes the system marginal cost, capacity congestion premium, and bandwidth premium into three node price components from the converged dual variables. This achieves real-time decoupling and transparent broadcasting of the marginal price of computing nodes, providing millisecond-level scarce signal transmission capability for inference request routing in computing power networks.
[0021] 5. This invention adapts the risk control and clearing mechanisms in the traditional financial market to the AI computing power service trading scenario by introducing a dynamic risk value model to calculate the trading margin and combining it with smart contracts to execute T+0 settlement. This effectively reduces counterparty risk, improves fund settlement efficiency, and enhances the overall stability and operability of the market.
[0022] 6. This invention establishes an institutional interface protocol device as middleware, using a technical abstraction layer to map heterogeneous billing rules and platform congestion status into a unified service unit, a market incentive layer to realize capacity pre-sales, brand premium monetization, and automatic allocation of secondary royalties, and a regulatory adaptation layer to provide asset grading, data residency routing, and zero-knowledge compliance audit capabilities. In this way, a scalable interface is built between technical implementation and institutional compliance, laying a technical foundation for the institutionalized operation of the AI inference service market and collaborative governance among multiple stakeholders. Attached Figure Description
[0023] Figure 1 This is a diagram of the overall system architecture of the present invention; Figure 2 This is a diagram of the internal structure of the ASU computing module of the present invention; Figure 3 This is a flowchart of the multimodal quality index generation process of the present invention; Figure 4 This is a flowchart of the C-LMP node price generation process of the present invention; Figure 5This is a schematic diagram of the hybrid transaction execution mechanism of the present invention; Figure 6 This is a diagram of the three-layer protocol architecture of the IIP device of the present invention. Detailed Implementation
[0024] Please see Figures 1-6 This invention discloses a standardized pricing and trading system, method, and apparatus for AI inference services. Its core lies in constructing a quality assessment framework adapted to different modalities and a unified cross-modal pricing mechanism, enabling text, image, video, voice, and multimodal services to be incorporated into the same trading and clearing system. The following examples 1 to 5 will first demonstrate the specific applications of this unified pricing mechanism in text inference, image generation, video generation, speech recognition, and multimodal services. Examples 6 to 9 will then demonstrate the implementation methods of general infrastructure such as hybrid trading, node marginal pricing, derivatives clearing, and regulatory interface protocols. It should be understood that the modal types in Examples 1 to 5 are merely illustrative; the ASU measurement framework proposed in this invention can be extended to any type of AI inference service.
[0025] Before applying the ASU calculation formula described in this invention to AI inference services for arbitrary modalities, it is necessary to first analyze each elasticity parameter. The estimation process includes the following steps: This step applies to all modalities of AI inference services and is used to eliminate multicollinearity among variables before parameter estimation, ensuring the stability of the estimation results: In estimating the hedonic regression model Before determining the parameters, multicollinearity detection and processing of the design matrix are required. The column vector of the design matrix is... The response variable is .
[0026] First, calculate the variance inflation factor for each column vector. ,in For the first The column is used to calculate the coefficient of determination for linear regression on the remaining columns. When any... When this occurs, it indicates severe multicollinearity, and Gram-Schmidt orthogonalization is performed: ,in For the original column vector, These are the column vectors after orthogonalization. After orthogonalization This eliminates the impact of collinearity on the stability of parameter estimation. Then, weighted least squares estimation is performed using the orthogonalized design matrix, and the optimal regularization parameter is selected through 10-fold cross-validation. Finally, stable elasticity parameter estimates are obtained. ; The process of 10-fold cross-validation is as follows: Divide the sample data into 10 equal parts, and alternately use 9 parts as the training set and 1 part as the validation set in the candidate parameter set. The mean squared error of the validation set is calculated separately, and the set with the smallest mean squared error is selected. As the optimal regularization parameter.
[0027] Example 1: ASU Calculation for Text Inference Service This example demonstrates how to generate a comprehensive quality index for text modality AI inference services. And calculate the standardized AI service unit value .
[0028] Step 1: Determine the modality type and evaluation dimensions This embodiment serves a large language model API (Service A). Because it belongs to the text modality, the evaluation dimensions are determined to be language understanding ability and code generation ability. The MMLU benchmark is selected as the evaluation indicator for language understanding, and the HumanEval benchmark is selected as the evaluation indicator for code generation.
[0029] Step 2: Collect raw performance scores The raw score data for Service A was obtained from a publicly available third-party benchmarking system: MMLU accuracy was 0.86, and HumanEval pass rate was 0.67.
[0030] Step 3: Standardization Processing Since MMLU accuracy and HumanEval pass rate are both within the [0,1] interval, they can be directly used as standardized scores. .
[0031] Step 4: Weighted aggregation to generate a comprehensive quality index Weights are assigned to each dimension based on the business scenario. Let language comprehension ability have a weight of 0.7 and code generation ability have a weight of 0.3. Then, the weighted aggregation yields the overall quality index of service A: ; It lies within the interval [0,1].
[0032] Step 5: Congestion Factor The determination Congestion factor This represents the real-time load level of the corresponding platform within the sampling time window. It is calculated as the ratio of the platform's current active requests to its maximum concurrent processing capacity, with a minimum value of 1. Let's assume that the platform hosting service A has 1200 current active requests and a maximum concurrent processing capacity of 2000. This indicates that the platform is currently under normal load.
[0033] Step 6: Equivalent Token Conversion of Original Billing Price Service A's native billing price is $15.00 per million tokens. Since it is billed in tokens, no modal conversion is needed, making it comparable in price. (USD / million equivalent tokens).
[0034] Step 7: Parameter Setting and ASU Calculation Assuming that the elasticity parameters of the hedonic regression model have been estimated using historical data, including mass elasticity... Delay elasticity Credit premium sensitivity Congestion elasticity Simultaneously set the latency coefficient for service A. Log-scale credit spillover .
[0035] Substitute into the ASU calculation formula: ; Calculate the terms in the denominator: , , , ; The product of the denominators is approximately ; .
[0036] Therefore, the standardized AI service unit value of service A is obtained. 19.2 ASR units (the specific value depends on the parameter calibration; this example only demonstrates the calculation process).
[0037] Example 2: ASU Calculation for Image Generation Service This example demonstrates how to generate a comprehensive quality index for an AI inference service for image generation modality. And calculate the standardized AI service unit value .
[0038] Step 1: Determine the modality type and evaluation dimensions The service target of this embodiment is a text-to-image model API (Service B). Since it belongs to the image generation modality, the evaluation dimensions are determined to be image fidelity and image-text matching degree. FID (Fréchet Inception Distance) is selected as the evaluation index for image fidelity, and CLIP Score is selected as the evaluation index for image-text matching degree.
[0039] Step 2: Collect raw performance scores The raw score data for service B was obtained from a publicly available third-party benchmarking system: the FID score was 15 (the lower the score, the higher the fidelity), and the CLIP Score was 0.85 (the higher the score, the higher the matching degree between text and images).
[0040] Step 3: Standardization Processing To eliminate the dimensional differences between different indicators and unify their polarity, the original scores are subjected to min-max normalization and mapped to the [0,1] interval.
[0041] For the FID metric: Assume that the optimal (minimum) FID value is known from publicly available industry data to be 10, and the worst (maximum) FID value is 30. Since lower FID is better, the formula is used... Perform forward normalization: ; For the CLIP Score: Assume the best (maximum) value is 0.95 and the worst (minimum) value is 0.70. Since a higher CLIP Score is better, the formula is used. Normalize: .
[0042] Step 4: Weighted aggregation to generate a comprehensive quality index Weights are assigned to each dimension based on business needs. Let image fidelity have a weight of 0.6 and image-text matching degree have a weight of 0.4. The weighted aggregation yields the overall quality index of service B: , It lies within the interval [0,1].
[0043] Step 5: Congestion Factor The determination Congestion factor This characterizes the real-time load level of the corresponding platform within the sampling time window. Assuming that the platform hosting service B currently has 450 active requests and a maximum concurrent processing capacity of 300, then... This indicates that the platform is currently overloaded.
[0044] Step 6: Equivalent Token Conversion of Original Billing Price Service B's native billing price is $0.12 per 1024×1024 image. To incorporate it into the unified ASU pricing framework, a pre-defined modality conversion factor is used to convert "per image" into a comparable price based on equivalent token units. The modality conversion factor is determined using a market price ratio method: using the output price of GPT-4o at a base time of 15 / million tokens as a reference anchor, combined with the market price per unit output of the target image generation service at the same base time, the number of tokens equivalent to the output of that modality is calculated in reverse. The conversion factor is set as follows: 1 1024×1024 image is equivalent to 2,667 equivalent tokens.
[0045] The comparable price of service B after discounting (USD / million equivalent tokens) is: ; Modal conversion factors bridge the gap between the native billing units of different modal services and a unified equivalent token unit of measurement. The core method involves using the unit output price of a specified reference platform at a base time as an anchor point and employing a market price ratio method to inversely calculate the equivalent token quantity corresponding to each modal unit output. Each modal conversion factor is periodically calibrated and updated by the system governance team based on market data. Through this conversion, the prices of services with previously heterogeneous billing units, such as image, video, and voice services, are unified under the same unit of measurement, "USD / million equivalent tokens," allowing for comparison and standardization using a unified ASU calculation formula.
[0046] Step 7: Parameter Setting and ASU Calculation The same elastic parameters as in Example 1 are used: , , , Set the latency factor for service B. Log-scale credit premium .
[0047] Substitute into the ASU calculation formula: ; Calculating the denominator yields: .
[0048] Therefore, the standardized AI service unit value of service B is 94.5 ASR units (the specific value depends on the parameter calibration; this example only demonstrates the calculation process).
[0049] Step 8: Example of cross-modal exchange rate generation Based on the calculation results of Examples 1 and 2, a cross-modal exchange rate can be generated. The exchange rate between text service A and image service B. This exchange rate represents the ASU value of 1 unit of service A as equivalent to the ASU value of 0.203 units of service B on a quality-adjusted standardized scale. When the system detects a cross-platform or cross-modal exchange rate deviation from a preset equilibrium threshold, it generates an arbitrage signal to guide the market back to a reasonable exchange rate.
[0050] Example 3: ASU Calculation for Video Generation Service This example demonstrates how to generate a comprehensive quality index for AI inference services in video generation modality. And calculate the standardized AI service unit value .
[0051] Step 1: Determine the modality type and evaluation dimensions The service target of this embodiment is a certain text-generated video model API (service C). Since it belongs to the video generation modality, the evaluation dimensions are determined to be video fidelity and temporal consistency. FVD (Fréchet Video Distance) is selected as the evaluation metric for video fidelity, and CLIPSIM is selected as the evaluation metric for temporal consistency.
[0052] Step 2: Collect raw performance scores The raw score data for Service C was obtained from a publicly available third-party benchmarking system: FVD score was 120, and CLIPSIM score was 0.78.
[0053] Step 3: Standardization Processing For the FVD metric: Assume that the optimal (minimum) value of FVD is known from publicly available industry data to be 50, and the worst (maximum) value is 300. Since lower FVD is better, a positive normalization is used: ; For the CLIPSIM metric: Assume the optimal (maximum) value is 0.95 and the worst (minimum) value is 0.60. Since a higher CLIPSIM is better, normalization is used. ; Step 4: Weighted aggregation to generate a comprehensive quality index Assuming a video fidelity weight of 0.6 and a temporal consistency weight of 0.4, the weighted aggregation yields the overall quality index of service C: .
[0054] Step 5: Congestion Factor The determination Congestion factor This characterizes the real-time load level of the corresponding platform within the sampling time window. If the platform hosting service C currently has 200 active requests and a maximum concurrent processing capacity of 500, then... This indicates that the platform is currently under normal load.
[0055] Step 6: Equivalent Token Conversion of Original Billing Price The native billing price for Service C is $0.50 per video segment with the above specifications (5 seconds long, 30fps, 1024×1024 resolution). Using the GPT-4o output price of 15 / million Tokens as a reference anchor, a modal conversion factor is determined using a market price ratio method. The conversion factor is set as follows: 1 second of 720p video is equivalent to 2,667 equivalent Tokens. Therefore, a 5-second video with a 1024×1024 resolution (different from 720p resolution; the conversion factor is scaled according to the resolution area ratio: 1024×1024 / 1280×720) is considered equivalent to 1024×1024 / 1280×724. (1.14 times) is equivalent to 5 × 2,667 × 1.14 15,200 equivalent tokens.
[0056] The comparable price of service C after discounting (USD / million equivalent tokens) is: .
[0057] Step 7: Parameter Setting and ASU Calculation Use the same elasticity parameters: , , , Set the latency factor for service C. Credit premium .
[0058] Substitute into the ASU calculation formula: ; The standardized AI service unit value for service C can be obtained through a similar calculation process. 32.89 / 0.512 64.2; Therefore, the standardized AI service unit value of service C is 64.2 ASR units (the specific value depends on the parameter calibration; this example only demonstrates the calculation process).
[0059] Example 4: ASU Calculation for Speech Recognition Service This example demonstrates how to generate a comprehensive quality index for voice modality AI inference services. And calculate the standardized AI service unit value .
[0060] Step 1: Determine the modality type and evaluation dimensions The service target of this embodiment is a speech recognition model API (Service D). Since it belongs to the speech modality, the evaluation dimension is determined to be recognition accuracy. Word error rate is selected as the core evaluation indicator. Real-time performance and other indicators can also be used for comprehensive evaluation.
[0061] Step 2: Collect raw performance scores and standardize them The WER of service D on the standard test set, obtained from a third-party benchmarking system, is 8.5%. Assume the industry's known best WER is 3.0% and worst is 25.0%. Since lower WER is better, positive normalization is used. ; Since this embodiment only uses WER as a single dimension, the overall quality index... If multi-dimensional indicators are introduced, they will be processed using a weighted aggregation method.
[0062] Step 3: Congestion Factor The determination Congestion factor This characterizes the real-time load level of the corresponding platform within the sampling time window. Assume that the platform hosting service D currently has 600 active requests and a maximum concurrent processing capacity of 800. This indicates that the platform is currently under normal load.
[0063] Step 4: Equivalent Token Conversion of Original Billing Price Service D's native billing price is $2.00 per kilosecond. Using the GPT-4o output price of 15 / million tokens at the baseline time as a reference anchor, a modal conversion factor is determined using the market price ratio method. The conversion factor is set as follows: 1 second of audio is equivalent to 7 equivalent tokens.
[0064] The comparable price of service D after discounting (USD / million equivalent tokens) is: .
[0065] Step 5: ASU Calculation Set the latency factor for service D Credit premium The same elastic parameters as in Example 1 are used: , , , Substitute into the ASU formula to complete the standardized measurement. .
[0066] Example 5: ASU computation for multimodal services This example demonstrates how to generate a comprehensive quality index for multimodal AI inference services. And calculate the standardized AI service unit value .
[0067] Step 1: Determine the modality type and evaluation dimensions This embodiment serves a multimodal large-scale model API (Service E), which supports the joint understanding and generation of multiple input modalities such as text, images, and audio. The evaluation dimensions cover cross-modal understanding and cross-modal generation capabilities. MMBench is selected as the evaluation metric for cross-modal understanding capabilities, and SEED-Bench is selected as the evaluation metric for cross-modal generation capabilities.
[0068] Step 2: Collect raw performance scores and standardize them Service E, obtained from third-party benchmarking systems, achieved a composite score of 0.72 (out of 1.0) on MMBEnch and a score of 0.65 (out of 1.0) on SEED-Bench. Both metrics fall within the [0,1] range and can be used directly or fine-tuned and normalized. Let the standardized values be... .
[0069] Step 3: Weighted aggregation to generate a comprehensive quality index Let the weight of cross-modal understanding ability be 0.5, and the weight of cross-modal generation ability be 0.5, then: .
[0070] Step 4: Congestion Factor The determination Congestion factor This characterizes the real-time load level of the corresponding platform within the sampling time window. The congestion factor of a multimodal service is determined by the overall computing power load of the platform hosting the service and does not change with the combination of input modalities in a single request. If the platform hosting service E currently has 700 active requests and a maximum concurrent processing capacity of 500, then... This indicates that the platform is currently overloaded.
[0071] Step 5: Equivalent Token Conversion of Original Billing Price Service E adopts a comprehensive billing model. The system converts it into a uniform equivalent token comparable price using a modal conversion factor. The conversion method adopts the market price ratio method, using the output price of GPT-4o at the base time of 15 / million tokens as the reference anchor. The text, image and other modal parts of the request are converted according to the corresponding modal conversion factor (text: 1 token = 1 equivalent token; image: 1 1024×1024 image = 2,667 equivalent tokens; video: 1 second 720p = 2,667 equivalent tokens; audio: 1 second = 7 equivalent tokens), and then summed to obtain the total equivalent token number, and then the comparable price is calculated. Substitute these values into the ASU formula to complete the standardized measurement. In this example, the congestion factor... Substitute into the ASU formula to complete the standardized measurement. In this example, the congestion factor... = 1.4 will reduce the platform's ASU value relative to a normal load platform, reflecting the decline in effective service quality caused by overload.
[0072] Example 6: Hybrid Transaction Execution Module This example demonstrates a hybrid trading mechanism in which the Central Limit Order Book (CLOB) and Automated Market Makers (AMM) operate in parallel.
[0073] 1. CLOB Market Maker Quotes Market makers generate quotes based on the Avellaneda-Stoikov optimal control model. Let the current midpoint be... Market maker risk aversion coefficient Price volatility ,time left Standardized inventory Half price difference parameter Order arrival strength .
[0074] Substitute into the formula ,in : Calculate the first premium: .
[0075] Calculate the second inventory adjustment: .
[0076] Best Price .
[0077] 2. AMM Constant Product Liquidity Pool Suppose that the initial liquidity pool for ASR tokens and stablecoin USDC contains... and constant product The current ASR price is .
[0078] When the user buys The amount of USDC to be paid at that time The actual average transaction price was approximately This results in approximately 1% price slippage.
[0079] The system calculates impermanent loss every 10 seconds. Assume the current ASR market price. Then the price ratio Impermanent loss This means that the LP lost 20% of the value compared to simply holding the asset.
[0080] 3. Intelligent liquidity routing The intelligent routing submodule monitors the CLOB ledger depth in real time. When the size of the order submitted by the user is less than When the order size exceeds a certain threshold, the order is routed to CLOB to obtain a better price; when the order size exceeds a certain threshold... If the CLOB spread is too large, the system will automatically split or route all orders to the AMM pool for execution to ensure fast order execution and liquidity.
[0081] Example 7: Computational Node Marginal Pricing Module (C-LMP) This embodiment demonstrates how the C-LMP module generates dynamic node marginal prices for each node in an AI inference computing network.
[0082] 1. Network Modeling Consider a simplified network with three computational nodes: nodes 1 and 2 are located in region A, and node 3 is located in region B. The cost function for each node is a piecewise quadratic convex function. .
[0083] The node parameters are shown in the table below ( The unit is RPS, which is the number of inference requests per second.
[0084]
[0085] Current system total requirements Demand utility function ,in marginal utility .
[0086] 2. ADMM Iterative Solution Set penalty parameters (Step size parameter used to balance primal feasibility and dual convergence speed), initialization ,all .
[0087] Taking the first iteration as an example, update the original variables for node 1. : intermediate variables of proximal operators .initial ,but .
[0088] Here, for any convex cost function This means adding a parameter about the point to the original objective function. The minimum value is calculated after the quadratic penalty term, thereby balancing the single-step update magnitude and global convergence during the optimization process. This embodiment addresses the piecewise quadratic convex cost function. The proximal operator has an analytical solution in closed form: Substituting the analytical solution of the proximal operator :
[0089] ;
[0090] ; The same calculation applies to nodes 2 and 3; In each iteration, after updating the original variables of all nodes, the dual variables need to be updated according to the degree of constraint violation, as follows:
[0091]
[0092]
[0093] in For penalty parameters, The function ensures the dual variable and The non-negativity of the congestion premium or bandwidth premium means that when the node load does not exceed its capacity or bandwidth limit, the corresponding congestion premium or bandwidth premium is zero; the premium only takes a positive value when the constraint is violated, thus accurately reflecting the scarcity of the node.
[0094] A typical convergence process requires 20 to 50 iterations. After convergence, the values of the dual variables are read, assuming the following: .
[0095] 3. Node Price Decomposition Real-time price of node 1 .
[0096] Real-time price of node 2 .
[0097] Real-time price of node 3 .
[0098] The price difference reflects the congestion level and bandwidth limitations of different nodes. Node 3 has the highest price, indicating that its region is experiencing a shortage of computing power or a severe bandwidth bottleneck. The scheduler should prioritize routing new inference requests to nodes 1 and 2, which have lower prices.
[0099] (iv) Price broadcasting and EMA smoothing Node prices are updated and broadcast at intervals not exceeding 100 milliseconds. To suppress transient fluctuations, the system applies dynamic EMA smoothing. Let the standard deviation of the price over the last 10 steps for node 1 be... Historical average .because Take the smoothing coefficient That is, it does not smooth out price changes and responds quickly to them. If a certain period of time... ,but To effectively smooth out noise; Smoothed node broadcast price Updated according to the exponential moving average formula: .when hour, That is, directly using the original calculated price; when At that time, the current price only contributes 20% of the weight, while the historical smoothed value accounts for 80%, thus effectively suppressing abnormal jumps.
[0100] Example 8: Derivatives and Clearing Module This example demonstrates the dynamic margin calculation and T+0 settlement process.
[0101] 1. Dynamic margin calculation A user holds value The ASR long position. The system uses historical simulation to calculate VaR: it takes the daily return series of the past 250 trading days, sorts them from smallest to largest, and takes the 2.5 percentile (i.e., the 7th worst return). Assuming this return is -8%, then... .
[0102] Setting the value-at-risk multiplier Position ratio buffer coefficient .
[0103] Dynamic minimum margin .
[0104] 2. Adjustment during periods of extreme volatility If the system detects that the price volatility in the past 24 hours exceeds three standard deviations of the historical average, it will automatically... Upgraded to 3.0. At this point... .
[0105] 3. Risk monitoring and forced liquidation The system tracks the net value of the user's margin account in real time. When the loss reaches 50% of the initial margin (21,000 USDC), i.e., 10,500 USDC, a margin call notice is sent to the user. When the loss reaches 80%, i.e., 16,800 USDC, the smart contract automatically triggers forced liquidation, selling 50% of the position at the market price to reduce risk exposure.
[0106] 4. T+0 smart contract settlement All transactions are settled in real time via smart contracts deployed on the blockchain. Once a transaction is confirmed, ASR tokens and stablecoins are exchanged in atomic swaps, and ownership of funds and assets is transferred instantly upon completion of the transaction, without waiting for the traditional settlement cycle.
[0107] Example 9: System Interface Protocol Device (IIP Device) This embodiment demonstrates the specific implementation of the three-layer middleware of the IIP device.
[0108] 1. Technical Abstraction Layer Transformation Operator It receives API responses from different AI service platforms and parses their heterogeneous billing rules. For example, one platform bills based on the number of input tokens, the number of output tokens, cache hit rate, and geolocation surcharge, while another platform bills based on image resolution and the number of generation steps. Operator Based on a pre-defined mapping table and real-time exchange rates, all billing elements are uniformly converted into ASU values. The parameter version management unit stores historical mapping snapshots in timestamped JSON format, supporting rollback to previous versions due to platform rule changes.
[0109] 2. Market Incentive Layer The capacity pre-sale feature allows computing power providers to sell computing power service quotas for a future period in the form of ASR forward contracts. Buyers can lock in future costs, while providers can obtain cash flow in advance. The ASR is automatically burned and the service quota is disbursed upon contract expiration.
[0110] Brand premium monetization is achieved through credit rating. Embedded ASU exchange rate calculation implementation. The high-reliability platform achieves lower costs due to its low failure rate and high SLA compliance rate. This value allows them to obtain a premium in cross-platform exchange rates, incentivizing platforms to improve service quality.
[0111] Secondary royalty distribution is automatically executed by a smart contract: whenever ASR tokens are transferred, the contract automatically transfers 1% of the transaction amount to the wallet address of the original issuing platform, thus achieving continuous value capture.
[0112] 3. Regulatory Adaptation Layer The asset tiering module classifies ASRs into utility-type (limited to redeeming AI inference services, non-transferable) and transaction-type (freely tradable) based on their functional attributes. These two types of ASRs correspond to different compliance requirements and transaction rules.
[0113] Geographic compliance is achieved through the DataTag mechanism: each inference request is accompanied by a data source geographic tag. When redeeming an ASR, the system automatically selects a computing node located in a compliant jurisdiction to execute the task based on the tag, ensuring that the data residency requirement is met.
[0114] The compliance audit employs zk-SNARK zero-knowledge proof technology. Without disclosing the specific content of user requests or their identity information, the platform can generate verifiable compliance certificates for regulatory agencies, proving that it has fulfilled its regulatory obligations, such as content review and data localization, as required.
Claims
1. A standardized pricing and transaction system for AI inference services, characterized in that, include: The AI service unit computing module is used to collect raw performance scores from third-party benchmarking systems in the corresponding domains for AI inference services of different modalities, and generate a unified quality index through a quality assessment model adapted to the modality. Based on the original billing price of each service and the aforementioned quality index, combined with the mode-related congestion factor... Calculate standardized AI service unit ASU values to establish a cross-modal, cross-platform comparable price measurement benchmark; The hybrid trading execution module is used to process trading orders for the assets associated with the AI service unit in parallel based on the Central Limit Order Book (CLOB) and the Automated Market Maker (AMM) mechanism; wherein, the CLOB generates quotes based on the market maker's optimal control model, and the AMM mechanism maintains liquidity and calculates impermanent loss based on a constant product formula; The C-LMP module for computing node marginal pricing is used to solve scheduling optimization problems with capacity and bandwidth constraints for AI inference computing networks with the goal of maximizing social net surplus. It solves the problem iteratively using the Alternating Direction Multiplier Method (ADMM), and after convergence, it directly decomposes the system marginal cost, capacity congestion premium, and bandwidth premium from the dual variables to form the node price and broadcast it to the system. The Institutional Interface Protocol (IIP) module, as middleware, provides a technical abstraction layer, a market incentive layer, and a regulatory adaptation layer. It is used to map heterogeneous billing rules and congestion factors of different modalities into standardized AI service units, execute capacity pre-sales and secondary royalty allocation, and implement data residency and compliance audits. The derivatives and clearing module is used to calculate the trading margin based on the real-time price and position data output by the hybrid trading execution module, through a dynamic risk value model, and to execute T+0 settlement through a smart contract.
2. The standardized pricing and transaction system for AI inference services according to claim 1, characterized in that, The process by which the AI service unit calculation module generates the quality index includes: Determine the modality type of the AI inference service to be evaluated, and select at least one third-party benchmark evaluation indicator for the corresponding capability dimension. The capability dimension includes, but is not limited to, language understanding, code generation, mathematical reasoning, and image quality assessment. The modality type includes text, image, video, voice, and multimodal. The raw performance scores of the AI service model on the capability dimension are collected from the selected third-party benchmarking system; The raw performance scores collected are standardized to eliminate the dimensional differences between different indicators; The standardized scores are then synthesized into a comprehensive quality index with values ranging from [0,1] using principal component analysis or weighted aggregation methods.
3. The standardized pricing and transaction system for AI inference services according to claim 2, characterized in that, The third-party benchmark evaluation indicators include, but are not limited to: For text modalities, at least benchmark metrics in the fields of language understanding, code generation, or mathematical reasoning should be included; For image generation modalities, at least benchmark evaluation metrics in the fields of image fidelity or image-text matching should be included; For video generation modalities, at least benchmark evaluation metrics in the fields of video fidelity or temporal consistency should be included; For multimodal models, at least benchmark metrics in the areas of cross-modal understanding and generation capabilities should be included.
4. The standardized pricing and transaction system for AI inference services according to claim 1, characterized in that, The AI service unit calculation module calculates the standardized AI service unit value. The method is as follows: in, The original billing price for the AI inference service is uniformly converted into a comparable price based on equivalent token units through a preset modality conversion factor. The modality conversion factor is determined by using the unit output price of a specified reference platform at a base time as an anchor point and the market price ratio method to inversely calculate the equivalent token quantity corresponding to each modality unit output. The quality index is mentioned above. This refers to the mass elasticity parameter; The delay factor is... This is a delay elasticity parameter; The credit premium is measured on a logarithmic scale. This is a parameter related to credit premium sensitivity. The congestion factor reflects the platform's real-time load level. This refers to the congestion resilience parameter.
5. The standardized pricing and transaction system for AI inference services according to claim 1, characterized in that, The scheduling optimization problem solved by the marginal pricing module of the computing nodes is based on the inference request volume allocated to each computing node. The decision variables are: demand utility and total cost at each node. The objective function is to maximize the difference between demand utility and total cost at each node. Constraints include node capacity limits. and bandwidth limit ; The node price The dual variables obtained from solving this problem are directly derived, and their decomposition is: ,in For the system's marginal cost, For capacity congestion premium, This is a bandwidth premium.
6. The standardized pricing and transaction system for AI inference services according to claim 1, characterized in that, The derivatives and clearing module calculates the dynamic minimum margin. The formula is: in, Value at Risk (VaR) This is the position ratio buffer factor. The value of holdings calculated based on historical simulation methods within a specific time window The maximum possible loss at a 99% confidence level. This is the real-time median price.
7. A standardized pricing and transaction method for AI inference services, characterized in that, Includes the following steps: S1. Data Acquisition and Quality Index Generation: For AI inference services of different modalities, raw performance scores are collected from third-party benchmark evaluation systems in the corresponding fields, covering multiple dimensions including but not limited to language understanding, code generation, mathematical reasoning, and image quality assessment. The collected raw performance scores are standardized, and a comprehensive quality index with a value range of [0,1] is generated through principal component analysis or weighted aggregation methods. ; S2. Parameter estimation: Obtain the original billing price of each AI inference service and convert it into a comparable price based on equivalent token units through a preset modality conversion factor; Based on the hedonic regression model Quality index Delay coefficient Credit premium and congestion factor Estimate the elasticity parameters. For comparable prices, > 0 represents the market clearing constant. for ; S3. ASU Calculation: Using the estimated parameters, through the formula... Calculate the standardized AI service unit value for each service; S4. Exchange Rate Generation and Arbitrage Detection: Generate cross-platform exchange rates based on the AI service unit value. ,in and They represent the first The and the first The standardized service unit value of an AI inference service platform, and generates arbitrage signals when the exchange rate deviates from a preset threshold; S5. Transaction Execution: Receive transaction orders for the assets associated with the AI service unit and match transactions based on a hybrid mode of central limit order ledger quotation mechanism and constant product automatic market maker mechanism; S6. Node Price Generation: For AI inference computing networks, with the goal of maximizing social net surplus, a scheduling optimization problem with capacity and bandwidth constraints is solved, and the marginal price of nodes is obtained by iteratively solving the problem using the alternating direction multiplier method. The obtained marginal price of nodes is decomposed into system marginal cost, capacity congestion premium and bandwidth premium, and broadcast to network nodes.
8. The standardized pricing and transaction method for AI inference services according to claim 7, characterized in that, The generation of a unified quality index The steps further include: Determine the modality type of the AI inference service to be evaluated, including text, image, video, voice, and multimodal. Based on the modality type, raw performance scores are collected from the corresponding third-party benchmark evaluation system on at least one capability dimension, including language understanding, code generation, mathematical reasoning, image fidelity, image-text matching, or video fidelity. The raw performance scores collected are standardized. The standardized scores are then synthesized into a comprehensive quality index with values ranging from [0,1] using principal component analysis or weighted aggregation methods. .
9. The standardized pricing and transaction method for AI inference services according to claim 7, characterized in that, The congestion factor This represents the real-time load level of the corresponding platform within the sampling time window, and is taken as the ratio of the platform's current active requests to its maximum concurrent processing capacity, with a minimum value of 1. As the platform load increases, the congestion factor increases, as indicated by the formula... This reduces the standardized AI service unit value of the platform, directing inference requests to nodes with lower loads.
10. An AI inference service system interface protocol device, characterized in that, include: Technical abstraction layer, containing a transformation operator It is used to map heterogeneous billing rules, platform congestion status, quality assessment results and service priority information of AI inference services of different modalities into a unified AI service unit value. The market incentive layer is used to execute capacity pre-sales in the form of AI service unit forward contracts, embed credit premium factors into cross-platform exchange rate calculations, and automatically allocate secondary market transfer royalties through smart contracts. The regulatory adaptation layer is used to classify AI service unit-related assets into utility and transaction categories, drive data residency routing based on data tags, and provide regulatory authorities with compliance audit proofs for cross-modal services through zero-knowledge proof technology.