System and method for managing token usage amount
The system manages token usage by authorizing departments and tracking usage limits, preventing excessive costs by ensuring compliance with departmental approval levels and updating token usage amounts.
Patent Information
- Application Number
- JP2025003399
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Existing systems struggle to manage token usage across different departments within a company, leading to potential excessive costs due to varying approval levels and usage requirements for generative AI systems.
A system and method that stores user accounts, organization identification data, and token management data, determining if the user's organization is authorized and if the estimated token usage is within limits before transmitting requests to the generative AI system, updating token usage amounts, and sending warnings when nearing depletion.
Effectively manages token usage, reducing costs by preventing excessive usage and ensuring compliance with departmental approval levels through precise token management.
Smart Images

Figure 2025109191000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a management system, and more particularly to a system for managing token usage. The present invention further relates to a method for managing token usage.
Background Art
[0002] Large language models (LLMs) such as those disclosed in Patent Document 1 have been rapidly developing in recent years. The number of users of generative AI systems (e.g., ChatGPT (registered trademark)) built using LLMs has also been increasing rapidly. The pricing plan of generative AI systems is a prepayment system where, upon paying a certain amount, a certain number of tokens can be used. If the number of tokens exceeds a certain amount, additional charges will be billed. The number of tokens is calculated based on the number of characters.
[0003] When a company enters into a corporate contract for the use of a generative AI system, there may be a situation where employees belonging to different departments within the company use the generative AI system, and there may also be a situation where the approval for the use of the generative AI system differs for each department. Therefore, managing the token usage status of different departments within a company and controlling the usage cost of the system are the problems that the present invention aims to solve.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Therefore, an object of the present invention is to provide a system and method for managing token usage.
Means for Solving the Problems
[0006] According to one aspect of the present invention, a system for managing token usage is configured to communicate with a generative AI system and a user terminal electronic device, and includes a storage unit and a processing unit connected to the storage unit. The storage unit stores a user account, organization identification data corresponding to the user account, and token management data. The token management data includes a plurality of authorized organization identification data respectively corresponding to a plurality of authorized organizations, and a plurality of remaining token usage amounts respectively corresponding to the authorized organizations. When the processing unit receives transmission-waiting data including a prompt from the user terminal electronic device that has logged in to the system for managing token usage via the user account, it determines whether the organization identification data corresponding to the user account matches one of the authorized organization identification data. If it is determined that the organization identification data corresponding to the user account matches one of the authorized organization identification data, based on the prompt, an estimated token usage amount is obtained, and it is determined whether the estimated token usage amount is less than or equal to one of the remaining token usage amounts corresponding to one of the authorized organization identification data that matches the organization identification data. If it is determined that the estimated token usage amount is less than or equal to one of the remaining token usage amounts corresponding to the one of the authorized organization identification data that matches the organization identification data, a response request including the transmission-waiting data is transmitted to the generative AI system, a response result corresponding to the response request and the current token usage amount are received from the generative AI system, the response result is transmitted to the user terminal electronic device, and the remaining token usage amount corresponding to the one of the authorized organization identification data that matches the organization identification data is updated using the current token usage amount.
[0007] According to another aspect of the present invention, a method for managing token usage is executed by a system for managing token usage configured to communicate with a generative AI system and a user terminal electronic device. The system for managing token usage includes a storage unit and a processing unit connected to the storage unit. The storage unit stores a user account, organization identification data corresponding to the user account, and token management data. The token management data includes a plurality of authorized organization identification data respectively corresponding to a plurality of authorized organizations, and a plurality of remaining token usages respectively corresponding to the authorized organizations.
[0008] The method for managing token usage includes the steps of: when the processing unit receives transmission-waiting data including a prompt from a user terminal electronic device that has logged in to the system for managing token usage via the user account, determining whether the organization identification data corresponding to the user account matches one of the authorized organization identification data; when the processing unit determines that the organization identification data corresponding to the user account matches one of the authorized organization identification data, obtaining an estimated token usage based on the prompt; determining whether the estimated token usage is less than or equal to one of the remaining token usages corresponding to one of the authorized organization identification data that matches the organization identification data; when the processing unit determines that the estimated token usage is less than or equal to one of the remaining token usages corresponding to the one of the authorized organization identification data that matches the organization identification data, sending a response request including the transmission-waiting data to the generative AI system; receiving, by the processing unit, a response result corresponding to the response request and the current token usage from the generative AI system; and sending, by the processing unit, the response result to the user terminal electronic device and updating the one of the remaining token usages corresponding to the one of the authorized organization identification data that matches the organization identification data using the current token usage.
Advantages of the Invention
[0009] The memory unit stores the remaining token usage corresponding to an authorized organization. The processing unit obtains an estimated token usage based on a prompt and transmits a response request including data waiting to be sent to the AI system only when the estimated token usage is less than or equal to the remaining token usage corresponding thereto. Thereby, the token usage can be managed, and the cost due to excessive token usage can be reduced.
[0010] Other features and advantages of the present invention will become apparent in the following detailed description of embodiments with reference to the accompanying drawings.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2A
Figure 2B
Modes for Carrying Out the Invention
[0012] As used herein, the terms "coupled" or "connected" mean that between a plurality of electrical devices / apparatus / facilities are directly connected by a conductive material (e.g., an electric wire), or that between two electrical devices / apparatus / facilities are indirectly connected by one or more other devices / apparatus / facilities or wireless communication.
[0013] Figure 1 shows a system 100 for managing token usage according to an embodiment of the present invention. The system 100 for managing token usage is configured to communicate with a generative AI system 200 (e.g., ChatGPT) and a user-end electronic device 300, and includes a storage unit 1 and a processing unit 2. The user-end electronic device 300 is, for example, a smartphone, a tablet computer, a laptop computer, or a desktop computer.
[0014] The storage unit 1 stores a plurality of user accounts, organization identification data corresponding to each user account, and token management data. The token management data includes a plurality of authorized organization identification data corresponding to a plurality of authorized organizations respectively, and a plurality of remaining token usage amounts corresponding to the authorized organizations respectively. The storage unit 1 may be realized as a non-volatile storage device such as a hard disk drive, a flash memory, etc.
[0015] The processing unit 2 is connected to the storage unit 1. The processing unit 2 may include at least one of a single-core processor, a multi-core processor, a dual-core mobile processor, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), but is not limited thereto.
[0016] Figures 2A and 2B show a method for managing token usage according to an embodiment of the present invention. The method for managing token usage is executed by the system 100 for managing token usage. The method for managing token usage includes steps S01 to S09.
[0017] In step S01, when the processing unit 2 receives data waiting to be transmitted including a prompt from the user terminal electronic device 300 that has logged in to the token usage management system 100 via one of the user accounts stored in the storage unit 1 (hereinafter referred to as the "logged-in user account"), it determines whether one of the organization identification data corresponding to the logged-in user account (hereinafter referred to as the "organization identification data of the affiliated organization") matches one of the approved organization identification data. If it is determined that the organization identification data of the affiliated organization matches one of the approved organization identification data, the process proceeds to step S02; otherwise, the flow is completed. For example, a user who requests the use of the generative AI system 200 uses the user terminal electronic device 300 to log in to the token usage management system 100 via the user's account (the logged-in user account). The organization identification data of the affiliated organization indicates the organization to which the user belongs. By determining whether the organization identification data of the affiliated organization matches one of the approved organization identification data, it is determined whether the organization to which the user belongs is an organization approved for the use of the generative AI system 200. If the organization to which the user belongs is not an approved organization, the user's token usage is rejected.
[0018] If it is determined that the organization identification data of the affiliated organization matches one of the approved organization identification data, in step S02, the processing unit 2 determines whether one of the remaining token usage amounts corresponding to one of the approved organization identification data that matches the organization identification data of the affiliated organization (hereinafter referred to as the "corresponding remaining token usage amount") is 0. If it is determined that the corresponding remaining token usage amount is 0, the flow is completed, avoiding excessive token usage. Otherwise, the process proceeds to step S03.
[0019] If it is determined that the corresponding remaining token usage amount is not 0, in step S03, the processing unit 2 obtains an estimated token usage amount based on the prompt.
[0020] In step S04, the processing unit 2 determines whether the estimated token usage amount is less than or equal to the remaining token usage amount of the corresponding one. If it is determined that the estimated token usage amount is less than or equal to the remaining token usage amount of the corresponding one, the process proceeds to step S05; otherwise, the flow is completed to avoid excessive token usage.
[0021] If it is determined that the estimated token usage amount is less than or equal to the remaining token usage amount of the corresponding one, in step S05, the processing unit 2 sends a response request including the data waiting to be sent to the generation AI system 200.
[0022] In step S06, the processing unit 2 receives the response result corresponding to the response request and the current token usage amount from the generation AI system 200.
[0023] In step S07, the processing unit 2 sends the response result to the user terminal electronic device 300 and updates the remaining token usage amount of the corresponding one using the current token usage amount. Specifically, the remaining token usage amount of the corresponding one is updated by subtracting the current token usage amount from the remaining token usage amount of the corresponding one.
[0024] In step S08, the processing unit 2 determines whether the updated remaining token usage amount of the corresponding one is lower than a predetermined warning value. If it is determined that the updated remaining token usage amount of the corresponding one is lower than the predetermined warning value, the process proceeds to step S09; otherwise, the flow is completed.
[0025] If it is determined that the updated remaining token usage amount of the corresponding one is lower than the predetermined warning value, in step S09, the processing unit 2 sends a warning notification indicating that the remaining token usage amount of the corresponding one is lower than the predetermined warning value. The warning notification is sent, for example, in the form of an email to the organizational administrator (such as the user's supervisor).
[0026] In summary, the storage unit 1 of the system 100 for managing the token usage amount of the present invention stores the remaining token usage amount corresponding to the authorized organization, and the processing unit 2 obtains the estimated token usage amount based on the prompt, and generates a response request including the data waiting to be sent to the generation AI system 200 only when the estimated token usage amount is less than or equal to the remaining token usage amount corresponding thereto. Thereby, the token usage amount can be managed, the cost due to excessive token usage can be reduced, and the object of the present invention can be surely realized.
[0027] In the above description, for the purpose of explanation, numerous specific details have been set forth in order to provide a complete understanding of the embodiments. However, it will be apparent to those skilled in the art that one or more other embodiments may be practiced without these specific details. Also, in the description indicating "one embodiment" or "an embodiment" in this specification, it should be understood that all descriptions accompanied by indications such as ordinal numbers may be included in specific implementations of the present invention having specific aspects, structures, and features. Furthermore, in this specification, sometimes a plurality of variations are incorporated into one embodiment, drawing, or their descriptions, which is for rationalizing this specification and for the purpose of understanding the multifaceted nature of the present invention, and one or more features or specific examples in one embodiment may, where appropriate, be implemented together with one or more features or specific examples in other embodiments in the implementation of the present invention.
[0028] As described above, the embodiments and variations of the present invention have been explained, but the present invention is not limited to these, and shall include all modifications and equivalent configurations as various configurations included within the spirit and scope of the broadest interpretation.
Explanation of Reference Numerals
[0029] 100 System for managing token usage amount 1 Storage unit 2 Processing unit 200 Generation AI system 300 User-terminal electronic device Steps S01 to S09
Claims
Claim 1 A system for managing token usage configured to communicate with a generation AI system and a user-end electronic device, comprising a storage unit and a processing unit connected to the storage unit, wherein the storage unit stores a user account, organization identification data corresponding to the user account, and token management data, and the token management data includes a plurality of authorized organization identification data respectively corresponding to a plurality of authorized organizations and a plurality of remaining token usage amounts respectively corresponding to the authorized organizations, the processing unit, when receiving transmission-waiting data including a prompt from the user-end electronic device that has logged in to the system for managing the token usage via the user account, determines whether the organization identification data corresponding to the user account matches one of the authorized organization identification data, if it is determined that the organization identification data corresponding to the user account matches one of the authorized organization identification data, obtains an estimated token usage amount based on the prompt, determines whether the estimated token usage amount is less than or equal to one of the remaining token usage amounts corresponding to one of the authorized organization identification data that matches the organization identification data, if it is determined that the estimated token usage amount is less than or equal to one of the remaining token usage amounts corresponding to one of the authorized organization identification data that matches the organization identification data, transmits a response request including the transmission-waiting data to the generation AI system, receives a response result corresponding to the response request and the current token usage amount from the generation AI system, transmits the response result to the user-end electronic device, and is configured to update the one of the remaining token usage amounts corresponding to one of the authorized organization identification data that matches the organization identification data using the current token usage amount, a system for managing token usage. Claim 2 When the processing unit determines that the organization identification data corresponding to the user account matches one of the approved organization identification data, it determines whether the remaining token usage amount corresponding to the one of the approved organization identification data that matches the organization identification data is 0. When it determines that the remaining token usage amount corresponding to the one of the approved organization identification data that matches the organization identification data is not 0, it is configured to obtain the estimated token usage amount based on the prompt. The system for managing token usage according to claim 1.
3. The processing unit determines whether the updated remaining token usage amount is lower than a predetermined warning value. When it determines that the updated remaining token usage amount is lower than the predetermined warning value, it is configured to send a warning notification indicating that the remaining token usage amount is lower than the predetermined warning value. The system for managing token usage according to claim 1.
4. A method for managing token usage executed by a system for managing token usage configured to communicate with a generative AI system and a user terminal electronic device, The system for managing token usage includes a storage unit and a processing unit connected to the storage unit. The storage unit stores a user account, organization identification data corresponding to the user account, and token management data. The token management data includes a plurality of approved organization identification data respectively corresponding to a plurality of approved organizations and a plurality of remaining token usage amounts respectively corresponding to the approved organizations. The method for managing token usage is as follows. When the processing unit receives transmission-waiting data including a prompt from the user terminal electronic device that has logged in to the system for managing token usage via the user account, it determines whether the organization identification data corresponding to the user account matches one of the approved organization identification data. When the processing unit determines that the organization identification data corresponding to the user account matches one of the approved organization identification data, it obtains an estimated token usage amount based on the prompt. The step in which the processing unit determines whether the estimated token usage amount is less than or equal to one of the remaining token usage amounts corresponding to one of the approved organization identification data that matches the organization identification data; When the processing unit determines that the estimated token usage amount is less than or equal to one of the remaining token usage amounts corresponding to the one of the approved organization identification data that matches the organization identification data, the step of transmitting a response request including the data waiting to be sent to the generating AI system; The step in which the processing unit receives a response result corresponding to the response request and the current token usage amount from the generating AI system; The step in which the processing unit transmits the response result to the user terminal electronic device and updates the one of the remaining token usage amounts corresponding to the one of the approved organization identification data that matches the organization identification data using the current token usage amount, including: A method for managing token usage amounts.
5. The step of obtaining the estimated token usage amount is When the processing unit determines that the organization identification data corresponding to the user account matches one of the approved organization identification data, determining whether the one of the remaining token usage amounts corresponding to the one of the approved organization identification data that matches the organization identification data is 0; When the processing unit determines that the one of the remaining token usage amounts corresponding to the one of the approved organization identification data that matches the organization identification data is not 0, obtaining the estimated token usage amount based on the prompt, including: The method for managing token usage amounts according to claim 4.
6. The step in which the processing unit determines whether the updated one of the remaining token usage amounts is lower than a predetermined warning value; When the processing unit determines that the updated one of the remaining token usage amounts is lower than the predetermined warning value, the step of transmitting a warning notification indicating that the one of the remaining token usage amounts is lower than the predetermined warning value, further including: The method for managing token usage amounts according to claim 4.
Citation Information
Patent Citations
Authority administrative server and authority administrative method
JP2015130073A
Using sample question embeddings to choose between an LLM interfacing model and a non-LLM interfacing model
US20240362213A1
Large language models in machine translation
US8332207B2
Cited By
Program, method, information processing device, and system
JP7792666B1
APPARATUS AND METHOD FOR Predicting ARTIFICIAL INTELLIGENCE Token Consumption and controlling Dynamic Resource Allocation Using Image Object Structure Analysis IN ARTIFICIAL INTELLIGENCE-BASED LEARNING ASSISTANCE SYSTEM
KR102965568B1