Method and System for Using AI to Invoke Prediction and Caching

By introducing a caching mechanism and predicting request methods in AI computing services, the system burden of AI computing services when processing requests during peak periods and waste of resources during off-peak periods is solved, and more efficient resource utilization and processing capabilities are achieved.

CN114026837BActive Publication Date: 2025-05-27VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980097937.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-07-05
Publication Date
2025-05-27
Estimated Expiration
2039-07-05

AI Technical Summary

Technical Problem

When AI computing services handle a large number of prediction requests during peak periods, it leads to increased system burden and waiting time, and is seriously wasted resources during off-peak periods, making it difficult to optimize to support high demand.

Method used

By introducing a caching mechanism, future requests to the AI ​​engine are predicted and these requests are created in advance to cache the results. Calling the predictive processing computer monitors the processor usage of the AI ​​computer, and when the usage is below the threshold, a request from the orchestration service computer can cache the request and send it to the AI ​​computer to generate and cache the output.

Benefits of technology

It effectively disperses the time to generate prediction results, reduces the pressure on AI computers during peak periods, improves the efficiency of processors, and reduces waste of resources during off-peak periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114026837B_ABST
    Figure CN114026837B_ABST
Patent Text Reader

Abstract

Methods and systems for predicting and processing cacheable AI calls are disclosed. One method includes determining by a call prediction processing computer that an AI computer is operating below a threshold processor utilization. Then, the method includes requesting, by the call prediction processing computer, a set of cacheable requests from an orchestration service computer and receiving, from the orchestration service computer, the set of cacheable requests. Then, the method includes sending, by the call prediction processing computer, a cacheable request from the set of cacheable requests to the AI computer. Then, the method includes receiving, by the call prediction processing computer, an output. The output is generated by the AI computer based on the cacheable request. Then, the call prediction processing computer stores the output in a data repository.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] None. Background Art

[0003] With the progress of artificial intelligence (AI) technologies such as machine learning, natural language processing, and knowledge representation, they have become increasingly important. AI intelligence can be used to analyze a wide variety of topics such as weather forecasting, business decisions, and astronomical theories. However, the running cost of AI models can be very high and requires a large amount of processing power, which may mean expensive servers, space, and energy. An individual may not have the resources to run complex machine - learning models and may thus rely on services with supercomputers or large computing clusters that can process data and run models for them. Thus, AI computing services may want to optimize their services. Specifically, they want to optimize to avoid wasting precious processing time and ensure that the service can scale to handle queries from a large number of clients. During peak periods, a large number of prediction requests may be received, and each prediction request may require one or more models to run for a long time. This can lead to system overload and long waiting times. Conversely, during off - peak periods, few requests may occur, leaving the system with a large amount of idle time. Increasing processing power during peak periods to support high demand is expensive and may increase the idle processing power during off - peak periods.

[0004] In previous systems, a user could send a request for a prediction to a computer. The computer could retrieve context data related to the prediction request and then send the prediction request to an AI computer. The AI computer could process the prediction request and generate an output. The output was then returned to the user device through the computer. Thus, the work of the AI computer was limited to the requests received from the user device.

[0005] The embodiments, individually and collectively, solve these and other problems. Summary of the Invention

[0006] One embodiment includes determining by a call prediction processing computer that an AI computer is operating below a threshold processor utilization rate. Then, the method includes requesting, by the call prediction processing computer, a set of cacheable requests from an orchestration service computer and receiving, by the call prediction processing computer, the set of cacheable requests from the orchestration service computer. Then, the method includes sending, by the call prediction processing computer, a cacheable request from the set of cacheable requests to the AI computer. Then, the method includes receiving, by the call prediction processing computer, an output, where the output is generated by the AI computer based on the cacheable request; and initiating, by the call prediction processing computer, storing the output in a data repository.

[0007] Another embodiment includes a system that includes a call prediction processing computer that includes a processor and a computer-readable medium that includes code for implementing a method. The method includes determining that an AI computer is operating below a threshold processor utilization. The method then includes requesting a set of cacheable requests from an orchestration service computer and receiving the set of cacheable requests from the orchestration service computer. The method then includes sending cacheable requests from the set of cacheable requests to the AI computer. The method then includes receiving an output, where the output is generated by the AI computer based on the cacheable requests, and initiating storing the output in a data repository.

[0008] Another embodiment includes receiving, by an orchestration service computer, a prediction request message from a user device that includes a prediction request in the form of an AI call, and determining, by the orchestration service computer, that the prediction request is a cacheable request. The method then includes retrieving, by the orchestration service computer, context data from a data repository and modifying the prediction request message to include the context data. The method further includes sending, by the orchestration service computer, the prediction request message to an AI computer, the prediction request message including the AI call and the context data. The method then includes receiving, by the orchestration service computer, an output in response to the prediction request message from the AI computer and sending the output to the user device. The method then includes initiating storing the prediction request and the output in the data repository by the orchestration service computer.

[0009] Another embodiment includes a system that includes an orchestration service computer that includes a processor and a computer-readable medium that includes code for implementing a method. The method includes receiving a prediction request message from a user device that includes a prediction request in the form of an AI call and determining that the prediction request is a cacheable request. The method then includes retrieving context data from a data repository and modifying the prediction request message to include the context data. The method then includes sending the prediction request message to an AI computer, the prediction request message including the AI call and the context data. The method then includes receiving an output in response to the prediction request message from the AI computer and sending the output to the user device. The method then includes initiating storing the prediction request and the output in the data repository.

[0010] For more details regarding embodiments of the present invention, see the detailed description and the drawings. Description of the Drawings

[0011] Figure 1 A block diagram showing a system according to an embodiment.

[0012] Figure 2 A block diagram showing an orchestration service computer according to an embodiment.

[0013] Figure 3 A block diagram showing a call prediction processing computer according to an embodiment.

[0014] Figure 4 A flowchart showing a prediction request process according to an embodiment.

[0015] Figure 5 A flowchart showing a caching process according to an embodiment.

[0016] Figure 6 A flowchart showing a call prediction process according to an embodiment. Detailed Description of the Embodiments

[0017] Embodiments can improve the efficiency of AI engine services by introducing caching. Caching results can improve the efficiency of artificial intelligence calculations, which can be expensive in terms of computing and financial resources. Additionally, embodiments introduce a system for predicting future requests to an AI engine, which can play a role in predicting future requests to the AI engine, creating those requests in advance, and then caching the results.

[0018] Cacheable requests can be requests that are repeatable and may give the same result if asked again within a reasonable time frame. Cacheable requests are requests where the context data is independent of time. An example of a cacheable request might be "Where is the best location to open a new pizza store in San Francisco?". In this example, "San Francisco" is an example of geographical input data and is stable over time. Thus, cacheable requests asked at two different times (e.g., a week apart) may give the same result.

[0019] Conversely, non-cacheable requests can be time-dependent requests and / or may depend on time-dependent data. For example, a non-cacheable request might be related to transaction fraud risk such as "What is the fraud score for this transaction?". Each transaction fraud risk determination may depend on the entities involved, transaction details, transaction time, etc. Thus, caching the transaction fraud risk response for a transaction fraud risk request may no longer be relevant after the transaction has occurred, as the transaction will only happen once.

[0020] Before discussing embodiments of the present invention, some terms can be described in further detail.

[0021] A "user device" can be any suitable electronic device that can process information and communicate the information to other electronic devices. The user device may include a processor and a computer-readable medium coupled to the processor, the computer-readable medium including code executable by the processor. The user devices may also each include an external communication interface for communicating with each other and other entities. Examples of user devices may include mobile devices, laptop or desktop computers, wearable devices, etc.

[0022] A "user" can include an individual or a computing device. In some embodiments, a user may be associated with one or more user devices and / or personal accounts. In some embodiments, a user may be a cardholder, account holder, or consumer.

[0023] A "processor" can include any suitable one or more data computing devices. The processor may include one or more microprocessors working together to achieve the desired function. The processor may include a CPU, the CPU including at least one high-speed data processor sufficient to execute program components for executing user- and / or system-generated requests. The CPU can be a microprocessor such as AMD's Athlon, Duron, and / or Opteron; IBM and / or Motorola's PowerPC; IBM and Sony's Cell processor; Intel's Celeron, Itanium, Pentium, Xeon, and / or XScale; and / or one or more similar processors.

[0024] A "memory" can be any suitable one or more devices that can store electronic data. Suitable memories may include non-transitory computer-readable media that store instructions executable by the processor to implement the desired method. Examples of memories may include one or more memory chips, disk drives, etc. Such memories may operate using any suitable electrical, optical, and / or magnetic operating modes.

[0025] A "prediction request" can be a request for a predicted answer to a question. For example, a prediction request can be a request for some information about future events, classification predictions, optimizations, etc. The prediction request may be in the form of natural language (e.g., as an English sentence) or may be in a computer-readable format (e.g., as a vector). The answer to the prediction request may or may not be determined using artificial intelligence.

[0026] A "prediction request message" can be a message with a prediction request sent to an entity (e.g., an orchestration service computer, an AI computer with an AI engine).

[0027] "AI call" can be the execution of a subroutine, specifically, the execution of a command of an artificial intelligence model. An AI call can be in the form of natural language (e.g., as an English sentence), or can be in a computer-readable format (e.g., as a vector). An AI call can specifically be a command to obtain a result from a certain form of AI engine (machine learning, natural language processing, knowledge representation, etc.).

[0028] "Call information message" can be a message having information about an AI call and / or a prediction request. A call information message can include a prediction request (or AI call) and an output generated based on the prediction request (or AI call). A call information message can also include information such as context data used to generate the output, an identifier of the context data, and a timestamp of the prediction request.

[0029] "Threshold processor usage" can be a threshold of the usage amount of computer processing power. For example, the threshold can be a percentage of the processor usage, such as 20%. As another example, the threshold processor usage can be based on multiple hardware units, such as multiple processor cores in use. In some embodiments, the threshold can be determined such that there is sufficient processor operation ability to execute call prediction processing while still processing incoming prediction requests without reducing the speed or performance of the processor.

[0030] "Cacheable request" can be a query that can be stored. A cacheable request can be a request that is static over time. This can be static because the input (and output) is independent of time, and thus the output can be saved for a period of time. A cacheable request can also be a prediction request. If the same prediction request is received within the same time period, the original output is still valid. The request can be static over a time period such as a week or a month (or longer or shorter than these times).

[0031] "AI engine" can include multiple modules with artificial intelligence capabilities. Each module can perform tasks for one or more different fields of artificial intelligence such as machine learning, natural language processing, and knowledge representation. The AI engine is capable of integrating the models to, for example, use a natural language processing module to process a question and then use a machine learning module to generate a prediction in response to the question.

[0032] "Context data" can include data providing the circumstances surrounding an event, entity, or item. Context data can be additional information that provides a broader understanding or more details about existing information or a request (e.g., a prediction request, an AI call). In an embodiment, context data can include transaction history, geographical data, census data, etc.

[0033] Figure 1A block diagram showing system 100 according to an embodiment. System 100 may include a user device 110, an orchestration service computer 120, an AI computer 130, a messaging system computer 140, an AI call information consumer 150, a data repository 160, and a call prediction processing computer 170. Figure 1 Any of the devices may communicate with each other via a suitable communication network.

[0034] The communication network may include any suitable communication medium. The communication network may be one and / or a combination of the following: direct interconnection; the Internet; a local area network (LAN); a metropolitan area network (MAN); operating tasks as nodes on the Internet (OMNI); a secure custom connection; a wide area network (WAN); a wireless network (e.g., employing protocols such as, but not limited to, the Wireless Application Protocol (WAP), i-mode, etc.). Figure 1 Messages between the entities, providers, networks, and devices shown may be sent using a secure communication protocol, such as but not limited to: File Transfer Protocol (FTP); Hypertext Transfer Protocol (HTTP); Secure Hypertext Transfer Protocol (HTTPS), Secure Sockets Layer (SSL), ISO (e.g., ISO 8583), etc.

[0035] The user device 110 may be a device that a user can use to request a prediction from the AI engine of the AI computer 130 via the orchestration service computer 120. Examples of the user device 110 may include a laptop computer, a desktop computer, and a mobile device.

[0036] The orchestration service computer 120 may process prediction requests. The orchestration service computer 120 may act as a router to route messages to the appropriate computer. The orchestration service computer 120 may also determine and / or retrieve context data associated with the prediction request. In some embodiments, the orchestration service computer 120 may also determine a prediction for future cacheable prediction requests.

[0037] The AI computer 130 may be a computer having an AI engine. The AI computer 130 may have multiple models that can be used to complete different tasks. For example, the AI computer 130 may have a fraud detection model, a business location model, and a natural language processing model. The AI engine of the AI computer 130 may use the models individually and / or in combination.

[0038] The messaging system computer 140 may include message-oriented middleware. For example, the messaging system computer 140 may operate with a RabbitMQ message queue or an Apache Kafka message queue. The messaging system computer 140 may be a server computer. The messaging system computer 140 may receive messages from entities called producers, which may include the orchestration service computer 120 and / or the call prediction processing computer 170. A producer may be an entity that generates messages having information that other entities may use later. The messaging system computer 140 may receive call information messages from the orchestration service computer 120. The call information messages may include cacheable requests, outputs based on the cacheable requests, and context data. The messaging system computer 140 may additionally receive other information, such as context data, from other producers. The messaging system computer 140 may process the information and arrange the call information messages into one or more queues or partitions.

[0039] The AI call information consumer 150 may read call information messages from the messaging system computer 140. After reading the call information messages, the AI call information consumer 150 may store the information in the call information messages in the data repository 160.

[0040] The data repository 160 may store information including cacheable requests, outputs generated from the cacheable requests, and context data. The data repository 160 may retain the information for a period of time (e.g., one day, one week, one month). The period of time for caching the output may be part of the output. The data repository 160 may also store context data such as transaction data, business data, geographical data, etc. The data repository 160 may be a non-relational database (NoSQL database).

[0041] The call prediction processing computer 170 may process predictions regarding future prediction requests (e.g., AI calls) that users will send to the AI computer 130. The call prediction processing computer may also send the prediction requests to the AI computer 130 and / or the orchestration service computer 120. The call prediction processing computer 170 may communicate with the AI computer 130 to monitor the processor usage rate.

[0042] Figure 2A block diagram showing an orchestration service computer 120 according to an embodiment. The orchestration service computer 120 may include a memory 122, a processor 124, a network interface 126, and a computer-readable medium 128. The computer-readable medium 128 may store code that can be executed by the processor 124 to implement some of all the functions of the orchestration service computer 120. The computer-readable medium 128 may include a data retrieval module 128A, a prediction request module 128B, a cache module 128C, and an invocation prediction module 128D.

[0043] The memory 122 may be implemented using any combination of any number of non-volatile memories (e.g., flash memory) and volatile memories (e.g., DRAM, SRAM) or any other non-transitory storage medium or combination of media.

[0044] The processor 124 may be implemented as one or more integrated circuits (e.g., one or more single-core or multi-core microprocessors and / or microcontrollers). The processor 124 may be used to control the operation of the orchestration service computer 120. The processor 124 may execute various programs in response to program code or computer-readable code stored in the memory 122. The processor 124 may include the function of maintaining multiple simultaneously executing programs or processes.

[0045] The network interface 126 may be configured to connect to one or more communication networks to allow the orchestration service computer 120 to communicate with other entities such as, for example, a user device 110, an AI computer 130, etc. For example, the communication with the AI computer 130 may be direct, indirect, and / or via an API.

[0046] The computer-readable medium 128 may include one or more non-transitory media for storage and / or transmission. Suitable media include, for example, random access memory (RAM), read-only memory (ROM), magnetic media such as a hard disk drive, or optical media such as a compact disc (CD) or a digital versatile disc (DVD), flash memory, etc. The computer-readable medium 128 may be any combination of these storage or transmission devices.

[0047] The data retrieval module 128A in combination with the processor 124 may retrieve context data from a data repository 160. The retrieved context data may be related to a prediction request. For example, for a prediction request regarding the best location of a new hotel in San Francisco, the data retrieval module 128A in combination with the processor 124 may retrieve context data such as tourism in San Francisco, existing hotel locations, meal and travel transaction history, etc.

[0048] The prediction request module 128B in combination with the processor 124 can process prediction requests. Processing a prediction request can include determining whether the prediction request is cacheable. If it is cacheable and already in the cache, the cache module 128C in combination with the processor 124 can retrieve the cached output of the prediction request. The prediction request module 128B can also, in combination with the processor 124, determine what context data the prediction request requires, and then work with the data retrieval module 128A to retrieve the data from the data repository. Alternatively, the prediction request module 128B can receive the context data only from the data retrieval module 128A without first determining the relevant context data (i.e., the data can be determined by the data retrieval module). The prediction request module 128B can also send the prediction request and the relevant context data.

[0049] The cache module 128C in combination with the processor 124 can cache and retrieve cacheable requests. The cache module 128C can initiate storing the cacheable request in the data repository. The cache module 128C can initiate storing the cacheable request by sending the prediction request, the output, and the context data to the messaging system computer. Additionally, the cache module 128C can retrieve the cached result from the data repository.

[0050] The invocation prediction module 128D in combination with the processor 124 can predict what cacheable requests might be made. The invocation prediction module 128D can store past prediction requests and the timestamps when the cacheable requests were made. The invocation prediction module 128D can use frequencies to determine when a prediction request can be made. For example, the invocation prediction module 128D can determine a specific cacheable request made by different users each week. The invocation prediction module 128D can also predict cacheable requests based on the date. For example, a cacheable request such as "Where is the best place to open a Halloween costume store?" might be received each fall. Alternatively, the invocation prediction module 128D in combination with the processor 124 can send requests for a set of predicted cacheable requests to an AI computer. The cacheable requests predicted by the invocation prediction module 128D can be used by the invocation prediction processing computer to generate a prediction request before the cacheable request is received from the user device by the orchestration service computer 120.

[0051] Figure 3 Depict an invocation prediction processing computer 170 according to an embodiment. The invocation prediction processing computer 170 can include a memory 172, a processor 174, a network interface 176, and a computer-readable medium 178. These components can be coupled with Figure 2The corresponding components in the orchestration service computer 120 are similar or different. The computer-readable medium 178 may store code that can be executed by the processor 174 to implement some or all of the functions of the call prediction processing computer 170 described herein. The computer-readable medium 178 may include a performance monitoring module 178A, a call prediction request module 178B, a prediction request module 178C, and a cache module 178D.

[0052] The performance monitoring module 178A, in conjunction with the processor 174, may monitor the processor utilization rate (e.g., CPU utilization) of the AI computer 130. The performance monitoring module 178A may determine the amount of processing power being used by the AI computer 130 (e.g., the number of processor cores, the amount of non-idle time), and may determine whether the processor utilization rate is below a predetermined threshold. Example thresholds may be percentages, such as 20%, 30%, etc. The threshold may be selected such that there is sufficient processor operating capacity to perform the call prediction processing without degrading the processor's speed or performance in handling incoming prediction requests.

[0053] In some embodiments, the performance monitoring module 178A may directly access the hardware of the AI computer 130. Alternatively, the performance monitoring module 178A may receive information about the processor utilization rate from the AI computer 130 (e.g., from the performance monitoring software of the AI computer 130). The performance monitoring module 178A may operate on a schedule. For example, the performance monitoring module 178A may check the processor utilization rate every 5 minutes. If the processor utilization rate is above the threshold, the call prediction processing computer 170 may wait until the next scheduled time. If the processor utilization rate is below the threshold, the performance monitoring module 178A, in conjunction with the processor 174, may cause the call prediction processing computer 170 to activate the call prediction request module 178B.

[0054] The call prediction request module 178B, in conjunction with the processor 174, may send a request to the orchestration service computer 120 for predictably cacheable requests. The call prediction request module 178B may receive a set of cacheable calls from the orchestration service computer 120. The call prediction request module 178B may specify multiple cacheable calls for the request. For example, the call prediction request module 178B may request the N most likely to be requested requests. For example, if it is December, the cacheable calls may include "Where can I find a gift for my wife?", "How late is the Acme store open on Christmas Eve?", and "What is the best meal to cook for Christmas?". The call prediction request module 178B may be triggered by the performance monitoring module 178A determining that the AI computer 130 is operating below the threshold processor utilization rate.

[0055] The prediction request module 178C in combination with the processor 174 may send a prediction request, specifically a cacheable request, to the orchestration service computer. The prediction request module 178C may select a specific cacheable request from a set of cacheable requests received from the orchestration service computer 120. For example, the prediction request module 178C may prioritize the set of cacheable requests. Alternatively, the prediction request module 178C may select a random cacheable call from the set of cacheable requests.

[0056] The cache module 178D in combination with the processor 174 may initiate storing the output of the cacheable request in the data repository 160. The cache module 178D in combination with the processor 174 may store the cacheable request and the output in response to the cacheable request in the data repository 160. Alternatively, the cache module 178D in combination with the processor 174 may send a call information message including the cacheable request and the output to the message system computer.

[0057] Figure 4 A flowchart showing the processing of a prediction request according to an embodiment.

[0058] In step S502, the user device 110 may send a prediction request message to the orchestration service computer 120. The prediction request message may include a prediction request, which may be in the form of a question. The prediction request may be in the form of an AI call. The AI call may be an input to the AI engine. An example prediction request may be "Where is the best place to build a hotel in Austin, Texas?" Another example of a prediction request may be "Is this transaction fraudulent?" The prediction request message may be sent to the orchestration service computer 120 via a network service, API, etc. For example, there may be a network interface that allows the user device 110 to send the prediction request message to the orchestration service computer 120.

[0059] In step S504, the orchestration service computer 120 may determine that the prediction request is a cacheable request. For example, the question "Where is the best place to build a hotel in Austin, Texas?" may be a cacheable request, while the question "Is this transaction fraudulent?" may be a non-cacheable request. If the output has been previously cached and the prediction request is a cacheable request, the orchestration service computer 120 may attempt to retrieve the output of the prediction request from the data repository 160. If the output is in the data repository 160, the orchestration service computer 120 may return the output to the user device 110. A cacheable request may be a request that is relatively independent of time.

[0060] If the output of a prediction request is not stored in the data repository 160, the orchestration service computer 120 may evaluate the prediction request and / or the context of the user device 110 to determine whether the prediction request is a cacheable request. For example, if the user device 110 is a computer of a processing network (e.g., a payment processing network), the orchestration service computer 120 may infer that the prediction request is a transaction fraud request rather than a cacheable request. Alternatively, if the user device 110 is a personal laptop computer, the orchestration service computer 120 may infer that the prediction request is a cacheable request, or the probability that the prediction request is a cacheable request is higher. In some embodiments, the orchestration service computer 120 may use the AI computer 130, e.g., using a machine learning model of the AI computer 130, to determine whether the prediction request is cacheable.

[0061] In step S506, the orchestration service computer 120 may retrieve context data related to the prediction request from the data repository 160. For example, if the prediction request is "Where is the best place to build a hotel in Austin, Texas?", the orchestration service computer 120 may retrieve information such as geographical information about Austin, the locations of existing hotels and tourist attractions in Austin, and the transaction patterns in Austin. In some embodiments, the prediction request message may include context data.

[0062] In step S508, the orchestration service computer 120 may send a prediction request message including an AI call (i.e., the prediction request) and the context data to the AI computer 130. The AI computer 130 may use an AI engine to process the AI call with an AI model or a combination of models and generate an output in response to the prediction request message. Examples of AI models may include machine learning models, neural networks, recurrent neural networks, support vector machines, Bayesian networks, genetic algorithms, etc. For example, the AI engine may analyze the context data (e.g., existing hotel locations, hospitality transaction patterns) with a business optimization learning model to determine where to build a new hotel. In some embodiments, the AI engine may use a natural language processing module to parse the prediction request.

[0063] In some embodiments, the output from the AI computer 130 may also include the period of time for which the output of the prediction request should be retained in the data repository. For example, the period of time may be 3 days, 1 week, 1 month, etc. The period of time may depend on the frequency of cacheable requests by the user. For example, cacheable requests made every hour may be stored for 3 days or a week, while cacheable requests made weekly may be stored for a month. The period of time may also depend on the context data associated with the cacheable request. For example, the output may depend on context data that is refreshed monthly. Thus, the output based on the context data may be cached for a period of one month.

[0064] Then, the orchestration service computer 120 can receive the output of the prediction request from the AI computer 130. The output can be, for example, the address of the optimized location of the new hotel, a list of addresses, the block where the new hotel is to be built, etc.

[0065] In step S510, the orchestration service computer 120 can send the output to the user device 110.

[0066] After receiving the output, the orchestration service computer 120 can initiate the storage of the prediction request and the output in the data repository 160. In some embodiments, initiating the storage of the prediction request can include, for example, Figure 5 the process, which will be described in further detail below. In other embodiments, the orchestration service computer 120 can directly store the prediction request and the output in the data repository. The orchestration service computer 120 can also store the context data and / or a reference to the location of the context data in the data repository 160.

[0067] The prediction request message can be a first prediction request message having a first prediction request in the form of a first AI call. Then, at a subsequent time but within the period when the output is stored in the data repository 160, the orchestration service computer 120 can receive a second prediction request in a second prediction request message. The second prediction request can be in the form of a second AI call. The second prediction request message can be received from a second user device (not shown) operated by a user who is the same as or different from the user of the first user device 110. In some embodiments, the second user device can be the same as the first user device 110. Alternatively, the second user device can be operated by a second user. Then, the orchestration service computer 120 can determine that the second prediction is a cacheable request using the process described in step S504.

[0068] Then, the orchestration service computer 120 can determine that the second prediction request corresponds to the first prediction request. In one embodiment, the orchestration service computer 120 can attempt to retrieve the second prediction request from the data repository 160. If the second prediction request is in the data repository 160 as the first prediction request, the orchestration service computer 120 can determine that the second prediction request corresponds to the first prediction request. In another embodiment, the orchestration service computer 120 can directly associate the second prediction request with the first prediction request.

[0069] Then, the orchestration service computer 120 can retrieve the output from the data repository 160 associated with the first prediction request and the second prediction request. Then, the orchestration service computer 120 can send the output to the second user device.

[0070] Figure 5A flowchart showing the use of a messaging system cache for output according to an embodiment. In some embodiments, Figure 5 The process of Figure 4 can occur immediately after the process of

[0071] Alternatively, the process can occur at the end of the day, hourly, or at some other suitable time.

[0072] In step S602, the orchestration service computer 120 can send a call information message to the messaging system computer 140. The messaging system computer 140 can receive the call information message and store the call information message in a queue. The call information message can include a prediction request, an output responsive to the prediction request, and context data. For example, the prediction request may be a question such as "Where is the best location to build a new restaurant in San Francisco?" Then, the output can be the location (e.g., address, neighborhood) determined by the AI engine as the best location to build the restaurant. The context data can include, for example, the geographical data of San Francisco, the locations of existing restaurants, and food transaction information for the past six months. In some embodiments, the call information message can include an identifier (e.g., a location in memory, a key) of the context data in the data repository 160 instead of the context data itself.

[0073] In step S604, the AI call information consumer 150 can be subscribed to the messaging system computer 140 and receive the call information message when it is published to the queue. The AI call information consumer 150 can alternatively or additionally retrieve information from the queue at will. For example, the AI call information consumer 150 can retrieve and store one call information message in the data repository 160 at a time. The AI call information consumer 150 can retrieve the call information message from the queue at a slower or faster rate than the orchestration service computer 120 sends the call information message.

[0074] Figure 6 A flowchart showing the execution of a call prediction process according to an embodiment.

[0075] In step S702, the invocation prediction processing computer 170 can determine that the AI computer 130 is operating below the threshold processor utilization rate. The invocation prediction processing computer 170 can perform this by monitoring the performance of the AI computer 130. For example, in some embodiments, the invocation prediction processing computer 170 can operate according to a schedule. For example, the invocation prediction processing computer 170 can check the AI computer 130 every 5 minutes, every hour, etc. The threshold can be, for example, 20% of the operating capacity of the AI computer. If the AI computer 130 is operating above the threshold, the invocation prediction processing computer 170 may go into a dormant state until its next check of the AI computer 130.

[0076] In step S704, after determining that the AI computer 130 is operating below the threshold, the invocation prediction processing computer 170 can request a set of cacheable requests from the orchestration service computer 120. The set of cacheable requests can include one or more cacheable requests. The orchestration service computer 120 can directly determine the set of cacheable requests, or can communicate with the AI computer 130 to determine the set of cacheable requests. For example, the orchestration service computer can maintain a list of the cacheable requests it receives, as well as the time and / or number of times each cacheable request has been made. Then, when the orchestration service computer 120 receives a request for the set of cacheable requests, it can use this information and frequency information to determine requests that are likely to be requested soon and are not in the cache. In some embodiments, the orchestration service computer 120 can maintain a record of the period of time (e.g., one day, one week) that a cacheable request remains in the cache. Alternatively, the orchestration service computer 120 can send a prediction request message to the AI computer 130 to determine the set of cacheable requests.

[0077] In some embodiments, when using the AI computer 130 to determine the set of cacheable requests, the prediction request can be "What are the N most likely cacheable requests to be requested?" and the context data can be, for example, the time and date and other environmental context information that may be suitable for determining cacheable requests. Then, the AI engine of the AI computer 130 can run an appropriate learning model to determine the set of cacheable requests. Then, the invocation prediction processing computer 170 can receive the set of cacheable requests from the orchestration service computer 120. In another embodiment, the invocation prediction processing computer 170 can directly send a prediction request for a set of cacheable requests to the AI computer 130.

[0078] In step S706, the call prediction processing computer 170 can select a cacheable request from a set of cacheable requests. For example, the call prediction processing computer 170 can select a first cacheable request from a set of cacheable requests. Alternatively, the call prediction processing computer 170 can prioritize a set of cacheable requests. For example, a set of cacheable requests can be differentiated in priority by the time when a user may request a cacheable request or the frequency of requesting a cacheable request. Then, the call prediction processing computer 170 can select a cacheable request that may be requested in the near future. The call prediction processing computer 170 can send the cacheable request to the AI computer 130 via the orchestration service computer 120 in a prediction request message. Alternatively, the call prediction processing computer 170 can directly send the cacheable request to the AI computer 130. The cacheable request can be in the form of an AI call.

[0079] In step S708, the orchestration service computer 120 can send the cacheable request (in the form of an AI call) to the AI computer 130. Before sending the cacheable request to the AI computer 130, in step S710, the orchestration service computer 120 can retrieve the context data of the cacheable request from the data repository 160. In some embodiments, the orchestration service computer 130 can determine and / or confirm that the cacheable request is cacheable. The AI computer 130 can use an AI engine to process the AI call with an AI model (e.g., a learning model including a neural network) or a combination of models, and generate an output in response to the prediction request message. After generating the output, the AI computer 130 can then send the output to the orchestration service computer 120.

[0080] The output can also include the time period for which the output should be retained in the data repository. For example, the time period can be 3 days, 1 week, 1 month, etc. The time period can depend on the frequency of the user requesting the cacheable request. For example, a cacheable request made hourly can be stored for 3 days or a week, while a cacheable request made weekly can be stored for a month. The time period can also depend on the context data associated with the cacheable request. For example, the output can depend on context data that is refreshed monthly. Thus, the output based on the context data can be cached for a time period of one month. In step S712, the orchestration service computer 120 can then return the output to the call prediction processing computer 170. Then, the call prediction processing computer 170 can receive the output.

[0081] In step S714, the invocation prediction processing computer 170 may store the output in the data repository 160. The invocation prediction processing computer 170 may also store the cacheable requests in the data repository 160. In some embodiments, the invocation prediction processing computer 170 may be a producer that can send information to the messaging system computer 140. When the invocation prediction processing computer 170 stores information in the data repository 160, the invocation prediction processing computer 170 may send an invocation information message to the messaging system computer 140.

[0082] Embodiments of the present invention provide several advantages. The advantages are as follows. By caching results, the time used to generate predictions is spread out. Thus, during peak periods, the AI computer can experience less stress because a portion of the results will be pre-computed and thus the AI engine does not need to compute all the results. As a result, the use of the processor is more persistent because during non-peak periods, the AI computer can be used to reduce the demand during peak periods. Additionally, predicting requests means that the cached information will be more useful. Using a messaging system can free up the orchestration service computer to continue processing a large number of prediction requests without waiting to confirm that the cacheable requests have been cached.

[0083] Any software component or functionality described in this application can be implemented as software code executed by a processor using, for example, conventional or object-oriented techniques and using any suitable computer language (e.g., Java, C++, or Perl). The software code can be stored as a series of instructions or commands on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), magnetic media such as a hard disk drive, or optical media such as a CD-ROM. Any such computer-readable medium can reside on or within a single computing device and can be on or within different computing devices in a system or network.

[0084] The above description is illustrative and not restrictive. Many variations of the present invention will become apparent to those skilled in the art upon review of this disclosure. Accordingly, the scope of the present invention may not be determined by reference to the above description, but may be determined by reference to the pending claims and their full scope or equivalents.

[0085] Without departing from the scope of the present invention, one or more features of any embodiment may be combined with one or more features of any other embodiment.

[0086] Unless specifically stated to the contrary, the recitation of "a / an" or "the" is intended to mean "one or more."

[0087] All patents, patent applications, publications, and descriptions mentioned above are incorporated by reference in their entirety for all purposes. It is not admitted that they are prior art.

Claims

1. A method, comprising: determining by a call prediction processing computer that an AI computer operates below a threshold processor utilization rate; based on the AI computer operating below the threshold processor utilization rate: sending, by the call prediction processing computer, a request to the AI computer, the request for determining a set of cacheable requests that are most likely to be requested by a user in the future, wherein, after receiving the request, the AI computer determines the set of cacheable requests, receiving, by the call prediction processing computer, the set of cacheable requests and context data respectively associated with each cacheable request in the set of cacheable requests, sending, by the call prediction processing computer, one or more cacheable requests from the set of cacheable requests and the context data respectively associated with each cacheable request in the one or more cacheable requests to the AI computer for processing, wherein receiving the one or more cacheable requests and the context data causes the AI computer to process the one or more cacheable requests and the context data to generate one or more outputs respectively responsive to the one or more cacheable requests, and generating a duration period as part of each output in the one or more outputs based on the context data corresponding to each cacheable request, and receiving, by the call prediction processing computer, the one or more outputs generated by the AI computer for the one or more cacheable requests; and storing, by the call prediction processing computer, each output in the one or more outputs corresponding to the one or more cacheable requests in a data repository for a duration period reserved by the AI computer for each cacheable request, wherein the data repository stores a plurality of outputs corresponding to a plurality of prediction requests previously processed by the AI computer, and the one or more outputs are added to the plurality of outputs, wherein, based on a subsequent prediction request received that is a cacheable request and included in the plurality of prediction requests, an output corresponding to the subsequent prediction request is retrieved from the data repository without sending the subsequent prediction request to the AI computer for processing, and wherein the cacheable request is a prediction request whose output does not change over time.

2. The method according to claim 1, wherein the threshold processor utilization rate is about 20% or less.

3. The method according to claim 1, wherein the output corresponding to the subsequent prediction request is received from an orchestration service computer.

4. The method according to claim 1, wherein the AI computer includes an artificial neural network.

5. The method according to claim 1, wherein the call prediction processing computer sends the one or more cacheable requests to the AI computer via an orchestration service computer.

6. The method according to claim 1, wherein the call prediction processing computer prioritizes the set of cacheable requests to select the one or more cacheable requests.

7. The method according to claim 1, wherein as long as the AI computer is below the threshold processor utilization rate, the call prediction processing computer continuously sends cacheable requests from the set of cacheable requests to the AI computer.

8. A system, comprising: A call prediction processing computer, the call prediction processing computer including a processor and a computer-readable medium, the computer-readable medium including code for implementing a method, the method including: Determine that the AI computer is operating below the threshold processor utilization rate; Based on the AI computer operating below the threshold processor utilization rate: Send a request to the AI computer, the request for determining a set of cacheable requests that are most likely to be requested by a user in the future, wherein, after receiving the request, the AI computer determines the set of cacheable requests, Receive the set of cacheable requests and context data respectively associated with each cacheable request in the set of cacheable requests, Send one or more cacheable requests from the set of cacheable requests and the context data respectively associated with each of the one or more cacheable requests to the AI computer for processing, wherein receiving the one or more cacheable requests and the context data causes the AI computer to process the one or more cacheable requests and the context data to generate one or more outputs respectively responsive to the one or more cacheable requests, and generate a duration period as part of each of the one or more outputs based on the context data corresponding to each cacheable request, and Receive the one or more outputs generated by the AI computer for the one or more cacheable requests; and Store each of the one or more outputs corresponding to the one or more cacheable requests in a data repository for a duration period reserved by the AI computer for each cacheable request, wherein the data repository stores a plurality of outputs corresponding to a plurality of prediction requests previously processed by the AI computer, and the one or more outputs are added to the plurality of outputs, wherein, based on a subsequent prediction request received that is a cacheable request and included in the plurality of prediction requests, the output corresponding to the subsequent prediction request is retrieved from the data repository without sending the subsequent prediction request to the AI computer for processing, and wherein the cacheable request is a prediction request whose output does not change over time.

9. The system according to claim 8, wherein the threshold processor utilization rate is 20% or less.

10. The system according to claim 8, wherein the set of cacheable requests includes cacheable requests that are most likely to be requested.

11. The system according to claim 8, wherein the output corresponding to the subsequent prediction request is received from an orchestration service computer.

12. The system according to claim 8, wherein the call prediction processing computer sends the one or more cacheable requests to the AI computer via an orchestration service computer.

13. The system according to claim 8, wherein the call prediction processing computer prioritizes the set of cacheable requests to select the one or more cacheable requests.

14. The system according to claim 8, wherein the call prediction processing computer continues to send cacheable requests from the set of cacheable requests to the AI computer as long as the AI computer is below the threshold processor utilization rate.

15. A method performed by an orchestration service computer, the method comprising: sending a request to an AI computer determined to be operating below a threshold processor utilization rate, the request for determining a set of cacheable requests that are most likely to be requested by a user in the future, wherein, after receiving the request, the AI computer determines the set of cacheable requests; receiving the set of cacheable requests and context data respectively associated with each of the cacheable requests in the set of cacheable requests; sending one or more cacheable requests from the set of cacheable requests and the context data respectively associated with each of the one or more cacheable requests to the AI computer for processing, wherein receiving the one or more cacheable requests and the context data causes the AI computer to process the one or more cacheable requests and the context data to generate one or more outputs respectively responsive to the one or more cacheable requests, and to generate a duration period as part of each of the one or more outputs based on the context data corresponding to each cacheable request; receiving the one or more outputs generated by the AI computer for the one or more cacheable requests; sending the one or more outputs respectively responsive to the one or more cacheable requests to a prediction processing computer, the prediction processing computer storing each of the one or more outputs corresponding to the one or more cacheable requests in a data repository for a duration period reserved by the AI computer for each cacheable request, wherein the data repository stores a plurality of outputs corresponding to a plurality of prediction requests previously processed by the AI computer, and the one or more outputs are added to the plurality of outputs; receiving a prediction request message from a user device including a subsequent prediction request in the form of an AI call; determining that the subsequent prediction request is a cacheable request, wherein the subsequent prediction request is included in the plurality of prediction requests; retrieving from the data repository an output corresponding to the subsequent prediction request without sending the subsequent prediction request to the AI computer for processing; and sending the output corresponding to the subsequent prediction request to the user device, wherein the cacheable request is a prediction request whose output does not change over time.

16. The method according to claim 15, wherein the context data includes transaction data.

17. The method according to claim 15, wherein determining that the subsequent prediction request is the cacheable request includes using the AI computer.

18. The method according to claim 15, wherein determining that the subsequent prediction request is the cacheable request includes evaluating the user device.

19. The method according to claim 15, wherein the user device is a computer that processes a network.

20. A system comprising: An orchestration service computer, the orchestration service computer includes a processor and a computer-readable medium, the computer-readable medium includes code for implementing a method, the method includes: Sending a request to an AI computer determined to be operating below a threshold processor utilization rate, the request for determining a set of cacheable requests that are most likely to be requested by a user in the future, wherein, after receiving the request, the AI computer determines the set of cacheable requests; Receiving the set of cacheable requests and context data respectively associated with each cacheable request in the set of cacheable requests; Sending one or more cacheable requests from the set of cacheable requests and the context data respectively associated with each of the one or more cacheable requests to the AI computer for processing, wherein receiving the one or more cacheable requests and the context data causes the AI computer to process the one or more cacheable requests and the context data to generate one or more outputs respectively in response to the one or more cacheable requests, and generating a duration period as part of each of the one or more outputs based on the context data corresponding to each cacheable request; Receiving the one or more outputs generated by the AI computer for the one or more cacheable requests; Sending the one or more outputs respectively in response to the one or more cacheable requests to a prediction processing computer, the prediction processing computer stores each of the one or more outputs corresponding to the one or more cacheable requests in a data repository for a duration period reserved by the AI computer for each cacheable request, wherein the data repository stores a plurality of outputs corresponding to a plurality of prediction requests previously processed by the AI computer, and the one or more outputs are added to the plurality of outputs; Receiving a prediction request message from a user device that includes a subsequent prediction request in the form of an AI call; Determining that the subsequent prediction request is a cacheable request, wherein the subsequent prediction request is included in the plurality of prediction requests; Retrieving the output corresponding to the subsequent prediction request from the data repository without sending the subsequent prediction request to the AI computer for processing; Sending the output corresponding to the subsequent prediction request to the user device, wherein the cacheable request is a prediction request whose output does not change over time.

21. The system according to claim 20, wherein the context data includes transaction data.

22. The system according to claim 20, wherein determining that the subsequent prediction request is the cacheable request includes using the AI computer.

23. The system according to claim 20, wherein determining that the prediction request is the cacheable request includes evaluating the user device.

24. The system according to claim 20, wherein the user device is a computer that processes the network.