A computer-implemented method for generating a response to a user input
A distributed network of specialized SLMs efficiently processes user inputs, addressing hardware and electricity costs, and reducing biasness and hallucination by utilizing a diverse network of external computers.
Patent Information
- Application Number
- PCT/MY2024/050014
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-09-04
AI Technical Summary
Current AI chatbots require substantial hardware and electricity costs, and suffer from issues such as biasness and hallucination due to their reliance on large language models with fixed architectures.
A distributed network of external computers, each with specialized knowledge, processes user inputs in a swarm of Small Language Models (SLMs) connected via a host server, dynamically selecting resources based on input complexity and cost, minimizing hardware investment and reducing biasness and hallucination.
This approach significantly reduces computational resource usage and costs while enhancing response accuracy and generalization, leveraging a diverse network of computers to counteract biased or inaccurate outputs.
Smart Images

Figure MY2024050014_04092025_PF_FP_ABST
Abstract
Description
A COMPUTER-IMPLEMENTED METHOD FOR GENERATING A RESPONSE TO A USER INPUT
[0001] The present invention relates to a computer-implemented method for generating responses to user queries or inputs in a conversational manner.
[0002] In recent years, automated chatbots using artificial intelligence (AI) to generate human-like responses to a user in a conversational manner have increased in popularity. It has become a new and interesting way for gathering information, completing some types of tasks, or even just passing the time by having a chat with an AI. It has generally been accepted that in order to make the responses seem more natural and human-like, more processing power is required. More processing power enables the chatbot to generate responses with more complexity. One solution created by OpenAI and known as ChatGPT uses layers of large language models (LLMs) to greatly increase the complexity of the responses it generates. With the latest iteration of ChapGPT, it is difficult to differentiate between a response it generated and one from a human. The more layers ChatGPT uses, the higher the amount of processing power and thus complexity it can generate in its responses.
[0003] The problem with this approach is the substantial amount of costly hardware required to populate all those layers. Furthermore, the graphical processing units (GPUs) that ChatGPT is configured to use is a particularly costly type, which is compounded when a large number is required for the whole system. Apart from hardware costs, there is also the cost of the electricity required to power all that hardware. For systems as large as the current or next generation of OpenAI’s chatbot, the electricity cost is very substantial.
[0004] Further to the issue of cost, there are still issues with accuracy and generalization capabilities when generating a response to a user input. One particular side effect of using LLM architecture is something known as biasness. Biasness occurs when an inaccurate, unpopular, or even entirely wrong point of view is learned by the LLM. Since an LLM is essentially one large unit that acts as a single module, any point of view that ‘catches hold’ may not find much resistance as there are no independent or opposing modules to check and balance this view, but instead only more positive reinforcement, thus allowing wayward opinions or beliefs to take hold in the LLM.
[0005] Yet another issue more common with LLMs is something known as hallucination. This is a negative phenomenon that occurs when an AI does not have complete information or knowledge of a topic. When a user then asks about this topic, the AI might try to fill the missing information with completely fictional data.
[0006] What is needed in the art is a chatbot that requires less hardware and having lower running costs, whilst still being able to generate natural, human-like responses with high accuracy, efficiency and generalization capabilities, and with a more robust defense against phenomena such as biasness and hallucination.
[0007] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential characteristics of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0008] In one aspect, this invention is a computer-implemented method for generating a response to a user input (which we shall refer to as a “chatbot”), comprising receiving an input from a user, then generating a response to that input which is accurate and natural, with as much information and knowledge as necessary, sending the response back to the user and then receiving the next input from the user, and so on, in a conversational fashion.
[0009] The chatbot of this invention can generate an accurate and natural response with far more efficient use of computational resources than currently known methods. It is able to do this in part due to use of a novel distribution mechanism, where inputs are sent from a host server controlled and operated by the provider / operator of the chatbot of this invention (chatbot provider), to a distributed network of external computers that are neither controlled nor operated by the chatbot provider. This is one of the main features of this invention which sets it apart from currently known methods, and will be described thoroughly here and in the detailed description. The distributed network comprises a plurality of external computers, each installed with software provided by the chatbot provider, that allows for communication and data transfer between the external computer and the host server. Each external computer is owned by a third-party entity such as an individual who has a computer at home that he doesn’t use all the time, or perhaps a company that has a decent amount of computing power that goes unused at night. These computers would be fully utilized after joining the network of this invention. Each external computer is persuaded to specialize in different knowledge topics to the other external computers, so that some external computers will be more proficient than others at handling any given user input or query. A list of specializations and level of capability in each specialization of all the external computers is stored and regularly updated on a database on the host server.
[0010] When the host server receives an input from the user, it processes the input in various ways, including ascertaining how complex the input is. Complex or challenging inputs are broken down into smaller units which we shall call phrases. This dividing of the original input is typically done along the lines of topic or category by the orchestrator model. The host server then selects a set of external computers from all that currently available on the network. The selection of external computers is based on how well and efficiently each external computer can handle or process the phrases. In essence, what we have here is a network that can cater to almost any size or amount of computational processing power required for a given user input or query. This makes the chatbot of this invention incredibly efficient in its use of computational resources, in addition to the much lower capital investment in hardware.
[0011] Regarding the problem of biasness seen with chatbots based on the LLM architecture, since the chatbot of this invention utilizes a large number of external computers which generally do not directly influence or affect each other, there is a great amount of opposing views and checks and balances that do not allow a rogue point of view to grow out of hand.
[0012] This same feature of the sheer number of external computers on the network of this invention also reduces the chances of hallucination. Since each external computer is persuaded to specialize in different topics, there would be an immense amount of knowledge comprehensively covering a wide variety of topics, readily available from the network, which all but eliminates any of the formidable-sized gaps in knowledge that opens the door to hallucination occurring.
[0013] Yet another advantage offered by the chatbot of this invention is a significant reduction in the computational processing required for handling similar tasks compared to chatbots that use the LLM architecture. In the present invention, the user input is broken up into smaller pieces which are processed separately by different external computers. The outputs of the external computers are then sent to an aggregator model located in the host server. The aggregator model has the very easy task of merely joining the pieces together in a coherent manner to form the response to the user.
[0014] In one embodiment of the chatbot of this invention, the external computers are provided with incentives that may include some form of currency that may be exchanged for real-world money.
[0015] This invention thus relates to a computer-implemented method for generating a response to a user input, the method comprising the following steps: receiving an input from a user onto a host server located on a computer network; the host server determining properties of the input, including its size and complexity; providing, on the host server, a database of external computers available on the network, the external computers having different capabilities and specializations from each other, which make some more proficient than others at handling any given input, the database includes information on the said capabilities and specializations; if the input is above a predefined degree of complexity or size, the host server dividing the input into two of more phrases; the host server selecting a set of external computers from the said database, the said selection based on any or a combination of the following:how proficient they are at processing the said input or phrases, measured by the quality of outputs generated,how efficient they are at processing the said input or phrases, measured by the amount of computational resources used, andthe cost of payment demanded by each external computer.The host server then sending a task offer to each external computer in the said set, or to the one external computer for the case where the input was not divided, the task requiring each external computer to process its input or phrase to produce an output that meets predefined standards. Next, the external computer or set of external computers completes the tasks and sending their said outputs to the host server; the host server generating a response based on the said outputs; and the host server sending the said response to the user.
[0016] In a preferred embodiment, the host server is controlled and operated by a provider of the computer-implemented method and the external computers are neither controlled nor operated by the provider of the computer-implemented method. In other words, the host server, which is controlled and operated by the entity that owns and / or provides the chatbot service receives the user input, the after some processing, sends it to a network of external computers. Each external computer is owned and / or controlled by third parties, and not by the entity that owns and / or provides the chatbot service. The external computers are given incentives or rewards for each task they complete for the host server. This allows the chatbot service to run on a relatively smaller amount of hardware and lower electricity costs. Furthermore, for every response generated, only the exact number of external computers required is used - the amount and capabilities of external computers can be sized according to the needs of each user input, resulting in an extremely efficient usage of computing hardware and other resources. In essence, the power, capacity and capabilities of the chatbot of this invention are tailored for each user input it receives. This use of third-party external computers also allows an organic growth of the chatbot’s capacity and capabilities, in the form of available external computers, as the task rewards persuade more and more third-party external computers to join the network.
[0017] In another preferred embodiment, the host server comprises two or more models arranged in a chained network configuration from a orchestrator model that receives the input from the user to a distribution hub, and wherein the external computers are connected to the distribution hub in a distributed network configuration, whereby each external computer is connected directly to the distribution hub without being connected to each other.
[0018] In another preferred embodiment, each external computer is installed with software that is provided by the chatbot operator that allows it to communicate with and transfer data to and from the host server. This software allows the external computer to function like an artificial neural network, comprising models, layers and nodes. Each external computer that receives a task offer from the host server may decline the task offer, in which case the host server selects a replacement external computer to offer the task to.
[0019] In another preferred embodiment, the host server checks that each external computer offered a task has enough computational power to handle the offered task within a specified amount of time.
[0020] In another preferred embodiment, the capabilities and specializations of each external computer is achieved and defined through training of datasets contained within each external computer. Each external computer learns from each output it generates by way of assigning weights to new information to improve on its said datasets, with information it considers more useful in the future given a higher weight and thus a larger effect on its datasets, such that, when given similar inputs or phrases to solve in the future, it does so with either better quality of output, less usage of computational, or both. Quality of output here may comprise its accuracy or naturalness, or both.
[0021] In another preferred embodiment, a content detective located on the host server, the content detective scanning the external computers every time a change is detected in any dataset located within the network, the said scanning designed to find any or a combination of the following: illegal or prohibited material, malware, and viruses.
[0022] In another preferred embodiment, each external computer is a Small Language Model (SLM) running on at least one CPU and at least one GPU. Because the user input is broken down into smaller units called phrases that are distributed to multiple SLMs, and each individual SLM is orders of magnitude less complex compared to existing chatbots such as OpenAI’s ChatGPT that use Large Language Models (LLM) with trillions of parameters, the cost of inference is significantly lower in the chatbot of this invention.
[0023] Additional features and advantages of the invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the invention. The features and advantages of the invention may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. These and other features of the present invention will become more fully apparent from the following description and appended claims, or may be learned by the practice of the invention as set forth hereinafter.
[0024] 1. Current AI chatbots require high capital expenditure on hardware and use massive amounts of electricity during operation.
[0025] 2. Undesirable behaviour in the chatbot’s responses such asbiasness, where the AI of the chatbot starts to form or believe in fake, unpopular or downright wrong opinions or beliefs, and hallucination, where the AI when asked about something that falls in a gap in its knowledge, it fills in the knowledge gaps with fake and / or fictional ‘knowledge’.
[0026] 1. The chatbot of this invention is able to customize and tailor instantly and quite precisely the size, capabilities and specializations of computational resources it uses for each user input, e.g. employing more computational resources to handle more complex inputs, and minimal computational resources to handle basic inputs. This is an entirely different approach to the currently available chatbots that use a fixed, inflexible network of computational resources. Most of the computational resources available to the chatbot of this invention is on a portion of the network that is not owned or operated by the provider or operator of the chatbot, so the capital investment needed for hardware is very low. Furthermore, since the external computers employed to assist in processing a response were selected specifically for that particular input, there is a very high efficiency, which translates to low electricity costs.
[0027] 2. The large number of external computers in the network of this invention provides more resilience to AI behaviour such as biasness and hallucinations. It provides many opposing points of view or beliefs, making it very unlikely that a single rogue opinion is reinforced over time to become something almost like fact to the AI. The large number of external computers also ensure that there are few, if any, sizable knowledge gaps, which reduces the chances of the hallucination phenomenon.
[0028] To further clarify the above and other advantages and features of the present invention, a more detailed description is provided with reference to specific embodiments, and which are illustrated in the appended drawings.
[0029] shows a diagram of a known method for automated chatting (prior art).
[0030] shows a general diagram of an embodiment of the method of the present invention.
[0031] shows a detailed diagram of an embodiment of the method of the present invention.
[0032] The present invention may be embodied in other specific forms without departing from its essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is, therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
[0033] Referring first to, which shows a typical neural network known in the industry as a Large Language Model (LLM), such as that used by OpenAI’s ChatGPT. The “neurons” or nodes are divided into layers, through which a user input will be passed. Nodes in each layer do not have contact with each other, but they do have contact with nodes from the preceding and subsequent layers. In this type of architecture, in general, the higher the number of layers provided, the more complex and accurate the responses can be. So the only way to improve on the quality of the responses is by adding more layers to the network, which is more computer hardware, and therefore an expense. More layers also means more computational processing, which increases electricity costs. More importantly, barring the afore-mentioned physically adding layers to the network, the number of nodes remains fixed, which makes it quite inflexible. User inputs run the entire gamut in size and complexity, and the same neural network is made to handle really basic queries, e.g. “What day is today?” where most of the network is not required, to extremely complex ones where the datasets aren’t of a good enough quality to generate a good quality response. This is clearly inefficient and less than ideal.
[0034] The computer-implemented method for generating a response to a user input (we will refer to this as “chatbot” for brevity) of this invention is implemented on a network of computers that has been given the name Swarm Artificial Intelligence Language Model (SAILM) network, as it is a swarm of AI language models, which may include Large Language Models (LLMs) and Small Language Models (SLMs) connected together to form this network. This network is diagrammatically shown in Figures 2 and 3.shows a broader, more high-level view of the network, highlighting the principle constituents and the method in general, whileshows a more detailed view of the network which includes everything essential for the method to be carried out.
[0035] Referring first to, there is shown a network connecting a user (1), a host server (10), and a plurality of external computers (50), the data connection being either wireless or a wired data cable such as ethernet. The user (1) sends an input to the host server (10), which may contain any or a combination of the following media types: text, audio, image, and video. The host server (10) processes the input into tasks, with each task being sent (or offered) to one or more external computers (50). The external computers (50) complete the task by processing the data it received from the host server (10) into an output. The host server (10) then compiles the outputs of each external computer (50) into a response that it sends back to the user (1).
[0036] For both Figures 2 and 3, the external computers (50), which are also known as SLMs (Small Language Models), are neither owned nor operated by the entity that is providing, managing and operating the chatbot of this invention. Instead, the external computers (50) are owned by third-party entities typically comprising individuals, groups, or corporations that may be persuaded to join the SLLM network in order to earn the incentives given for completing tasks within the network. These third-party entities register their computers with the chatbot operator, who then provides them with proprietary software to be installed on the said computers, the software creates a virtual Small Language Model (SLM) within each computer (known as external computers within the SLLM network). In this way, the SLLM network can be regarded as a network of neural networks, since each SLM is a neural network. In terms of hardware, each external computer (50) must be equipped with at least a CPU and a GPU with sufficient computational processing power as mandated by the chatbot operator. Ideally, each external computer (50) is trained with a different set of capabilities and specializations to other external computers. The owner / operator of each external computer (50) is responsible for training the external computer(s) (50) under their care. The training is done on datasets embedded within each external computer (50).
[0037] Referring now to, there is shown the same network ofin greater detail, connecting a user (1), a host server (10), and a plurality of external computers (50). Only three external computers (51, 52, 53) are shown here in the interest of brevity, when in reality, there would be at least forty, and up to tens of thousands or more of external computers (50) in the network. The user (1) sends an input to an orchestrator model (12) located on the host server (10). The orchestrator model (12) manages the overall execution of the network, coordinates the flow of data and control between the external computers (50), handles input and output formats, manages intermediate results, and ensures that the external computers (50) execute in the correct sequence. The orchestrator model (12) also determines if the input is above a predetermined level of complexity or size, in which case it breaks the input down into two or more phrases. The phrases follow a template that makes it easily understood and processed by an external computer (50). The orchestrator model (12) stores a database (13) of available external computers (50) including information on the specializations and capabilities of each external computer (50). The orchestrator model (12) also uses a confidence matrix to determine the degree of confidence of each external computer (50) regarding a particular input or phrase. Referring to this database (13) and the confidence matrix, a distribution hub (14) selects one or more external computers (50) for processing the input or phrases into an output that includes information or data that will eventually be used to compile the intended response to the user. The external computers (50) selected for any one input are those that maximize the quality of the eventual response to the user, or minimize the computational resources required, or both. A third factor considered by the distribution hub (14) in selecting external computers (50) for a task is the amount of payment to the external computer for completing the task. The selected external computers (50) are then sent a task offer each via a distribution hub (14) located within the host server (10).
[0038] The distribution hub (14) functions as a hub that connects to all the external computers (50) in a distributed network. The distribution hub (14) also handles the selection of the external computers (50) for handling the processing of phrases into outputs, as well as coordinating data from the orchestrator model (12) to the external computers (50).
[0039] An external computer (50) that is offered a task by the distribution hub (14) can decline the offer; in which case the orchestrator model (12) reselects the next best set of external computers (50) for the phrases in hand.
[0040] The external computers (50) that do accept the task offers complete the tasks by processing the phrases they received from the distribution hub (14) into outputs that meet a predefined standard, and generally speaking in a form closer to the eventual response that is to be sent back to the user. Each external computer (50) then sends their output to an aggregator model (18) located within the host server (10). The aggregator model (18) compiles the outputs of each external computer (50) into a response that is sent back to the user (1).
[0041] Also seen inis a content detective (16) model, which scans all external computers (50) in the network every time a change is detected in any dataset contained within the network. The content detective (16) also does a scan on external computers (50) that have newly joined the network. The scans are designed to detect illegal or prohibited content, malware and viruses. If the content detective (16) detects any of these in an external computer (50), it blocks or quarantines the content so that it is not able to move within the network. The owners / operators of the offending external computer may also be penalized or blacklisted to dissuade them from allowing prohibited material onto their datasets.
[0042] The following are examples showing how the method would handle certain user inputs.
[0043] Example 1:
[0044] User input: “Create a two player 2-D game based on football using JavaScript”. External computer A (51) has good JavaScript capabilities (8 / 10), external computer B (52) has experience in designing and coding computer games (6 / 10), and external computer C (53) has knowledge of the game of football (10 / 10). The host server (10) might calculate that the best (least costly and / or produces the highest quality output) solution in this case would be to create and offer tasks to each of the three external computers. After accepting the tasks, each external computer passes their knowledge to the aggregator model (18), which then compiles a JavaScript code for the 2-D football game. Alternatively, external computer B (52) might have “internet access” and “disk space for new knowledge available” listed as its capabilities, and the host server (10) might then calculate that is works out less costly or that a higher quality response would be generated if the input is split into two phrases and two tasks offered to external computer A (51) and external computer B (52).
[0045] Example 2:
[0046] User Input: “Make me an audio file of a recipe for Korean fried chicken in Shakespearean English and rapped in a 90s hip hop style”. External computer A (51) has a highly specialized knowledge (9 / 10) of Korean fried chicken and 90s hip hop (8 / 10). External computer B (52) has an excellent knowledge of Shakespeare (10 / 10) and a passable knowledge of Korean food (6 / 10). External computer C (53) has a good knowledge of Shakespeare’s works (8 / 10) and 90s hip hop (8 / 10). There doesn’t seem to be any obvious two external computers for the host server (10) to select. However, the host server (10) still runs the calculations of cost vs response quality for all promising combinations, and uses the results to select the combination of external computers it selects to offer tasks to.
[0047] User (1)
[0048] Host server (10)
[0049] Orchestrator model (12)
[0050] Database (13)
[0051] Distribution hub (14)
[0052] Content detective (16)
[0053] Aggregator model (18)
[0054] External computers / SLMs (50)
[0055] External computer A (51)
[0056] External computer B (52)
[0057] External computer C (53)
Claims
A computer-implemented method for generating a response to a user input, the method comprising:receiving an input from a user (1) onto a host server (10) located on a computer network;providing, on the host server (10), a database (13) of external computers (50) available on the network, the external computers (50) having different capabilities and specializations, which make some more proficient than others at handling any given input, the database (13) includes information on the said capabilities and specializations;if the input is above a predefined degree of complexity or size, the host server (10) dividing the input into two of more phrases;the host server (10) selecting a set of external computers (50) from the said database (13), the said selection based on any or a combination of the following: how proficient they are at processing the said input or phrases, measured by the quality of outputs generated, how efficient they are at processing the said input or phrases, measured by the amount of computational resources used, and the cost of payment demanded by each external computer (50);the host server (10) sending a task offer to each external computer (50) in the said set, the said task requiring each external computer (50) to process its input or phrase to produce an output that meets predefined standards;the set of external computers (50) completing the tasks and sending their said outputs to the host server (10);the host server (10) generating a response based on the said outputs; andthe host server (10) sending the said response to the user (1).The computer-implemented method of claim 1, wherein the host server (10) is controlled and operated by a provider of the computer-implemented method and the external computers (50) are neither controlled nor operated by the provider of the computer-implemented method.The computer-implemented method of claim 1, wherein the host server (10) comprises two or more models arranged in a chained network configuration from a orchestrator model (12) that receives the input from the user (1) to a distribution hub (14), and wherein the external computers (50) are connected to the distribution hub (14) in a distributed network configuration, whereby each external computer (50) is connected directly to the distribution hub (14) without being connected to each other.The computer-implemented method of claim 1, wherein each external computer (50) is installed with software that allows it to function like an artificial neural network, comprising models, layers and nodes.The computer-implemented method of claim 1, wherein each external computer (50) that receives a task offer from the host server (10) may decline the task offer, in which case the host server (10) selects a replacement external computer (50) to offer the task to.The computer-implemented method of claim 1, wherein the said capabilities and specializations of each external computer (50) is achieved and defined through training of datasets contained within each external computer (50)The computer-implemented method of claim 6, wherein each external computer (50) learns from each output it generates by way of assigning weights to new information to improve on its said datasets, with information it considers more useful in the future given a higher weight and thus a larger effect on its datasets, such that, when given similar inputs or phrases to solve in the future, it does so with either better quality of output, less usage of computational, or bothThe computer-implemented method of claim 1, further comprising a content detective (16) located on the host server (10), the content detective (16) scanning the external computers (50) every time a change is detected in any dataset located within the network, the said scanning designed to find any or a combination of the following: illegal or prohibited material, malware, and viruses..
Citation Information
Patent Citations
Method and system for high-throughput distributed computing of computational jobs
EP4184325A1
Information processing method using network, information processing system, information processing program and recording medium
JP2002351681A
System and method for high-speed clustering
KR1020170039397A
Method for registering edge computing access and an edge computing node device
US20210344705A1
Managing and deploying applications in multi-cloud environment
US20230168874A1