Response output device and system

The response output device enhances user interaction by selecting appropriate language models based on instruction categories, addressing the inadequacies of existing AI response output technologies.

JP2025163614APending Publication Date: 2025-10-29MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024067055
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-17
Publication Date
2025-10-29

AI Technical Summary

Technical Problem

Existing response output technologies using artificial intelligence do not adequately consider suitable configurations for user interactions.

Method used

A response output device that includes a control unit to generate and transmit instruction sentences to large-scale language models, and a display unit to show responses, with the ability to select an appropriate language model based on the category of the instruction.

Benefits of technology

Provides a more suitable response output technique by enhancing user interaction through selective use of language models, improving response relevance and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025163614000001_ABST
    Figure 2025163614000001_ABST
Patent Text Reader

Abstract

To contribute to sustainable development goals (SDGs), namely "9 Industry, Innovation, and Infrastructure" and "11 Sustainable Cities and Communities" by providing a more suitable artificial intelligence response output technique.SOLUTION: A response output device comprises: a control section for generating an instruction sentence on the basis of a user's input, transmitting the instruction sentence to a large-scale language model (LLM), and acquiring a response sentence as a response generated by the LLM from the LLM; and a display section for displaying the instruction sentence and the response sentence. When there are a plurality of usable LLMs as LLM, the control section acquires a response sentence as a response generated by the LLM selected from the plurality of LLMs according to the category of the instruction sentence.SELECTED DRAWING: Figure 3A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a response output device and system. [Background technology]

[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special table 2019-528512 publication Summary of the Invention [Problem to be solved by the invention]

[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.

[0005] An object of the present invention is to provide a more suitable response output technique. [Means for solving the problem]

[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes multiple means for solving the above problem, and one example thereof may be configured as follows: A response output device, comprising: a control unit that generates an instruction sentence based on a user's input, transmits the instruction sentence to a large-scale language model (LLM), and acquires an answer sentence from the LLM as a response generated by the LLM; and a display unit that displays the instruction sentence and the answer sentence, wherein if there are multiple LLMs that can be used as the LLM, the control unit acquires the answer sentence as a response generated by an LLM selected from the multiple LLMs according to the category of the instruction sentence. [Effects of the Invention]

[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments. [Brief explanation of the drawings]

[0008] [Figure 1A] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 1B] 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 1C] 1 is a diagram showing an example of the operation of an artificial intelligence response output device and system according to an embodiment of the present invention; [Figure 2A] 1 is a diagram illustrating an example of a response output device and a system according to an embodiment of the present invention. [Figure 2B] 1 is a diagram illustrating an example of a response output device and a system according to an embodiment of the present invention. [Figure 2C] FIG. 10 illustrates an example of a field / category classification according to one embodiment. [Figure 2D] FIG. 1 is a diagram illustrating an outline of each embodiment. [Figure 2E] FIG. 10 illustrates an example of processing by a control unit according to an embodiment. [Figure 2F] FIG. 10 illustrates an example of a process for analyzing directives according to one embodiment. [Figure 2G] FIG. 10 is a diagram illustrating an example of the presentation of LLM information used in a response according to one embodiment. [Figure 2H] FIG. 10 illustrates an example of changing the LLM used in a response according to one embodiment. [Figure 2I] FIG. 10 illustrates an example of an LLM recommendation for use in an answer, according to one embodiment. [Figure 2J] 1 illustrates an example of a response output device and a system according to an embodiment; [Figure 3A] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 3B]FIG. 10 illustrates an example of a category classification process according to one embodiment. [Figure 3C] FIG. 10 illustrates an example of an LLM selection process according to one embodiment. [Figure 3D] FIG. 10 is a diagram illustrating a specific example of processing according to an embodiment. [Figure 3E] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 3F] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 3G] FIG. 10 illustrates an example of a category classification process according to one embodiment. [Figure 3H] FIG. 10 illustrates an example of a category classification process according to one embodiment. [Figure 4A] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 4B] FIG. 10 is a diagram illustrating a specific example of processing according to an embodiment. [Figure 4C] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 4D] FIG. 10 illustrates an example screen according to one embodiment. [Figure 4E] FIG. 10 illustrates an example screen according to one embodiment. [Figure 4F] FIG. 10 illustrates an example screen according to one embodiment. [Figure 5A] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 5B] FIG. 10 is a diagram illustrating a specific example of processing according to an embodiment. [Figure 5C] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 6A] FIG. 10 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. [Figure 6B] FIG. 10 illustrates an example screen according to one embodiment. [Figure 6C] FIG. 10 illustrates an example screen according to one embodiment. [Figure 6D] FIG. 10 illustrates an example screen according to one embodiment. [Figure 6E] FIG. 10 illustrates an example screen according to one embodiment. [Figure 6F] FIG. 10 illustrates an example screen according to one embodiment. [Figure 6G] FIG. 10 illustrates an example screen according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.

[0010] Note that if the AI ​​response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI ​​response output device has an audio output function, it may be referred to as an audio output device. The AI ​​response output device may simply be referred to as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI ​​response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI ​​response output device or the display output of the AI ​​response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI ​​response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI ​​response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.

[0011] Example 1 As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.

[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI ​​response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system including the AI ​​response output device 10010 and the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.

[0013] In the example of FIG. 1A, the AI ​​response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Furthermore, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which the pixels emit light themselves. Furthermore, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.

[0014] 1A, the audio output unit 1140 provided in the AI ​​response output device 10010 is configured with a speaker. The AI ​​response output device 10010 also has a microphone 1139, which can pick up the user's voice. By audio input from the microphone 1139 or operation input from the user via an operation input unit (described later), the AI ​​response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.

[0015] The AI ​​response output device 10010 may be provided with a local large-scale language model within the AI ​​response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or an audio output from the audio output unit 1140.

[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.

[0017] Alternatively, the AI ​​response output device 10010 may also include a local large-scale language model, and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI ​​response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.

[0018] The configuration when the AI ​​response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions, or may go via a router or repeater. The AI ​​response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI ​​response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI ​​response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as one system.

[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI ​​response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.

[0020] 1A shows an example in which a display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of FIG. 1A, the prompt display area 10051 displays an icon 10052 representing a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of FIG. 1A, the artificial intelligence response display area 10061 displays an icon 10062 representing an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI ​​response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI ​​response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.

[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment. These large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in these artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called Transformer. Reference 1, for example, has been published as an example of the learning of these models.

[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf

[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI ​​response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI ​​response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001 and the AI ​​response output device 10010. In this way, instead of performing pre-learning for each application or terminal, replicating the large-scale language model, which is the base model generated by large-scale pre-learning, and using it on individual servers and terminals allows the resources used for learning to be shared, resulting in good resource efficiency.

[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.

[0025] Furthermore, large-scale language models can be pre-trained on natural languages ​​and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. In FIG. 1A, a server having a multimodal large-scale language model is shown as large-scale language model server 20001. Specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. Note that these multimodal large-scale language models are AI models generated by large-scale pre-training on natural language contained in numerous documents and texts in the human world, as well as types of information other than natural language text information (e.g., images, videos, audio, etc.). In addition to this, there are also models that have undergone reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.

[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf

[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.

[0028] The AI ​​response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a nonvolatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, and an imaging unit 1180. The AI ​​response output device 10010 may have a large screen, such as a monitor or television.

[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by a user's finger. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire the user input that serves as the basis for a prompt to the large-scale language model, which is the artificial intelligence.

[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet (registered trademark) connection interface as hardware and communicate using a LAN-type communication method. This allows the AI ​​response output device 10010 to communicate with various servers connected to the Internet 19000.

[0031] The AI ​​response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.

[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI ​​response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.

[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller or the like, and inputs a signal regarding an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. In addition to a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may also be used by, for example, an administrator to operate the AI ​​response output device 10010. By the user's operation input via the operation input unit 1107, the AI ​​response output device 10010 can acquire a user input that serves as the basis for a command sentence (prompt) to a large-scale language model, which is an AI. Note that a modified configuration in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107 is also possible.

[0034] The video signal input unit 1131 is connected to an external video output device and inputs video data. The video signal input unit 1131 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces, etc.

[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.

[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.

[0037] The microphone 1039 is a microphone that picks up sounds around the AI ​​response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI ​​response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.

[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI ​​response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.

[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.

[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 controls video switching, such as which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also control image processing of the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling, which enlarges, reduces, or deforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.

[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI ​​response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.

[0042] The nonvolatile memory 1108 stores various data used by the AI ​​response output device 10010. The data stored in the nonvolatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI ​​response output device 10010, display icons, object data for user operations, layout information, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.

[0043] The local LLM processing unit 10028 has a memory capable of holding a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI ​​response output device 10010.

[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI ​​response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, which is a speaker or the like.

[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 of the AI ​​response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and responses are obtained from these large-scale language models. All of this control can be performed by the control unit 1110.

[0046] The storage unit 1170 may also store a fixed response phrase database (which may be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction sentences from the AI ​​response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI ​​response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI ​​response output device 10010.

[0047] Furthermore, in the example of the fixed response phrases in the database shown in FIG. 1C, if multiple fixed response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the fixed response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the fixed response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.

[0048] Next, an example of condition number 5 shown in Fig. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I might not know about that." By responding in this way, the user can be prompted to input again, and the corrected user input can be waited for.

[0049] Next, an example of condition number 6 shown in Fig. 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI ​​response output device 10010 shown in Fig. 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the fixed response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI ​​response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.

[0050] The AI ​​response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI ​​response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).

[0051] 1C described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI ​​response output device 10010. However, the response template database (response template DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI ​​response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).

[0052] In the above explanation, it has been explained that the AI ​​response output device 10010 has a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.

[0053] 1A and 1B, an example has been described in which the AI ​​response output device 10010 includes the display unit 10011. However, the AI ​​response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI ​​response output device 10010 may be configured to accept input from a user to the AI ​​via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the user input via the voice output unit 1140.

[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference of the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.

[0055] <Example 2> As a second embodiment of the present invention, a response output device and a system for outputting a response from a large-scale language model artificial intelligence will be described. The basic configuration of the second embodiment is the same as or common to the first embodiment, and the following mainly describes the components that are different from the first embodiment.

[0056] In Example 2, when there are multiple LLMs (e.g., corresponding LLM servers) each with its own LLM, the system's response output device, for example, selects an LLM to use for the answer based on the content of the user's instruction and controls the selected LLM to generate the answer. In particular, each unique LLM is an LLM that specializes in learning in a specific field / category (sometimes referred to as a specialized LLM). This specialized LLM can also be a non-general-purpose LLM, a field-specific LLM, a category-specific LLM, or a specialized LLM.

[0057] [Issues related to Example 2] Issues related to Example 2 will be explained. LLMs include general-purpose LLMs (LLMs) that have learning models trained without limiting the field / category / area, etc., and specialized LLMs (LLMs) that have learning models trained (in other words, fine-tuned) specifically for a specific field / category / area, etc. In Example 2, a general-purpose LLM and a specialized LLM are used. Assume that multiple LLMs, including these, exist as candidates that can be used to respond (answer) to a directive. In this case, it is important to determine which LLM to use to obtain a more appropriate answer. Domain specialization through fine-tuning of a specialized LLM enables improved accuracy in that domain, but can also lead to a decrease in versatility due to overfitting. Therefore, it is important to achieve both versatility and improved accuracy.

[0058] Currently, many of the response output devices using LLMs are based on general-purpose LLMs with general-purpose learning models. While general-purpose LLMs can generate answers to a variety of user instructions (questions), they have issues with the accuracy and precision of the answers. Hallucinations occur when the LLM places too much emphasis on events that it has not learned much about or on context compatibility.

[0059] On the other hand, using a specialized LLM that has been trained (fine-tuned) for a specific field / category can improve the accuracy and precision of answers in that field / category, but conversely, it falls short of a general-purpose LLM in terms of generality. With a specialized LLM, over-training can lead to a decline in performance when instructing fields / categories other than the one it was trained for.

[0060] Increasing the reliability of answers in all fields using a general-purpose LLM, which is a general-purpose model, is practically limited in terms of physical resources (in other words, time, cost, energy, etc.). Therefore, it is conceivable to create a system that combines a general-purpose model with specialized LLMs, which are models specialized for each field / category, and configures the system to select and switch the model (LLM) used for the answer according to the user's instructions. This will realize a response output system that has a good balance of both versatility and reliability.

[0061] In Example 2, the system selects an appropriate LLM from multiple candidate LLMs in response to a user instruction and uses it to answer. In other words, in Example 2, the system selects an LLM corresponding to an appropriate field / category in accordance with the user instruction and controls to switch to a response using that LLM.

[0062] Specific solutions / methods in the second embodiment include the following:

[0063] (1) The response output device analyzes the user's instruction and selects the LLM to use for the response. The system classifies the instruction based on keywords and other information in the instruction to determine which field / category it belongs to and assigns flag information. The system associates each part (e.g., word, phrase, or sentence) with a field / category or the LLM information corresponding to that field / category. A general-purpose LLM may be assigned to parts that cannot be classified. The system may also prioritize flags based on the results of keyword analysis. For example, for each part of the instruction, one or more fields / categories (or corresponding LLMs) with the highest number or ratio of keywords may be assigned, with priority assigned (see below). A maximum number of flags (corresponding categories or LLMs) may be set for each instruction or part. While this example illustrates assigning flag information based on the results of keyword analysis, the instruction may also be classified based on vector analysis, or keyword analysis and vector analysis may be used in combination.

[0064] The system selects an LLM to use based on the classified flag information in the instruction. For example, for each part of the instruction (e.g., a sentence), a specialized LLM corresponding to the category with the most flags is selected. When multiple extracted keywords in a instruction are assigned multiple flags from multiple categories, the system prioritizes the categories with the most flags, stores them, and associates them with the LLM to use. Furthermore, when a certain instruction / part is assigned multiple flags from multiple categories and cannot be identified as belonging to a single category, the system associates a generalized LLM as a multi-category item. Alternatively, if the system does not use a generalized LLM, it may simply associate multiple specialized LLMs within an upper limit with multiple flags from multiple categories and select them.

[0065] The system may present the user with information on the field / category to which the instruction applies and the corresponding LLM information indicating which LLM was used in the answer for each instruction / part (described below). By viewing the answer and the corresponding information, the user can recognize which LLM was used in the answer. The system may also allow the user to change the LLM used in the answer. If the answer and the corresponding information differ from what the user expected, the user can provide feedback on changes to the LLM used in the answer (described below).

[0066] (2) The response output device 10010 analyzes the user's instruction, weights it based on, for example, the ratio of flag information, and combines multiple LLMs to create and configure an AI model (in other words, a persona, etc.) to be used for the answer. The system sends instructions to multiple LLMs that make up the AI ​​model and receives responses from each LLM. The system then synthesizes / merges the multiple responses to create an answer to present to the user. For example, if the instruction relates to categories A, B, and C, the system combines specialized models A, B, and C corresponding to those categories with weights based on the flags (in other words, the ratios used for the answer) to create and configure an AI model. When configuring an AI model, only specialized models may be used, or general-purpose models and specialized models may be used together. Users can use an AI model configured to their liking as their own personal assistant / advisor.

[0067] (3) Based on the user's instruction, the system may select a first LLM in the first stage, obtain a primary answer, and output it. Then, depending on the user's request, select another second LLM in the second stage, obtain a secondary answer, and output it. For example, the system may select a general-purpose model in the first stage and obtain a primary answer from the general-purpose model, and then select a specialized model related to the instruction in the second stage and obtain a secondary answer from the specialized model. The system may skip the second stage if the user is satisfied with the primary answer. If the user views the primary answer and requests more detailed information, the system may perform the second stage of interaction to provide a secondary answer. Based on an analysis of the user's instruction, the system may associate and flag the primary answer with a specialized LLM corresponding to the relevant field / category. Then, if the flag is set, the system applies a processing flow to provide a secondary answer corresponding to the detailed information.

[0068] (4) The user may select the LLM to use for the response to a given instruction. The user may select an LLM that corresponds to a similar category based on the content of the instruction. The user may also select multiple LLMs to use for a given instruction. The user may also set / select an AI model that combines multiple LLMs to be used. The user may also set / select the ratio / weight of use among the multiple LLMs to be used. AI models can be set as personified characters / personas. The user can set and use an AI model tailored to the user. The user can develop an appropriate AI model through dialogue with the LLM.

[0069] In addition, for an LLM or AI model selected by the user, the system may select and present a recommended LLM / AI model from the instruction text (described below).

[0070] Furthermore, the LLM that a user can use may be turned on / off based on the user's billing information, etc. Furthermore, the LLM / AI model used by a user may be made publicly available and shared not only by a single user but also between a user and others, or among multiple people. For example, a user may upload an AI model that he or she created and provide it to others. For example, a user may download and use an AI model created by others.

[0071] [Multiple LLM servers] FIG. 2A shows an example of a system including a response output device according to a second embodiment. In the system of FIG. 2A, parallel servers are installed outside a response output device 10010, and each server has its own LLM (large-scale language model). In FIG. 2A, multiple LLM servers 19001, such as LLM server 19001A, LLM server 19001B, ..., LLM server 19001Z, are connected to the Internet 19000. For example, as their own LLMs, LLM server 19001A has LLM-A, LLM server 19001B has LLM-B, and LLM server 19001Z has LLM-Z. While the Internet 19000 is shown here, it may be replaced by a local network such as an intranet.

[0072] FIG. 2B shows an example of a system including a response output device of Example 2. The system of FIG. 2B is an example in which a parallel server is installed outside the response output device 10010 and is connected to other LLM servers via a general-purpose LLM server. In FIG. 2B, the parallel servers connected to the Internet 19000 include a general-purpose LLM server 19001G, to which other LLM servers, LLM server 19001A, LLM server 19001B, ..., LLM server 19001Z, are connected. The general-purpose LLM server 19001G has LLM-G as a general-purpose LLM. A, B, etc. are identifiers for explanatory purposes.

[0073] In Figure 2B, each of the LLMs on LLM server 19001A, LLM server 19001B, ..., LLM server 19001Z is a specialized LLM, which is an LLM specialized in a specific field, while the LLM on general LLM server 19001G is a general LLM, which is not specialized in a specific field. A specialized LLM is one that has been studied in advance with specialization in a specific field / category. Furthermore, the study of a specialized LLM is not necessarily limited to one field / category; it may include multiple fields / categories as long as it is limited to a specific field / category.

[0074] Note that one LLM server in Fig. 2A, for example, LLM server 19001A, may be general-purpose LLM server 19001G in Fig. 2B. Also, each LLM server may be a multimodal LLM server that can handle multiple types of data, such as text, images, and audio.

[0075] [LLM Classification] Figure 2C shows an example of an LLM classification in a table format when multiple LLMs are provided, each specialized for a specific field / category ("Specialized LLM"). In the example of Figure 2C, a hierarchical classification system is established, with large, medium, and small categories, such as "Major Categories" 2C01, "Medium Categories" 2C02, and "Small Categories" 2C03. Examples of "Major Categories" 2C01 include "Natural Sciences," "Social Sciences," and "Humanities." Examples of "Medium Categories" 2C02 include Science, Engineering, Agriculture, and Medicine. Examples of "Small Categories" 2C03 include Mathematics, Physics, Chemistry, and Biology. Each of these categories / categories has a corresponding specialized LLM. For example, a specialized LLM may be provided for each category in at least one of the large, medium, and small categories. Furthermore, a specialized LLM may be prepared by pre-training an LLM limited to multiple categories in each field.

[0076] Each of the multiple LLMs (LLM server 19001) in Figure 2A or Figure 2B may be an LLM (specialized LLM) specialized in a relatively large category, such as "major category" 2C01 or "medium category" 2C02. Alternatively, a relatively detailed LLM (specialized LLM) such as "minor category" 2C03 may be prepared. For example, an LLM specialized in "medicine" or "economics" may be prepared. Furthermore, the "major category" may be a general-purpose LLM, and the "medium category" and below may be specialized LLMs, etc.

[0077] Furthermore, multiple LLMs for each category as shown in Figure 2C do not have to be comprehensively installed, and only a few LLMs may be provided, for example, for a specific field, including at least one general-purpose LLM and one or more specialized LLMs.For fields / categories for which no specialized LLM exists, a general-purpose LLM (general-purpose LLM server 19001G) may be used to provide a response.

[0078] The system may also manage and store information, such as a data table, that associates classification information like that shown in Figure 2C with multiple candidate specialized LLMs that can be used (association information 2E05 in Figure 2E). When the system presents classification information like that shown in Figure 2C to the user, the user may set which level / hierarchy of classification to present.

[0079] [Summary of Several Examples] FIG. 2D is a table summarizing the outline of several examples belonging to Example 2.

[0080] In the illustrated embodiment 2A, the response output system including the response output device 10010 analyzes the user's instruction and selects one or more LLMs to use in the answer. Either one or multiple LLMs can be selected as the LLM to use in the answer, and either is possible. When multiple LLMs are selected, multiple answers from the multiple LLMs are used, for example, by combining the multiple answers to create a single answer to be presented to the user. Furthermore, the usage ratio of the multiple LLMs to be used, in other words, the weight to be reflected in the answer, may be selected and set.

[0081] The illustrated Example 2B is a variation of Example 2A, in which the user selects and specifies the LLM to be used in the answer using a predetermined GUI. The system provides a predetermined GUI for this purpose. For example, the user can select specialized LLMs (one or more) corresponding to a field / category that the user considers to be close / suitable for the instruction. The system then presents the user with an answer based on the selected specialized LLM or other LLM.

[0082] In the illustrated embodiment 2C, a response output system including a response output device 10010 analyzes a user's instruction and creates or selects an AI model (in other words, an AI persona, AI character, AI assistant, etc.) that combines multiple LLMs as the LLM to be used in the response. For example, the multiple LLMs to be used are weighted based on the ratio of keywords in each field / category in the instruction, and an AI model that combines the multiple LLMs is obtained. Furthermore, this system may pre-configure and prepare multiple such AI models and select and use them.

[0083] The illustrated Example 2D is a variation of Example 2C, in which the user selects and specifies the AI ​​model to be used for the answer using a predetermined GUI. The system provides a predetermined GUI for this purpose. The user can select an AI model that they consider to be close / suitable for the instruction. The system presents the user with an answer based on the selected AI model. The user may also configure and adjust the AI ​​model to their own preferences. The user selects multiple LLMs to use and sets the ratios (weights) to be reflected in the answers in the selected LLMs, thereby creating an AI model that combines the multiple LLMs, and saving the configuration information for the AI ​​model. The saved AI model can then be selected and used by the user.

[0084] The illustrated Examples 2E and 2F are examples and modifications of the above Examples 2A to 2D, which are related to other components and viewpoints. In Example 2E, the response output device 10010 (particularly the control unit 1110) analyzes the instruction sentence (e.g., extracts keywords), classifies / determines the field / category, and selects the LLM or AI model to be used.

[0085] In Example 2F, instead of the response output device 10010, an external LLM such as a general-purpose LLM analyzes and parses the instruction sentence, for example classifying / judging the field / category, and selecting the LLM or AI model to use.

[0086] The illustrated example 2G is, in outline, realized by a two-stage instruction (question)-answer (response) exchange in response to a user instruction. In example 2G, an answer (in other words, a primary answer) is provided using a general-purpose model in the first stage, and further, if the user instructs, an answer (in other words, a secondary answer) is provided using a specialized LLM in the second stage. The primary answer and secondary answer may be provided sequentially or in parallel.

[0087] For example, the system analyzes a user's instruction text and associates it with a specialized model to be used for the answer, and then in the first stage of interaction, obtains a primary answer using a general-purpose model and provides it to the user. If there is a specialized model associated with the instruction text, the system flags the instruction text in advance for a secondary answer using the specialized model. Based on the flagging, the system applies a flow to the user for requesting a secondary answer. The system checks whether the user requires more detailed information (i.e., a secondary answer) for the primary answer. If the system receives an input requesting more detailed information, it obtains a secondary answer using the flagged specialized model associated with the request in the second stage of interaction and provides it to the user.

[0088] Details of each of the above embodiments are provided below.

[0089] [Control Unit] FIG. 2E illustrates a configuration example of a control unit 1110 according to a second embodiment. In the control unit 1110, a prompt processing unit 2E02 creates a prompt 2E03 based on user input information 2E01 obtained through an input / output device. A prompt analysis unit 2E04 analyzes the prompt 2E03. This analysis includes keyword analysis, etc. (described later). The prompt analysis unit 2E04 performs this analysis while referencing information 2E05 that associates fields / categories with LLMs. As a result of the analysis, for example, a prompt 2E06 categorized and flagged for each part is obtained. An LLM selection unit 2E07 selects an LLM to be used in a response, for example, for each part, based on the association information 2E05 and the flag information of the prompt 2E06. The LLM selection unit 2E07 creates a prompt 2E08 for the selected LLM and sends the prompt to the selected LLM. A response processing unit 2E09 waits for a response from the selected LLM and acquires the response. The response processing unit 2E09 outputs a reply (user output information) 2E10 corresponding to the acquired response to the user.

[0090] [Analysis of instruction sentences] FIG. 2F shows an example of instruction statement analysis processing by the control unit 1110 (instruction statement analysis unit 2E04 in FIG. 2E). Instruction statement 2F01 is an example of instruction statement 2E03 (FIG. 2E) created based on user input. Here, the specific sentence content is illustrated in an abstract form. Instruction statement 2F01 has three sentences, for example, sentence 1 "AAAA BBBB ...", sentence 2 "HHHH IIII ...", and sentence 3 "OOOO PPPP ...". "AAAA" and the like are words or phrases. The control unit 1110 divides instruction statement 2F01 into parts, for example, three sentences. The control unit 1110 extracts field / category keywords and the like for each sentence, classifies the sentences into fields / categories, and assigns flags corresponding to the fields / categories.

[0091] For example, in sentence 1, the word "CCCC" is extracted as a keyword corresponding to category A. Similarly, the word "EEEE" is extracted as a keyword corresponding to category B. In sentence 2, the words "JJJJ" and "LLLL" are extracted as keywords corresponding to category C. The word "NNNN" is extracted as a keyword corresponding to category D. In sentence 3, the words "PPPP" and "QQQQ" are extracted as keywords corresponding to category E. The word "TTTT" is extracted as a keyword corresponding to category F.

[0092] The control unit 1110 associates the instruction 2F02, which is assigned a category classification flag, with the LLM to be used for the answer. The instruction 2F02 and association information 2F03 are examples of this association. For example, for Part 1 corresponding to Sentence 1, there are flags for Categories A and B, but the category cannot be identified as a single category, so the control unit 1110 associates it as using a general-purpose LLM. For Part 2 corresponding to Sentence 2, there are two flags for Category C and one flag for Category D, with Category C having the most flags. Therefore, the control unit 1110 identifies Part 2 as Category C and associates the specialized LLM-C, which is associated with Category C, as the LLM with the first priority. For Part 3 corresponding to Sentence 3, there are two flags for Category E and one flag for Category F, with Category E having the most flags. Therefore, the control unit 1110 identifies Part 3 as Category E and associates the specialized LLM-E, which is associated with Category E, as the LLM with the first priority.

[0093] The system may set an upper limit on the number of LLMs (particularly specialized LLMs) associated with each instruction or each part. For example, if the upper limit on the number of specialized LLMs to be used for each instruction is set to two, the following applies: As shown in the figure, the control unit 1110 selects a general-purpose LLM for Part 1, a specialized LLM-C with priority 1 for Part 2, and a specialized LLM-E with priority 1 for Part 3. This instruction uses two specialized LLMs. As another example, if the upper limit on the number of specialized LLMs to be used for each part of a statement is set to two, the following applies: For Part 2, the control unit 1110 selects a specialized LLM-C with priority 1, as well as a specialized LLM-D with priority 2 that corresponds to Category D. Two specialized LLMs are used in Part 2. For Part 3, the control unit 1110 selects, in addition to the specialized LLM-E with priority 1, the specialized LLM-F corresponding to Category F as the specialized LLM with priority 2. Two specialized LLMs are used in Part 2.

[0094] In addition, if a directive or part has flags for multiple categories and cannot be identified as belonging to a single category, the system may treat it as spanning multiple categories and associate it with a general-purpose LLM as the LLM to be used. Sentence 1 is such an example.

[0095] In the above example, the upper limit for the number of specialized LLMs used is set to two, but this is not limiting and the upper limit may be set according to, for example, the number of sentences or paragraphs that make up the directive. The upper limit may also be set according to the user's billing information, or according to the number of times the user uses the program, the duration of use, or the frequency of use.

[0096] Furthermore, if two responses (answers) are generated for the same sentence using multiple LLMs, for example, two LLMs, it is possible to present these two responses (answers) to the user in parallel, but it is preferable to allow the user to select only one response (answer) or to combine or merge them into one response (answer) so that it is easier for the user to understand.

[0097] [Screen example] Figure 2G shows an example of a screen used to present and output instructions and answers to a user, specifically displaying information about the category of the instruction and the LLM information selected and used to generate the answer (corresponding LLM information). This screen includes a "Category and LLM Used for Answer" column 2G03 in addition to columns for instruction 2G01 and answer 2G02. This column 2G03 displays, for example, for each part of a statement in an instruction, the corresponding category classification result information and the LLM information selected by the system as the LLM used in the answer. On this screen, the user can view and confirm the instruction 2G01, answer 2G02, and the corresponding LLM information in column 2G03.

[0098] In this example, a "Change LLM to use in answer" button 2G04 is also provided. The user can view the instruction statement 2G01, answer statement 2G02, and the corresponding LLM information in 2G column 03 on this screen, and if they do not match their own expectations, they can change the LLM to be used in the answer. When the user operates button 2G04, a GUI such as a list box is used to display options for the LLM to be used, for example, for each part, and the user can select from the options. If the user then issues another instruction (in other words, updates the answer) after the change, an answer using the changed LLM will be obtained and displayed in the answer statement 02 column.

[0099] [Screen example] Figure 2H shows another example screen in which a user selects an LLM / category to use in their answer. The screen may be configured to select a category or an LLM. When selecting a category, the LLM associated with the selected category is automatically selected. The example screen in Figure 2H includes a "Select LLM / Category to Use in Answer" field 2H03 associated with instruction 2H01. Field 2H03 includes at least one of a GUI for selecting a category to use in the answer and a GUI for selecting an LLM to use in the answer. For example, the GUI for selecting an LLM may include a list box, and the user can use a cursor or the like to select an LLM to use in their answer from the general-purpose LLM and specialized LLM options displayed.

[0100] [Screen example] Figure 2I shows another example of a screen where the system presents the user with recommended LLMs for use in their answers. The screen example in Figure 2I includes a column for instruction 2I01, a "Category and Recommended LLM" column 2I02, and a "Select LLM to Use in Answer" column 2I03. The "Category and Recommended LLM" column 2I02 displays, for each part, the system's categorization results and recommended LLMs for use in the answer, corresponding to instruction 2I01 (e.g., similar to column 2G03 in Figure 2G). The "Select LLM to Use in Answer" column 2I03 also displays a GUI that allows the user to select the LLM to use in their answer for each part, corresponding to the "Category and Recommended LLM" column 2I02. For example, for each part, a list box displays the same information as column 2I02 as the default display, and the user can use a cursor or other tool to select the LLM to use in their answer from the general-purpose LLMs and specialized LLMs displayed as options.

[0101] [Local LLM] FIG. 2J illustrates a modified example of the system in which an LLM is provided within the response output device 10010. Regarding the LLM used in each embodiment, in FIGS. 2A and 2B, an LLM server 19001 external to the response output device 10010 is used, but this is not limited to this. As shown in FIG. 2J, one or more LLMs may be provided as local LLMs within the response output device 10010 and used as appropriate. In the example of FIG. 2J, an LLM server 2J01 is provided as a local LLM that can be referenced by the control unit 1110. Examples of the LLM server 2J01 include an LLM server 2J01α that includes an LLM-α and an LLM server 2J01β that includes an LLM-β. To generate a response, the control unit 1110 can select a local LLM (LLM server 2J01) in addition to an external LLM (LLM server 19001). If the external LLM and the local LLM contain specialized LLMs that include the same field / category, the local LLM is used.

[0102] [Sequence (1): Examples 2A and 2E] 3A is a sequence diagram showing an example of processing for selecting an LLM to use based on the content of a directive (keywords, etc.) and generating a response in Example 2 (particularly Examples 2A and 2E). The right part of this sequence diagram shows the processing and operation of the control unit 1110 (FIG. 1B) of the response output device 10010, and the left part shows the processing and operation of, for example, the multiple LLM servers 19001 in FIG. 2B, such as general-purpose LLM server 19001G, LLM server 19001A (specialized LLM-A), LLM server 19001B (specialized LLM-B), ..., LLM server 19001Z (specialized LLM-Z).

[0103] In step 3A01, the control unit 1110 of the response output device 10010 generates an instruction statement (in other words, a question / request) from a user input. In step 3A02, the control unit 1110 performs a process for classifying the content of the instruction statement by field / category (see FIG. 2F above). For example, keywords for each field / category are extracted from the instruction statement. The field / category to which the instruction statement belongs is determined based on the number and ratio of keywords, etc. In step 3A03, the control unit 1110 performs a process for selecting an LLM (here, a specialized LLM) from multiple LLM servers 19001 that corresponds to the classification by category. In step 3A04, the control unit 1110 performs a process for transmitting the instruction statement to the selected LLM. In the example of FIG. 3A, the selected LLM is specialized LLM-A (LLM server 19001A).

[0104] In step 3A05, the LLM server 19001 of the selected LLM (e.g., LLM server 19001A) receives the instruction sentence and performs processing to generate a response (an LLM-generated response including natural language). In step 3A06, the LLM server 19001 performs processing to transmit the response to the response output device 10010. In step 3A07, the response output device 10010 receives the response, and the control unit 1110 performs processing corresponding to the response, such as outputting and presenting the response (in other words, an answer sentence) to the user.

[0105] When the system transmits or transfers the response of step 3A06 from the selected LLM server 19001 to the response output device 10010, it may also transmit or transfer information about the corresponding LLM (specialized LLM-A in the example of FIG. 3A), i.e., information indicating the used or selected LLM that generated the response, along with the response text. The corresponding LLM information may be, for example, information such as the name, ID, or summary of the specialized LLM-A.

[0106] In addition, it may take a certain amount of time for the response output device 10010 to obtain a response to the instruction statement from the selected LLM server 19001. The response generation process in the LLM may take a long time. Therefore, while the selected LLM is generating a response, information about the LLM used to generate the response may be displayed and presented to the user first. For example, the response output device 10010 may output information about the selected LLM (corresponding LLM information) used to generate the response to the user first using the display unit 10011 of FIG. 2B or the like. The processing in this case ("prior output") is performed at a timing such as step 3A08 shown in the figure (before step 3A07).

[0107] [Category classification processing] The categorization process of step 3A02 is, in other words, a pre-processing performed on the user's instruction sentence. This categorization process can be realized using, for example, the following techniques. The above-mentioned FIG. 2F is an example of this categorization process. This categorization process may be realized using any known technique.

[0108] FIG. 3B is an explanatory diagram summarizing, in a table format, examples of category classification processing that can be applied in this embodiment.

[0109] (Example 1) The process in Example 1 is to identify the associated category and LLM from the keywords and phrases of the input instruction based on the keywords and phrases set in advance for each LLM (corresponding field / category), and to set a flag for that LLM. The flag is flag / control information that indicates the associated field / category or the selected LLM (i.e., LLM of a certain specialized field / category) to which the instruction will be sent in response to that field / category.

[0110] (Example 2) The processing in Example 2 involves identifying the LLM with the most number of relevant keywords, etc., as the associated LLM (the LLM to be used in the answer) based on the keywords and phrases in the input instruction that have been set in advance for each LLM (corresponding field / category), and then flagging that LLM.

[0111] (Example 3) As with Example 2, the processing in Example 3 first involves identifying the single LLM with the most number of matching keywords, etc., based on keywords and phrases pre-defined for each LLM (corresponding field / category) from the keywords and phrases in the input instruction. The processing in Example 3 then involves identifying up to a pre-defined number (one or more) of LLMs, starting from the single LLM with the most number, as associated LLMs (LLMs to be used in the answer). The processing in Example 3 then involves flagging the identified LLMs.

[0112] (Example 4) In Example 4, if the keywords or phrases in the input instruction do not correspond to the keywords or phrases pre-defined for each LLM (corresponding field / category), in other words, if the instruction does not contain keywords for each LLM, the process selects a general-purpose LLM without selecting a specialized LLM and sets a flag. This flag is flag / control information indicating that the selected LLM to which the instruction is sent is a general-purpose LLM. If no general-purpose LLM exists, the specialized LLM that includes the most fields / categories may be selected and used.

[0113] (Example 5) The processing of Example 5 first includes a process of identifying, as the associated LLMs (LLMs to be used for the answer), up to a specified number of LLMs, starting from the LLM with the largest number of applicable keywords, etc., as similar to Example 3. Then, if there is no significant difference in the number of applicable keywords, etc., between the LLMs, the processing of Example 5 includes a process of selecting a general-purpose LLM and flagging the general-purpose LLM, as it is not possible to select a specific specialized LLM.

[0114] (Example 6) The process in Example 6 involves identifying the associated LLM (LLM to be used for the answer) for each part of the input instruction from the keywords / phrases of that part of the instruction based on keywords, etc., previously set for each LLM (corresponding field / category), and setting a flag for that LLM for each part. A part is a part into which the instruction is divided, such as a paragraph, sentence, or phrase.

[0115] (Example 7) The process of Example 7 includes a process of setting a flag in the LLM (LLM to be used for the answer) set by the user for the input instruction sentence.

[0116] (Example 8) The process of Example 8 includes a process of setting a flag in the LLM (LLM to be used for the answer) set by the user for each part in the input instruction sentence.

[0117] (Example 9) The processing in Example 9 includes a process of calculating the keyword relevance of an input instruction sentence with keywords, etc., set in advance for each LLM (corresponding field / category), and recording the keyword relevance information. For example, there may be cases where there are no registered keywords / phrases that match the keywords / phrases detected from the input instruction sentence. In such cases, the relevance is calculated and recorded based on the similarity between the detected keywords / phrases and the registered keywords / phrases. Even when classification is performed using vector analysis, the relevance is calculated and recorded based on the similarity with the registered vectors. Furthermore, there may be cases where a keyword / phrase is registered in multiple categories. In such cases, the relevance of each keyword category may be calculated and recorded, taking into account the categories to which other keywords / phrases belong.

[0118] (Example 10) The processing in Example 10 includes flagging multiple LLMs (LLMs to be used in the answer) set by the user for the input instruction, setting the usage ratio (ratio / weight to be reflected in the answer) for each of the multiple LLMs, and recording the usage ratio information.

[0119] (Example 11) The process of Example 11 involves a process of classifying an input instruction statement by an LLM (e.g., a general-purpose LLM) to which field / category the instruction statement or keywords therein correspond / belong. In other words, the process involves a process of querying an LLM with the instruction statement to determine which category the instruction statement corresponds to and obtaining a response.

[0120] FIG. 3G is an explanatory diagram showing an example of keyword flagging in the above-described categorization process, corresponding to Example 1. This system, including the response output device 10010, maintains a database (DB) in which keywords, etc., for each field / category are registered in advance. This DB is referred to as the keyword DB 10081. During the categorization process of step 3A02 in FIG. 3A, the control unit 1110 analyzes and extracts keywords and phrases from the user's instruction 3G01, compares the extracted keywords, etc. 3G02 with registered keywords, etc. 3G03 registered in the keyword DB 10081, and extracts and determines corresponding keywords, etc. 3G04. The registered keywords, etc. 3G03 are set in association with fields / categories. The control unit 1110 flags or quantifies each extracted corresponding keyword, etc. 3G04 to indicate which field / category it relates to. For example, a flag 3G05 indicating the field / category is assigned to each corresponding keyword, etc. 3G04. The flags shown in FIG. 2F are such flags.

[0121] Figure 3H is an explanatory diagram of a case where an instruction for category classification is queried from an LLM, corresponding to Example 11. This system including the response output device 10010 realizes the category classification process by exchanging instructions and answers with an LLM inside or outside the response output device 10010. For example, the response output device 10010 creates an instruction regarding to which field / category the user's instruction or keywords therein relate, sends it to an LLM inside or outside the response output device 10010, and obtains an answer from the LLM.

[0122] The example of FIG. 3H shows a case where an inquiry is made to an external general-purpose LLM (general-purpose LLM server 19001G). In step 3H01 of the categorization process, the control unit 1110 creates an instruction (referred to as a "category instruction") 3H02 from the user's instruction, indicating which field / category the instruction relates to, and transmits it to the general-purpose LLM server 19001G. This category instruction 3H02 may be created as a category instruction for each part (multiple category instructions) according to sentences, keywords, etc. extracted from the user's instruction. In step 3H03, the general-purpose LLM of the general-purpose LLM server 19001G generates a response (referred to as a "category response") 3H04 indicating which field / category the user's instruction corresponds to in response to the input of the category instruction 3H02. The general-purpose LLM server 19001G transmits the response 3H04 to the response output device 10010. In step 3H05, the control unit 1110 obtains from the response 3H04 the result of categorization, indicating to which field / category the user's instruction sentence falls.

[0123] [LLM Selection Process] FIG. 3C is an explanatory diagram summarizing in table form processing examples that can be applied as the LLM selection processing in step 3A03 of FIG. 3A.

[0124] (Example 1) The process in Example 1 is a process of selecting a flagged LLM. The flag of the category classification result described above (e.g., flag 3G05 in FIG. 3G) represents a classification / category and also represents an LLM associated with the classification / category based on the correspondence between the classification / category and the LLM (FIG. 2E or 2F). In this way, the LLM indicated by the flag may be selected.

[0125] (Example 2) In Example 2, a general-purpose LLM is selected if no flag is attached. As mentioned above, if the field / category indicated by the flag is not specified / cannot be specified, a general-purpose LLM may be selected.

[0126] (Example 3) In Example 3, when there are multiple flagged LLMs, all of the flagged LLMs are selected. For example, if a directive is flagged for categories A, B, and C, specialized LLMs A, B, and C corresponding to categories A, B, and C are selected.

[0127] (Example 4) In Example 4, when there are multiple flagged LLMs, a predetermined number of LLMs (such as the upper limit value mentioned above) are selected from the multiple flagged LLMs. For example, if a directive is flagged for categories A, B, and C and the upper limit value is 2, two LLMs are selected from specialized LLMs-A, B, and C corresponding to categories A, B, and C.

[0128] (Example 5) The process of Example 5 is a process of selecting one or more LLMs based on the keyword relevance information of the input instruction (Example 9 in Figure 3B). For example, the corresponding LLM is selected based on the relevance information (matching degree) between the keywords detected from the instruction and the registered keywords. Alternatively, if the keywords in the input instruction exist in multiple categories, an LLM is selected based on the keyword relevance, taking into account the category information to which the other keywords in the instruction belong.

[0129] (Example 6) The process of Example 6 is a process of selecting an LLM to be used for each part of an input instruction sentence. For example, as shown in FIG. 2F, an LLM may be selected for each sentence.

[0130] (Example 7) The process of Example 7 corresponds to Example 7 in Figure 3B, and is a process for selecting the LLM set by the user. As will be described later, the user can select and specify the LLM to be used for the answer on the screen.

[0131] (Example 8) The process of Example 8 corresponds to Example 10 in Figure 3B, and is a process of selecting an LLM based on the usage ratio information set by the user. As will be described later, the user can select and specify multiple LLMs to use in the answer and their usage ratios (weights) on the screen.

[0132] (Example 9) The processing in Example 9 corresponds to Example 8 in Figure 3B, and is processing to select the LLM set by the user for each part of the instruction statement. As will be described later, the user can select and specify the LLM to be used for each part on the screen.

[0133] (Example 10) The process in Example 10 is a process for selecting an LLM based on usage ratio information set by others (other users). A user can share multiple LLMs (AI models described below) configured with usage ratio information set by others.

[0134] (Example 11) The processing in Example 11 is processing for selecting an LLM based on usage ratio information recommended by the system including the response output device 10010. The user can use multiple LLMs (AI models described below) configured with usage ratio information recommended by the system.

[0135] [Processing example (1)] FIG. 3D shows a processing example based on FIG. 3A. In this example, the control unit 1110 or LLM of the response output device 10010 searches for keywords, etc. from the sentences input by the user (sentences constituting the question sentence / instruction sentence) and extracts the keywords, etc. This search may be performed word by word, phrase by phrase, sentence by sentence, etc. The control unit 1110 or LLM performs a classification process (step 3A02) of the instruction sentence based on the extracted keywords, etc.

[0136] In FIG. 3D, the vertical axis represents time, and the horizontal axis represents, from left to right, user instruction input, response output, display content of the user interface (GUI), and processing by the control unit 1110 or LLM. The GUI is displayed, for example, on the display screen of the display unit 10011 in FIG. 2A. A user instruction input is, for example, "Since waking up this morning, I've had a runny nose and a slight fever. What are these symptoms?" The control unit 1110 or LLM receives such an instruction and searches for keywords, etc. from the instruction. The control unit 1110 or LLM extracts parts such as "Since waking up this morning," "I have a runny nose," and "I have a slight fever" as the [explanation part] from the instruction, and extracts the part "What are these symptoms?" as the [instruction part]. The control unit 1110 or LLM extracts "this morning" as the keyword from the part "Since waking up this morning," extracts "runny nose" as the keyword from the part "I have a runny nose," and extracts "slight fever" as the keyword from the part "I have a slight fever." Furthermore, "symptoms" is extracted as a keyword from the part "What are these symptoms?" Such extracted keywords correspond to, for example, registered keywords etc. 3G03 in keyword DB10081 in FIG. 3G.

[0137] Each keyword is associated with a predetermined field / category in advance (FIG. 3G). The control unit 1110 or the LLM classifies the instruction sentence into a field / category based on the relevant keyword, etc. (3G04) and assigns a flag (3G05).

[0138] In the example of Fig. 3D, three keywords related to illness ("runny nose," "slight fever," and "symptoms") and one keyword related to time ("this morning") are matched. The control unit 1110 or LLM infers from the instruction that the most keywords related to illness are matched, and therefore, that the instruction is related to "illness."

[0139] The control unit 1110 or the LLM selects an LLM to use based on the classification of the instruction (step 3A03 in FIG. 3A). In the example of FIG. 3D, an LLM-X including medical care is selected as a specialized LLM (here, LLM-X) from the instruction related to "illness." The LLM-X has a predetermined name (e.g., "Medical-LLM"). The control unit 1110 or the LLM transmits or conveys the instruction to the selected LLM-X.

[0140] If there is no LLM available as a specialized LLM that can be associated with a command classification, such as "illness," a general-purpose LLM may be selected, as described above. Examples of such cases include when there is no physically applicable LLM, when there is an applicable LLM but the load is too high to communicate, or when the bill for using the applicable LLM is insufficient and the LLM cannot be used.

[0141] The selected LLM (LLM-X) generates a response (user response output 1) in response to the instruction and returns the response. The response output device 10010 receives the response from the used LLM and transmits / outputs the response to the user via a specified user interface. The display content of this response may be, for example, "This is a question about your symptoms. Medical-LLM will respond to you." (notification of corresponding LLM information), and the response output device 10010 receives this response (answer text). Next, the response (user response output 2) may be, for example, "Your symptoms are thought to be the early symptoms of a cold. It is recommended that you visit a hospital as soon as possible or take over-the-counter medicine and get plenty of rest. Nearby hospitals that accept reception are as follows: 1) XX Hospital, 2) XX Hospital, ..." (answer text).

[0142] The above-described categorization process may be performed as follows: If the sentence entered by the user is a compound question, classification may be performed for each question sentence. Also, if the question spans multiple parts within a single sentence, new instruction sentences may be generated to separate the sentence, and each may be classified separately.

[0143] [Sequence (2): Example 2F] Figure 3E shows the sequence of another processing example. This example corresponds to Example 2F. This example is an example in which a general-purpose LLM (general-purpose LLM server 19001G) selects an LLM (particularly a specialized LLM) to use for the answer based on the content of the user's instruction (keywords, etc.), and the selected LLM generates the answer. It also shows a case in which the answer response is sent directly (in other words, without going through the general-purpose LLM) from the selected LLM to the response output device 10010.

[0144] In step 3E01, the control unit 1110 generates an instruction statement based on user input. In step 3E02, the control unit 1110 sends the instruction statement to the general-purpose LLM server 19001G. In step 3E03, the general-purpose LLM server 19001G receives the instruction statement and performs category classification processing. The general-purpose LLM identifies the field / category to which the instruction statement is associated by inference. In step 3E04, the general-purpose LLM server 19001G selects an LLM (particularly a specialized LLM) to use for the field / category of the instruction statement. In step 3E05, the general-purpose LLM server 19001G sends the instruction statement to the selected LLM, for example, specialized LLM-A (LLM server 19001A). Note that this instruction statement is accompanied by information indicating that the question source is the response output device 10010.

[0145] In step 3E05, the selected LLM, for example, specialized LLM-A (LLM server 19001A), performs processing to generate a response to the received instruction. In step 3E06, the specialized LLM-A (LLM server 19001A) transmits the generated response to the response output device 10010, which is the source of the query, without going through the general-purpose LLM server 19001G. Note that this response may be accompanied by information about the selected LLM (for example, specialized LLM-A) that generated the response as corresponding LLM information. In step 3E07, the control unit 1110 receives the response and presents or outputs it to the user. The control unit 1110 may present or output the corresponding LLM information to the user along with the response. It may also take a certain amount of time from when the response output device 10010 transmits the instruction until it receives an answer (response). Therefore, the response output device 10010 may provide the user with information about the LLM used to generate the answer while the selected LLM is generating the response sentence. In this case, the information about the selected LLM may be sent directly from step 3E04 to step 3E07 as corresponding LLM information.

[0146] [Sequence (3): Example 2F] Figure 3F shows the sequence of another processing example. This example corresponds to Example 2F. This example is an example in which a general-purpose LLM (general-purpose LLM server 19001G) selects an LLM (particularly a specialized LLM) to use for the answer based on the content of the instruction (keywords, etc.), and the selected LLM generates the answer. Also shown is a case in which the answer response is sent from the selected LLM to the response output device 10010 via the general-purpose LLM.

[0147] Steps 3F01 to 3F05 are the same as those in FIG. 3E. The selected LLM, for example, specialized LLM-A (LLM server 19001A), transmits the response generated in step 3F05 to the general-purpose LLM server 19001G. In step 3F06, the general-purpose LLM server 19001G receives the response and performs a response preparation process. This response preparation process is a process for providing the response generated by the selected specialized LLM-A and information corresponding to the specialized LLM-A (corresponding LLM information) to the response output device 10010. In step 3F07, the general-purpose LLM server 19001G transmits the response generated by the specialized LLM-A, along with the corresponding LLM information indicating the selected LLM (specialized LLM-A) used in the response (answer), to the response output device 10010. At this time, the general-purpose LLM server 19001G may first transmit the corresponding LLM information to the response output device 10010 while waiting for a response from the specialized LLM-A. In step 3F08, the control unit 1110 receives the response generated by the specialized LLM-A and the corresponding LLM information from the general-purpose LLM server 19001G, and first presents and outputs the corresponding LLM information to the user. In step 3F09, the control unit 1110 presents and outputs the answer based on the response generated by the specialized LLM-A to the user. Note that information on the selected LLM may be transmitted directly from step 3F04 to step 3F08 as the corresponding LLM information, or the transmission and output of the corresponding LLM information as in step 3F08 may be omitted.

[0148] [Sequence (4): Example 2F: Selection of Multiple LLMs] FIG. 4A is a sequence diagram showing another processing example. FIG. 4A corresponds to Example 2F. This example shows a case where multiple LLMs are selected according to the context of the instruction statement, and multiple responses (answers) from the multiple LLMs are synthesized and presented to the user. This example also shows a case where a request and response are performed via a general-purpose LLM (general-purpose LLM server 19001G). Steps 4A01 to 4A03 are the same as those in FIG. 3F.

[0149] In step 4A04, the general-purpose LLM of the general-purpose LLM server 19001G selects multiple LLMs to use for the answer based on the category classification of the instruction. For example, assume that the category classification results of the instruction are Category A, Category B, and Category C in descending order of applicability (category indicates field / category / area, etc.). The general-purpose LLM selects a specialized LLM associated with Category A, a specialized LLM associated with Category B, and a specialized LLM associated with Category C. In this example, assume that specialized LLM-A (LLM server 19001A) and specialized LLM-B (LLM server 19001B) are selected. For example, the instruction may be divided into multiple parts based on category classification, and a specialized LLM to be used may be selected and associated with each part.

[0150] The general-purpose LLM server 19001G sends each instruction to each selected specialized LLM (e.g., specialized LLM server 19001A and specialized LLM server 19001B). For example, in steps 4A05-1, 4A05-2, ..., and 4A05-N, each specialized LLM (e.g., specialized LLM-A and specialized LLM-B) performs a process to generate a response and transmits the generated response to the general-purpose LLM server 19001G. In step 4A06, the general-purpose LLM server 19001G receives the responses from each LLM and performs a response preparation process. This response preparation process includes a process of synthesizing the responses from each specialized LLM and creating a synthesized response, i.e., a response to be provided to the response output device 10010. Synthesis can, for example, involve composing a single sentence from multiple answer sentences. A synthesis of responses may be inferred by using the responses from each specialized LLM and the information obtained in step 4A03 when identifying the fields / categories that can be associated by inference from the instruction.

[0151] In step 4A07, the general-purpose LLM server 19001G transmits the synthesized response to the response output device 10010. The synthesized response may be accompanied by corresponding LLM information indicating information on the multiple LLMs used (e.g., specialized LLM-A, specialized LLM-B). In step 4A08, the control unit 1110 receives the synthesized response and first presents and outputs the corresponding LLM information to the user. In step 4A09, the control unit 1110 presents and outputs an answer sentence based on the synthesized response to the user. The transmission and output of the corresponding LLM information as in step 4A08 may be omitted.

[0152] In addition, in this example, when presenting a synthesized response (answer sentence) in response to a user's instruction sentence, the system may present corresponding LLM information indicating which part of the answer sentence corresponds to which LLM. As an example, the synthesized response is displayed in a different color for each corresponding LLM.

[0153] [Processing example (4)] FIG. 4B shows an example of processing related to FIG. 4A. In this example, the content of the user-input instruction (designated Instruction 1) is, "Please tell me about the value enhancement for companies through SDG initiatives, including examples of environmental and market evaluation." The general-purpose LLM categorizes the instruction. For example, it may be classified into two categories: (1) the value enhancement of SDG initiatives from an environmental perspective, and (2) the value enhancement of SDG initiatives from a market evaluation perspective. The first category may be associated with, for example, "environment," and the second category may be associated with, for example, "market." The general-purpose LLM selects a first specialized LLM (e.g., Name: Environment LLM) for the first category "environment" for (1) environmental aspects, and a second specialized LLM (e.g., Name: Marketing LLM) for (2) market evaluation, corresponding to the second category "market."

[0154] The General LLM separates Instruction 1 into content corresponding to the selected Specialized LLM according to the classification. In other words, the General LLM creates instructions corresponding to the selected Specialized LLM from Instruction 1. In this example, Instruction 2 is created for the first Specialized LLM and Instruction 3 is created for the second Specialized LLM. An example of Instruction 2 would be, "Please tell me about the environmental aspects of how companies can improve their value through their efforts toward the SDGs, including examples." An example of Instruction 3 would be, "Please tell me about the market valuation aspects of how companies can improve their value through their efforts toward the SDGs, including examples."

[0155] The general-purpose LLM sends instruction 2 to the first specialized LLM and receives answer 2 as a response. The general-purpose LLM sends instruction 3 to the second specialized LLM and receives answer 3 as a response. The general-purpose LLM combines answer 2 and answer 3 to generate a single answer. The general-purpose LLM transmits the combined answer corresponding to instruction 1 to the response output device 10010 as a response. The response content might be, for example, "We will provide an answer, including examples, about the effect that efforts toward the SDGs have on improving corporate value. (1) Environmental effects: In terms of the environment, reducing carbon emissions can... (2) Market valuation effects: In terms of market valuation, the following effects can be achieved in the stock market..." The (1) environmental aspect portion of the answer is created from the response from the first specialized LLM, and the (2) market valuation portion is created from the response from the second specialized LLM.

[0156] In addition, when presenting information about the LLM used in the answer, it may be possible to output, for example, "(1) Environmental effects: (Answer based on the environmental LLM)."

[0157] In this example, Instructions 2 and 3 were created by separating Instruction 1. However, this is not limited to this. The same Instruction 1 may be sent to each specialized LLM to obtain an answer. Even if the instruction is the same, since the specialized LLMs have learned different categories, it is expected that each specialized LLM will provide a different answer according to the category. Note that if the same instruction is sent to each LLM in this way, there may be overlapping answers for the instruction in each LLM. Therefore, the general-purpose LLM performs a synthesis process to prioritize the answer of the LLM selected by the classification process (or the LLM with the highest ratio / weight) for overlapping answer parts. For example, for parts related to the category "environment," the answer from the first specialized LLM corresponding to "environment" is prioritized.

[0158] [Sequence (5): Examples 2E and 2C: Selection of Multiple LLMs] Figure 4C is a modified example of Figure 4A and corresponds to Examples 2E and 2C. This example shows a sequence in which the response output device 10010 selects multiple LLMs according to the context of the instruction statement, synthesizes multiple responses (answers) from the multiple LLMs, and presents them to the user. This example also shows a case in which an AI model (persona) is configured from multiple LLMs. This example shows a case in which a request-response is performed without going through a general-purpose LLM (general-purpose LLM server 19001G).

[0159] In step 4C01, the control unit 1110 of the response output device 10010 generates an instruction statement based on user input. In step 4C02, the control unit 1110 performs a category classification process on the instruction statement. In step 4C03, the control unit 1110 performs a process of selecting multiple LLMs (specialized LLMs) to use based on the classification of the instruction statement. The control unit 1110 creates an instruction statement for each selected LLM. In step 4C04, the control unit 1110 sends the instruction statement to each selected LLM. The instruction statement to be sent may be the same instruction statement sent to each LLM, or the instruction statement may be divided into multiple parts according to category classification and each part sent to the specialized LLM to be used.

[0160] In this example, if multiple LLMs are selected corresponding to multiple classifications through the category classification process in step 4C02, in step 4C03, the control unit 1110 selects the adoption ratio (in other words, the ratio / weight) for the multiple LLMs from the instruction text. This corresponds to the persona composition described below (FIG. 4D). The control unit 1110 creates an instruction text for each LLM according to the adoption ratio.

[0161] In steps 4C05-1 to 4C05-N, each selected LLM (LLM server 19001) performs processing to generate a response to the instruction statement. In step 4C06, each LLM transmits the generated response to the response output device 10010. The response may be accompanied by corresponding LLM information indicating the LLM used. In step 4C07, the control unit 1110 first presents and outputs the corresponding LLM information to the user. In step 4C08, the control unit 1110 combines the responses from each LLM to obtain a single answer statement. In step 4C09, the control unit 1110 presents and outputs the combined response (answer statement) to the user.

[0162] In this example, in the response synthesis process of step 4C08, the control unit 1110 synthesizes the responses (answer sentences) from each LLM based on the adoption rate of step 4C03 to create a single answer sentence by the persona. The answer sentence does not necessarily have to be created using all of the configured LLMs; an answer sentence may be created by selecting from among the configured available LLMs. This persona is an AI model configured by combining multiple selected LLMs. Furthermore, in step 4C09, the control unit 1110 may present the answer sentence to the user on a GUI screen (described below) so that it is an answer sentence by the persona. Furthermore, as described above, when presenting the synthesized response (answer), information may be presented regarding which part of the answer sentence is generated by which specialized LLM.

[0163] [Persona screen example (1)] As an example of FIG. 4C, FIG. 4D shows an example screen in which a persona (i.e., an anthropomorphized personality) is created and configured as an AI model composed of a combination of one or more selected LLMs to be used when there are multiple LLMs (e.g., specialized LLMs in FIG. 2A or 2B) that can be used for a response, and an answer is generated and output using the persona. The system creates and configures a persona corresponding to a single LLM, or a persona composed of multiple LLMs combined with assigned ratios / weights. For example, the system creates an answer by exchanging instructions and responses with the multiple LLMs that make up the persona, as shown in FIG. 4C, and synthesizing multiple responses.

[0164] The screen example of FIG. 4D shows an example in which an AI model corresponding to a persona constructed based on multiple selected LLMs is set as a personal advisor / assistant that responds to a user's instructions. A screen such as that shown in FIG. 4D is displayed as a GUI screen on the display unit 10011 of the response output device 10010. The control unit 1110 creates various GUI screens. The screen of FIG. 4D includes an image, video, icon, etc. representing the persona / advisor / assistant, as shown in the "Dedicated Advisor" column 4D02. The image of the persona may be a humanoid or character image, or may be arbitrarily set by the user. Furthermore, this GUI is not limited to display only, and may also include audio output. The persona may provide answers, guidance, etc., not only in text but also in audio.

[0165] The number and ratio of LLMs to be used as personas among multiple LLMs (for example, the general-purpose LLM, specialized LLM-A, ..., specialized LLM-Z in Figure 2B) are selected and set by the system taking into consideration the content of the instruction. The system may also use past history information to call up and use a previously set persona.

[0166] In addition, in this system, each persona configured by combining each LLM may be set in advance. The control unit 1110 may select and call the persona to be used in response to an instruction. Alternatively, the user may select the persona to be used in response to an instruction. Furthermore, the user may create and set a persona in advance by combining each LLM, save it as their own personal advisor / assistant, and call it up as desired. Furthermore, the system may display the setting information of a persona prepared in advance as a default, etc., on a GUI screen, and the user may confirm the persona on the GUI screen and adjust or update the setting information of the persona, thereby saving and using it as a personalized, customized persona. The various set personas are stored in the system's resources (e.g., the memory of the response output device 10010) and can be called up as desired for use. A user can set and use their own customized persona, which can be used as an expert / advisor / assistant.

[0167] Examples of the composition of LLMs that make up a persona include: One persona is created with 70% generic model and 30% medical model. This persona is a composite model that evokes the image of a generalist with medical experience. The medical model is an LLM specialized in the medical field, and corresponds to "Medicine" in the classification of Figure 2C. In another example, a persona is created with 80% medical model and 20% generic model. In another example, a persona is created with 20% generic model, 40% legal model, and 40% economic model.

[0168] The example screen in FIG. 4D includes a "AI for Instructions" field 4D01. The "AI for Instructions" field 4D01 includes a "Dedicated Advisor" field 4D02, a "LLM Settings Used" field 4D03, a LOAD button, a SAVE button, and the like. The "Dedicated Advisor" field 4D02 displays an image or the like representing a persona, which is a dedicated advisor. In addition to the persona image, the persona's name / ID, a description of its personality, characteristics, and the like may also be displayed. The "LLM Settings Used" field 4D03 displays one or more LLMs used that make up the persona and their usage ratios (ratios / weights). For example, as shown in the figure, the LLM usage ratios may be displayed in the form of a graph such as a radar chart. In this example, the usage ratios of five types of LLMs are displayed, but this is not limited to five types and may be increased or decreased as needed.

[0169] In the example screen of FIG. 4D, the persona corresponding to the "Dedicated Advisor" column 4D02 is set as follows: General LLM 50%, Specialized LLM-A 5%, Specialized LLM-B 10%, Specialized LLM-C 20%, and Specialized LLM-D 15%. GUIs for users to set personas include, for example, a GUI for selecting an LLM and a GUI for selecting the adoption rate of the selected LLM. The LOAD button is a GUI for calling up a saved persona. The SAVE button is a GUI for saving the persona on the screen.

[0170] 4C, the control unit 1110 sets the adoption rate (ratio / weight) of each LLM to be used based on the result of categorizing the instruction sentence input by the user. Alternatively, the control unit 1110 may select and call the persona having the most suitable adoption rate from the personas stored in the system in accordance with the result of categorizing the instruction sentence.

[0171] Furthermore, when issuing instructions and responses to each LLM constituting a persona, the adoption ratio (ratio / weight) of the LLMs used may be varied for each instruction sentence. For example, for part A of the entire instruction sentence, LLM-A and LLM-B may be used in a ratio of 7:3, and for part B of the instruction sentence, LLM-A and LLM-B may be used in a ratio of 2:8. In this case, if the answers obtained from each LLM for the same instruction sentence (part) differ between the LLMs, the control unit 1110 prioritizes and adopts the answer from the LLM with the higher adoption ratio (ratio / weight) during the synthesis process of step 4C08. Furthermore, during the synthesis process, the control unit 1110 creates or calls a persona as shown in FIG. 4D, in which the answers from each LLM are adopted according to the adoption ratio.

[0172] The settings for the adoption ratio of each LLM created by the control unit 1110 or the user of this system are saved as personas and can be recalled and used when issuing other instructions. The settings that provide the answers desired by the user can be used repeatedly as personas. This eliminates the need to set the settings again when issuing subsequent or related instructions.

[0173] Furthermore, the personas used by a user are not limited to those set by the user himself, but may also be shared and available to others (other users of the system). For example, when the LOAD button is pressed on the screen in Figure 4D, personas set by others are listed among the options, and the user can select and use one of them. The number of personas available may increase depending on the user's billing status, etc., or depending on the number of times or period of use.

[0174] The control unit 1110 also presents to the user various personas with adoption rates set by the system as candidates. The control unit 1110 may select one or more suitable personas as candidates from the personas set by the system for the user or in response to an instruction, and present them as recommendations. The user can select a persona to use from the recommendations (described below).

[0175] [Persona screen example (2)] FIG. 4E shows another example screen with a persona. The difference between FIG. 4E and FIG. 4D is that the "LLM Settings" field 4E03 has a different GUI than field 4D03. Specifically, each LLM that constitutes the persona is displayed on a separate row. The LLM information for each row 4E04 includes an ON / OFF button, an LLM (LLM name), and a percentage (ratio / weight). The ON / OFF button allows you to select whether or not to use each LLM row. The percentage value allows you to set the percentage of that LLM. The percentage value may be entered directly or via GUI operations such as up / down buttons or a bar. In the illustrated example, a certain persona is configured with 50% general-purpose LLM, 30% specialized LLM-B, and 20% specialized LLM-E, totaling 100%.

[0176] FIG. 4F shows another example of a screen in which a persona issues a response. In this example, the "AI Response" field 4F01 displays an image of the dedicated advisor, which is the persona being used, and the "LLM Used" field 4F02 displays information about multiple LLMs (e.g., category, LLM, and percentage) that make up the persona. In this example, a speech bubble GUI appears from the image of the dedicated advisor, which is the persona (e.g., a character image), and the persona's response statement 4F03 is displayed in the speech bubble GUI. The user can receive the response to their instruction as the persona's response. When the response is displayed, audio may be used in combination with the display, or audio alone may be used. Audio may be set for each dedicated advisor image, or the user may select from multiple options.

[0177] [Sequence (6): Two-stage response: Example 2G] FIG. 5A shows a sequence diagram of an embodiment corresponding to Example 2G. This example shows a case where a two-stage request-response exchange is performed using a general-purpose LLM and a specialized LLM. In the first stage, this system sends an instruction from the response output device 10010 to the general-purpose LLM to receive a primary response, and then, if necessary, in the second stage, sends an instruction from the response output device 10010 to the specialized LLM to receive a detailed response as a secondary response. Here, different LLMs are used for the primary and secondary responses. A secondary response may be available depending on the user's billing, number of uses, and usage period.

[0178] For the first stage of instruction-answer, this system generates an answer (primary answer) using a general-purpose LLM and presents and outputs this primary answer (i.e., the first answer) to the user. In parallel with this answer, for the second stage of instruction-answer, this system also generates an answer (secondary answer) using a specialized LLM corresponding to the category classification and presents and outputs this secondary answer (i.e., the second answer) to the user. While the first stage of instruction-answer and the second stage of instruction-answer are typically performed sequentially, this is not a requirement. For greater efficiency, it is advisable to process the second stage in parallel while processing the first stage. When responding to the first stage of the primary answer, this system informs the user that an additional response (secondary answer) can be requested. However, if a specialized LLM capable of providing an additional response (secondary answer) does not exist or cannot be used, the system does not inform the user that an additional response (secondary answer) can be requested. When the user sees the primary answer, confirms that an additional response (secondary answer) can be requested, and inputs a request for the additional response (secondary answer), the system presents and outputs the secondary answer to the user using the specialized LLM that was generated in advance. In addition, when requesting an additional response (secondary answer), the system may notify the user in advance of the corresponding LLM information, indicating which specialized LLM corresponds.

[0179] In FIG. 5A, in step 5A01, the control unit 1110 of the response output device 10010 generates an instruction statement (hereinafter referred to as first-stage instruction statement 1) input by the user. In step 5A02, the control unit 1110 transmits instruction statement 1 to the general-purpose LLM server 19001G. In step 5A03, the general-purpose LLM of the general-purpose LLM server 19001G performs a category classification process on the received instruction statement 1. In step 5A04, the general-purpose LLM performs a process to generate a response (hereinafter referred to as first-stage response 1, corresponding to the primary response) according to the category classification result. In step 5A05, the general-purpose LLM server 19001G transmits response 1 to the response output device 10010. Response 1 may be accompanied by corresponding LLM information indicating the general-purpose LLM that generated response 1. In step 5A06, the control unit 1110 presents and outputs the answer statement of response 1 (primary response) to the user as a process corresponding to the received response 1.

[0180] In step 5A07, the control unit 1110 performs processing to request an additional response if an additional response (corresponding to a secondary response) is possible. The availability of an additional response is determined by determining whether a specialized LLM corresponding to the category classified in step 5A03 exists or whether the corresponding specialized LLM is available. The control unit 1110 performs processing to request an additional response, for example, when a user requests an additional response in response to a user input operation on a screen. The control unit 1110 creates a second-stage instruction statement (hereinafter referred to as instruction statement 2) for the additional response. In this embodiment, at this stage, a specialized LLM to be used for the secondary response has not yet been selected. In instruction statement 2, the LLM to be used has not yet been identified or designated. In step 5A08, the control unit 1110 sends instruction statement 2 to the general-purpose LLM server 19001G. In step 5A09, the general-purpose LLM of the general-purpose LLM server 19001G performs processing to select a specialized LLM to be used for the second-stage response (secondary response) in response to the received instruction statement 2. For example, suppose that the Specialized LLM-B is selected based on the results of the previous categorization.

[0181] In this processing example, one specialized LLM is selected in the second stage, but multiple specialized LLMs to be used in the answer may be selected in the second stage, as in the above-mentioned embodiment.

[0182] The general-purpose LLM server 19001G transmits instruction 2 to the selected specialized LLM, for example, specialized LLM-B. The transmitted instruction 2 has the same content as that received in step 5A08, but the request originates from the response output device 10010 and the request destination is specialized LLM-B. Here, instruction 1 and response 1 may be transmitted in step 5A08 together with instruction 2. In step 5A10, the specialized LLM server 19001B of the selected specialized LLM-B performs processing to generate a second-stage response (response 2, corresponding to a secondary response) in response to the received instruction 2. In step 5A11, the specialized LLM server 19001B transmits response 2 to the response output device 10010. In step 5A12, the control unit 1110 presents and outputs the response text of response 2 (secondary response) to the user as processing in response to the received response 2.

[0183] 5A illustrates the first stage processing (steps 5A01 to 5A06) and the second stage processing (steps 5A07 to 5A12) as being performed sequentially, but as described above, the second stage processing may be started in parallel midway through the first stage processing. For example, before the user requests an additional response in step 5A07, the control unit 1110 causes the general-purpose LLM to perform the LLM selection process in step 5A09, generates response 2 using the selected specialized LLM, and acquires this response 2. Then, when the control unit 1110 receives a request for an additional response from the user, it presents a reply sentence using the acquired response 2.

[0184] [Processing example (6)] Figure 5B shows a processing example related to Figure 5A. This example is similar to Figure 3D up to the instruction sentence classification process. In Figure 5B, the control unit 1110 sends instruction sentence 1 to the general-purpose LLM, which performs instruction sentence keyword search and instruction sentence classification process to generate response 1. The control unit 1110 receives response 1 from the general-purpose LLM. The control unit 1110 presents the answer sentence based on response 1 to the user and also indicates that additional answers are possible if a specialized LLM corresponding to the classified category exists and is available. If a specialized LLM corresponding to the classified category does not exist and the corresponding specialized LLM is unavailable, only the answer sentence based on response 1 is presented to the user. An example of the answer sentence in the user interface is, "Your symptoms are suspected to be a cold, influenza, or novel coronavirus infection. *You can add more detailed information to obtain a specialized answer."

[0185] The display of the answer also includes a GUI for requesting an additional response, such as a "Request Additional Response" button. The GUI for requesting an additional response allows the user to input additional information (i.e., detailed information / related information). This additional information is to be passed to the LLM along with the instruction. Examples of additional information include "body temperature 37.7 degrees, sore throat" and attaching an image of the throat. In step 5A07 of FIG. 5A, the control unit 1110 uses the additional information input by the user through the GUI to create instruction 2 and transmits it to the general-purpose LLM or the selected specialized LLM. The general-purpose LLM selects the LLM to be used in step 5A09 in response to instruction 2. For example, specialized LLM-B (name: Medical-LLM), which includes medical-related fields, is selected. The general-purpose LLM transmits instruction 2 to the selected specialized LLM-B. The general-purpose LLM may also transmit instruction 1 and its response 1 together with instruction 2 to the selected specialized LLM-B.

[0186] In addition, when selecting an LLM, if there is no suitable specialized LLM corresponding to the classified category, a general-purpose LLM may be used without selecting a specialized LLM, and a second-stage response 2 may be generated using additional information from the user.

[0187] Furthermore, while waiting for response 2, the control unit 1110 first presents the corresponding LLM information to the user. For example, a message such as "Medical-LLM will respond to your questions regarding the details of your symptoms" is output on the user interface.

[0188] The selected specialized LLM-B generates response 2 in response to instruction sentence 2 and transmits response 2 to response output device 10010. Control unit 1110 receives response 2 and presents response 2 to the user. Response 2 may be displayed, for example, as follows: "Based on your symptoms of fever, sore throat, and runny nose, it is highly likely that you have influenza. You will need to be administered anti-influenza medication. It is recommended that you undergo a viral test at the nearest hospital to identify the disease."

[0189] In this example, the system first obtains a first-stage primary answer from the general-purpose LLM. If the user determines that the answer is insufficient or that an additional answer is desired, the system obtains a second-stage secondary answer from the selected specialized LLM to elicit more specialized information. If the user is satisfied with the first-stage primary answer (no request for an additional answer is entered), or if an additional answer is not possible, the second-stage request-response can be omitted.

[0190] Furthermore, if there is a specialized LLM in the field / category of the instruction, the system may indicate this in advance and ask the user whether to request an additional response / detailed response (a second-stage response using the specialized LLM). For example, at step 5A01 of FIG. 5A, this information may be displayed on the screen, allowing the user to input whether to request an additional response. If there is an input requesting an additional response, the control unit 1110 may create and transmit an additional instruction for requesting the additional response in addition to instruction 1. Alternatively, the additional instruction or information may be written in instruction 1, or the additional instruction or information may be written in metadata, etc. Alternatively, the response output device 10010 (control unit 1110) may request additional information from the user that is expected to be necessary for the additional response, and the user may input the additional information in response to the request. In this case, the control unit 1110 transmits the input additional information along with the additional instruction.

[0191] [Sequence (7): Example 2G] FIG. 5C shows a modified example of FIG. 5A, in which a response (primary response) is obtained from a general-purpose LLM in the first stage and a response (secondary response) from a specialized LLM in the second stage is obtained in advance, reducing the time lag for the user. The processing example of FIG. 5C, in other words, is a case in which the second stage processing is performed in parallel from the middle of the first stage processing. In FIG. 5C, for the first stage answer, Answer 1 is generated using a general-purpose LLM and returned to the user of the response output device 10010. In parallel with the first stage processing, the system also generates Answer 2 using a specialized LLM corresponding to the category classification as the second stage processing, and obtains Answer 2 as soon as possible. When presenting Answer 1, the system notifies the user that an additional answer can be requested. If the user further inputs a request for an additional answer, the system presents Answer 2, which has been generated and obtained in advance, to the user of the response output device 10010. This reduces the time lag experienced by the user. When requesting an additional answer, the user may be informed in advance of which specialized LLMs are supported (supported LLM information).

[0192] In Figure 5C, the first stage processing from step 5C01 to step 5C06 is the same as in Figure 5A. The general-purpose LLM of the general-purpose LLM server 19001G performs processing to generate response 1 in step 5C04, transmits response 1 in step 5C05, and performs LLM selection processing to select a specialized LLM to use based on the category classification in step 5C09. Depending on the performance of the general-purpose LLM server 19001G, the processing of step 5C09 may be performed in parallel with the processing of step 5C04, or may be processed sequentially as shown. The general-purpose LLM server 19001G transmits instruction 1 (or instruction 2 created based on instruction 1) to the selected specialized LLM, for example, specialized LLM-B. In step 5C10, the selected specialized LLM-B performs processing to generate response 2 in response to instruction 1 (or instruction 2) and transmits response 2 to the general-purpose LLM server 19001G. In step 5C11, the general-purpose LLM server 19001G receives and temporarily stores response 2.

[0193] In a modified example, response 2 may be transmitted from the specialized LLM to the response output device 10010.

[0194] Meanwhile, after processing corresponding to response 1 in step 5C06, the control unit 1110 of the response output device 10010 performs additional response request processing in step 5C07. If the user inputs a request for an additional response or detailed information, the control unit 1110 sends a corresponding additional response request to the general-purpose LLM server 19001G. If no such request is received, the process returns to step 5C01. Upon receiving the request, the general-purpose LLM server 19001G sends the temporarily stored response 2 (secondary response) together with the corresponding LLM information to the response output device 10010 in step 5C12. In step 5C13, the control unit 1110 receives response 2 and presents / outputs the corresponding LLM information and secondary response to the user.

[0195] [Sequence (8): Examples 2B and 2D: User's selection of LLM to use] FIG. 6A shows a sequence diagram of an embodiment corresponding to embodiments 2B and 2D. This example shows a process example in which the user of the response output device 10010 selects an LLM to be used for the answer based on an instruction statement created by the user. In this process example, the user selects and specifies the LLM to be used to output the answer based on the content of the instruction statement created by the user. One or more LLMs may be selected as the LLM to be used for the answer. For example, one or more may be selected from the multiple available candidate LLM servers 19001 in FIG. 1A or FIG. 1B. FIG. 6A shows an example of selecting multiple specialized LLMs.

[0196] When specifying multiple LLMs, the following methods and processing examples can also be applied.

[0197] (1) You can select and specify an LLM based on the weight (or ratio or adoption rate) and priority of each LLM among the multiple LLMs you use. The weight here refers to the weight that affects the response, and the priority refers to the priority of the LLM to be reflected in the response. For example, if the category or keyword information of a directive is registered in multiple LLMs, you can select and specify the LLM based on the weight and priority information.

[0198] (2) You may select and specify the LLM to be used for each part of the instruction.

[0199] (3) If the user does not make any selection or designation, the system (control unit 1110) may select the LLM to be used from among the corresponding LLM candidates based on the category and keyword information of the instruction statement.

[0200] (4) If the system determines that there is a more suitable / optimal alternative LLM available for the LLM selected by the user, the system may present the more suitable / optimal alternative LLM to the user as the system's recommended LLM. The system may also present the user with an answer using the recommended LLM.

[0201] The information used to generate the answer, such as the LLM selection and the weighting of the LLM, can be recorded in the system and read out for reuse or used by others. In the case of allowing others to use the information, the system can upload the information to a server, for example, so that others can access the server and use the information.

[0202] In FIG. 6A, in step 6A01, the control unit 1110 of the response output device 10010 generates an instruction statement input by the user. In step 6A02, the control unit 1110 provides a GUI screen (described below) that allows the user to select an LLM to use in response to the instruction statement, and the user selects and specifies the LLM to use as needed on the GUI screen. On the GUI screen, the user selects one or more LLMs from multiple available candidate LLMs. In this example, multiple specialized LLMs are selected. In step 6A03, the control unit 1110 transmits the instruction statement, together with information on the LLM to be used (corresponding LLM information) selected by the user, to, for example, the general-purpose LLM server 19001G.

[0203] In step 6A04, the general-purpose LLM server 19001G receives the instruction and performs pre-processing. This pre-processing involves sending the instruction to one or more selected or specified LLMs (particularly specialized LLMs). Here, the instruction to be sent may be the same instruction to each LLM, or the instruction may be divided into multiple parts based on category classification, and each part may be sent to the specialized LLM to be used. In steps 6A05-1 to 6A05-N, each specialized LLM (LLM server 19001A to 19001Z) that received the instruction performs response generation processing and sends the generated response to the general-purpose LLM server 19001G. In step 6A06, the general-purpose LLM server 19001G receives responses from each specialized LLM and performs response preparation processing. This response preparation processing is, for example, a synthesis process that combines one or more responses from specialized LLMs into a single answer. Here, the general-purpose LLM server 19001G infers a synthesis of responses based on the response sentences obtained from each specialized LLM, based on the weights and priority information set by the user. In step 6A07, the general-purpose LLM server 19001G transmits the synthesized response to the response output device 10010, along with corresponding LLM information indicating the LLM actually used in the response. In step 6A08, the control unit 1110 receives the response and presents / outputs the corresponding LLM information and the response sentence to the user.

[0204] [Screen example: Select LLM to use] FIG. 6B shows an example of an input screen provided to the user for the processing example of FIG. 6A. In step 6A02, the control unit 1110 presents a screen such as that shown in FIG. 6B. On this screen, the user can input an instruction and select and set the LLM to be used in the answer. The user can also not select an LLM to use. When selecting multiple LLMs to use when setting the LLM to use, the ratio / weight / adoption rate and priority of each LLM to be used can be selected and set.

[0205] The example screen in Figure 6B includes a "User Input Instructions" field 6B01, a "Select LLM to Use" field 6B02, and an "AI-Generated Response" field 6B03. The "User Input Instructions" field 6B01 displays instructions entered by the user. An example of such instructions might be, "I'm planning to launch an e-commerce website for the product shown in the photo. Please provide information about the market size of this product and competing products. Please also provide a sample HTML code for a page introducing this product, including the following: (1) Product Features, (2) Differences from Similar Products, (3) Important Notes, etc." The "User Input Instructions" field 6B01 also allows for text entry, as well as image, video, and audio input, with corresponding buttons provided for each. In the example above, the "Image Input" button is used to attach a photo of the product.

[0206] The "LLM Selection" column 6B02 displays information about the LLMs available as options for the LLMs, with each row displaying information that can be selected or deselected using the ON / OFF button. Additionally, the "Percentage" field allows you to set the weight, ratio, or adoption rate to be reflected in the answer for the selected LLM. For example, a general LLM might be assigned 50%, a specialized LLM-B 30%, and a specialized LLM-E 20%, totaling 100%. For example, a specialized LLM-B is an LLM in the field of economics, and a specialized LLM-E is an LLM in the field of IT.

[0207] The "AI-generated answer" field 6B03 displays the answer provided by the selected LLM to the instruction. An example of an answer might be, "The product in the photo is XXX. The annual sales of this product will be XXX billion yen in 2023. The top-selling product in 2023 is from ABC Company." See the sample code below.<!DOCTYPE html> ..."

[0208] Furthermore, if a user's question is a compound question, in other words, a question with parts relating to multiple fields / categories, it is desirable to use an LLM associated with each field / category corresponding to each part of the instruction in the answer to the question. In this case, a GUI is provided on the screen of FIG. 6B that allows the user to select an LLM to use for each instruction or for each part. For example, as a first GUI / method, the user enters a first sentence corresponding to a first category in field 6B01, taking into account the grouping of fields / categories, and selects an LLM to use in field 6B02. Next, the user enters a second sentence corresponding to a different second category in field 6B01, and selects an LLM to use in field 6B02. In this manner, the user may select the corresponding category and LLM for each sentence. As a second GUI / method, the user enters an instruction in field 6B01, and the user or control unit 1110 divides the instruction into parts, taking into account the grouping of fields / categories. For example, the instruction is divided into two sentences or paragraphs. Then, in field 6B02, the user selects the LLM to be used for each part.

[0209] [Screen example: Select persona] Figure 6C shows another example of an input screen related to the processing example of Figure 6A. This screen example shows a case where the LLM to be used for the answer is selected by specifying a persona, an AI model composed of multiple LLMs, rather than directly specifying one or more LLMs as in Figure 6B. The system pre-configures persona settings, including the LLM selected from candidate LLMs and the usage ratio of each LLM. In the example screen of Figure 6C, the "Answer AI Selection" field 6C02 displays information on multiple available candidate personas. Each of these personas generates an answer by combining the answers of the selected LLMs according to the configured LLM usage ratio. Selectable personas may also be added based on the user's billing information, etc., or based on the user's usage frequency, usage time, and usage frequency.

[0210] The system generates an answer for a persona by combining multiple answers generated by each LLM used in the persona, using a set ratio as a guideline. It is not necessary to generate an answer using all of the configured LLMs; an answer can be generated by selecting from the configured available LLMs. Alternatively, the system divides the instruction into parts (e.g., sentences), assigns the answer to the LLMs taking into account the respective ratios of each part, and then combines multiple answers to generate an answer for the persona. Furthermore, these persona settings can be read and used by others, or users can upload their own settings and allow others to use them.

[0211] In Figure 6C, the example instructions and answers are the same as in Figure 6B. In the "Answering AI Selection" column 6C02, for example, personas 6Ca to 6Cf are displayed as candidate personas. The display information for each persona includes the persona's character image, name, selection button, and so on. The user can select the persona to use in the answer by turning the button on or off. For example, persona 6Ca, whose name is "John," is selected. Detailed information for each persona can also be viewed. For example, when the user performs a predetermined operation (e.g., double-clicking) in the display area for persona 6Ca, information about the used LLM that constitutes persona 6Ca is displayed as detailed information about the persona. The display of this detailed information may be similar to column 6B02 in Figure 6B, for example. Furthermore, unavailable personas may not be displayed, or their display state may be changed, depending on the user's billing information or usage information.

[0212] [Screen example: Persona settings] FIG. 6D shows another example of an input screen related to the processing example of FIG. 6A. This example screen allows the user to set a persona, which is an AI model to be used in the answer. On this screen, the user can set the LLMs (type and name) that make up the persona and the adoption rate, etc. In another example, the rate setting may be omitted and only the used LLM may be set. When only the used LLM is set as the persona setting, the system generates the persona's answer using the used LLM of the selected persona in accordance with the instruction (a suitable LLM may be selected and used). When the persona setting includes the used LLM and adoption rate (ratio / weight), the system generates the persona's answer using the used LLM of the selected persona at the set adoption rate in accordance with the instruction.

[0213] In another example, when the system uses a persona selected by the user for a command statement, the control unit 1110 may select an appropriate LLM from among the multiple LLMs that make up the persona, which is appropriate for the command statement, to generate a response. For example, suppose a persona set by the user uses specialized LLMs-A, B, and C, which correspond to categories A, B, and C. For a command statement, the control unit 1110 determines, as a result of categorization, that categories A and B are highly relevant. In this case, the control unit 1110 selects specialized LLMs-A and B as appropriate LLMs from the composition of the user's persona, and generates a response.

[0214] Furthermore, in the persona settings shown in Figure 6D, the user may be able to set not only the LLMs to be used and their ratios, but also the response style desired when creating a response (in other words, the characteristics of the response). Examples of the response style (characteristics of the response) include the tone / voice of the response, the length of the response, and whether or not examples are provided.

[0215] In the example screen of Figure 6D, an "AI setting screen" 6D01 has an "AI character" field 6D02 and a "Select LLM to use" field 6D03. The "AI character" field 6D02 displays an image of the AI ​​model / persona / character to be set, basic settings such as name, gender, and age, and other characteristics, and these settings can be changed according to user input operations. Examples of other characteristics include "friendly tone of voice," "emphasis on short summaries in responses," and "not assertive tone for low-confidence responses."

[0216] In this way, in this embodiment, the user can tune the AI ​​model used for answering questions, and an AI assistant that is tailored to the user can be created.

[0217] [Screen example: Persona settings] FIG. 6E shows an example of a modified screen from FIG. 6D. The "Used LLM Selection" display may be displayed as a bar graph or radar chart. The "Used LLM Selection" field 6E03 in FIG. 6E is displayed as a bar graph, unlike the row display of the "Used LLM Selection" field 6D03 in FIG. 6D. In the bar graph 6E04, (1), (2), and (3) indicate the used LLMs and adoption rates, which are components of the AI ​​model, as displayed in the table below. The rates are arranged in descending order of (1), (2), and (3). The user may, for example, manipulate the table to change the used LLMs and rates. Alternatively, the user may, for example, manipulate the (1), (2), and (3) displays in the bar graph 6E04 with a cursor or the like to change the rates.

[0218] [Screen example: Select LLM to use for each part] FIG. 6F shows another example of an input screen for the processing example of FIG. 6A. In this example, a GUI is provided that allows users to select the LLM to use for each part, e.g., sentence, of a command line when inputting the command line. When inputting the command line, the user can set the LLM to use for each part, e.g., sentence, of the command line. At this time, the user can select and set which LLM to use for the answer for each part, e.g., sentence, of the command line. Furthermore, for parts of the command line not selected by the user (in other words, for parts for which no LLM was specified), the system may automatically set the system to use a general-purpose LLM or a predetermined LLM. For answers to complex commands, it is desirable to use an LLM (especially a specialized LLM) corresponding to each part.

[0219] The example screen in Figure 6F includes a "User-Input Instructions" column 6F01, a "Select LLM to Use" column 6F02, and an "AI-Generated Response" column 6F03. In the "User-Input Instructions" column 6F01, the instruction is divided into multiple parts. The system or the user divides the instruction into multiple parts. This division can be performed by a user entering a separator, or the control unit 1110 can analyze the instruction and automatically divide it into units such as sentences. For example, the first sentence might be "I'm thinking of launching an e-commerce site related to the product shown in the photo." The second sentence might be "Please tell me about the market size of this product and information about competing products." The third sentence might be "Please also provide three sample HTML code examples for a page introducing this product. The page should include the following: (1) Product Features, (2) Differences from Similar Products, (3) Notes, etc." While the divisions in this example are by sentence, they can also be by paragraphs or clauses.

[0220] In this example, the user can select the LLM to use for each delimiting sentence in the instruction sentence in the "Select LLM to Use" field 6F02. For example, a general-purpose LLM is selected for the first sentence, specialized LLM-B for the second sentence, and specialized LLM-E for the third sentence. Alternatively, to obtain a response that takes into account the content of the surrounding sentences, the surrounding sentences may be sent to the selected LLM.

[0221] An example of how a response is generated using the LLM used for each sentence is shown below. First, the response generated for the first sentence is as follows. For the attached input image corresponding to the "product in the photo" in the first sentence, a general-purpose LLM is used to infer what kind of item it is through image recognition. For example, an inference result indicating what kind of item it is is obtained. If the inference result does not indicate what category the item belongs to, the general-purpose LLM is used to classify the category. If the category is specified in advance, such as in an instruction, a specialized LLM that matches that category can be selected to generate a response.

[0222] Next, the response generated for the second sentence is as follows. From the second sentence, a specialized LLM that excels in market information for the competing product in question is selected. For example, a specialized LLM educated with a focus on marketing theory is used as specialized LLM-B. This allows proposals to be made that utilize methods and concepts for seizing marketing opportunities, such as market research, segmentation, target market selection, and analysis of customer purchasing behavior, while suppressing the interference of other information.

[0223] The response generated for the third sentence is as follows. Since the third sentence is an instruction for creating an actual web page, for example, a specialized LLM trained specifically in web design techniques is used as the specialized LLM-E. This allows us to obtain an example of a web page that can serve as a sample for the user's needs.

[0224] [Screen example: Persona settings] FIG. 6G shows a variation of the example persona setting screen of FIG. 6D. The "Select LLM to Use" field 6G03 in FIG. 6G is displayed differently from the "Select LLM to Use" field 6D03 in FIG. 6D. Specifically, in field 6G03, for each row of LLM information, those that can be used to generate answers are displayed normally, and those that cannot be used are displayed in gray. In the example shown, the available LLMs are "General LLM," "Specialized LLM-B," "Specialized LLM-D," and "Specialized LLM-E."

[0225] The LLMs available for generating responses to prompts may be varied based on information such as: (1) User billing information (2) User Contribution (3) User frequency

[0226] The user's billing information in (1) above is billing information when charging a user for using the functions / services of this system. For example, the larger the billing amount, the more LLMs that can be used. It may also be possible to purchase usage rights for each LLM.

[0227] The user's contribution level in (2) above can be determined by entering into a contract to provide the system developer with information such as the user's usage status and settings data when using the system's functions / services, and the level of contribution can be determined based on the degree of provision. The period of use in the system usage status is also one example.

[0228] The frequency of use in (3) above is a perspective regarding the LLM that the user prefers to use. Even if there are multiple candidate LLMs across the entire service, the LLMs used may be limited and biased depending on the user. This system assumes that some LLMs will be used more frequently and selects the LLM to use taking into account the user's preferences / directions.

[0229] Based on the above information for each user, the system determines which of the multiple candidate LLMs are available and which are unavailable. Unavailable LLMs are displayed as grayed out on the screen, as shown in Figure 6G, making them unselectable. In the example of Figure 6G, rows such as "Specialized LLM-A" and "Specialized LLM-C" are grayed out and unselectable depending on the user's status. This prevents the user from selecting the unavailable LLMs and forces the user to select the LLM to use from the available LLMs that are normally displayed. While the example of Figure 6G illustrates a case in which the user selects an LLM to use in the composition of a persona, this is not limiting. Similarly, when selecting an LLM to use in response to a command, as shown in Figure 6B, the display of available and unavailable LLMs can be controlled depending on the user's status.

[0230] The example in Figure 6G shows a case where the user selects an LLM to use. However, even when the system automatically selects an LLM to use, it may change the available and unavailable LLMs based on the above information (factors). Furthermore, the system may present a recommended LLM based on the instruction text for the LLM selected by the user. In this case, the system may select and present a recommended LLM from the available LLMs based on the available and unavailable LLMs based on the above information (factors).

[0231] The technology according to this embodiment makes it possible to provide a more suitable AI response output technology. It is expected that such AI response output technology will be introduced into higher quality, more reliable infrastructure. The introduction of this technology into infrastructure will contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This will contribute to the achievement of "Build resilient infrastructure, promote inclusive and sustainable industrialization, inclusive and sustainable technological development," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.

[0232] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into public transportation facilities to improve access to transportation systems for vulnerable people. The introduction of this technology into public transportation can contribute to improving traffic safety through the expansion of public transportation and realizing access to a safe, affordable, and easily usable sustainable transportation system for all people. This contributes to "Sustainable cities and communities," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.

[0233] Although various embodiments have been described above in detail, the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments are detailed descriptions of the entire system to clearly explain the present invention, and the present invention is not necessarily limited to a system including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.

[0234] The configurations, functions, processes, etc. described in the above embodiments may be partially or entirely implemented in hardware, for example, by designing an integrated circuit, a general-purpose processor, or an application-specific processor. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. Furthermore, functions, processes, etc. may be implemented in software by the processor interpreting and executing a program that implements each function. Furthermore, the scope of software implementation is not limited, and hardware and software may be used together. [Explanation of symbols]

[0235] 10010...Response output device (artificial intelligence response output device), 10011...Display unit, 1110...Control unit, 19001...LLM server.

Claims

1. A response output device, a control unit that generates an instruction sentence based on a user's input, transmits the instruction sentence to a large-scale language model (LLM), and obtains an answer sentence from the LLM as a response generated by the LLM; a display unit that displays the instruction sentence and the answer sentence; Equipped with When there are multiple LLMs available as the LLM, the control unit acquires the answer sentence as a response generated by an LLM selected from the plurality of LLMs according to a category of the instruction sentence. Response output device.

2. 2. The response output device according to claim 1, The plurality of LLMs that can be used as the LLM include specialized LLMs that are specialized and learned for specific categories. Response output device.

3. 3. The response output device according to claim 2, The LLMs that can be used include a general-purpose LLM that is not specialized in a specific category. Response output device.

4. 2. The response output device according to claim 1, the control unit analyzes the category of the instruction sentence and selects an LLM associated with the category of the analysis result as an LLM to be used in response to the instruction sentence. Response output device.

5. 4. The response output device according to claim 3, The control unit transmits the instruction to the general-purpose LLM, The general-purpose LLM analyzes a category from the instruction sentence, selects an LLM associated with the category of the analysis result as an LLM to be used for answering the instruction sentence, and transmits the instruction sentence to the selected LLM. Response output device.

6. 4. The response output device according to claim 3, The control unit transmits a request to the general-purpose LLM to analyze the category of the instruction sentence, receives a response of the analysis result of the category by the general-purpose LLM, selects an LLM associated with the category of the analysis result as an LLM to be used in response to the instruction sentence, and transmits the instruction sentence to the selected LLM. Response output device.

7. 4. The response output device according to claim 3, the control unit analyzes the category of the instruction sentence, and if the instruction sentence cannot be classified into a specific category, selects the general-purpose LLM as an LLM to be used in response to the instruction sentence. Response output device.

8. 2. The response output device according to claim 1, the control unit acquires the answer sentence as a response generated by a plurality of LLMs selected from the plurality of LLMs up to an upper limit value according to a category of the instruction sentence. Response output device.

9. 2. The response output device according to claim 1, the control unit acquires the answer sentence as a response generated by a plurality of LLMs selected from the plurality of LLMs in accordance with a priority order according to a category of the instruction sentence. Response output device.

10. 2. The response output device according to claim 1, the control unit provides a screen displaying information about the selected LLM to be used in answering the prompt or information about the LLM used in answering the prompt. Response output device.

11. 2. The response output device according to claim 1, The control unit provides a screen that displays the analysis result of the category of the instruction sentence or information on the LLM associated with the category. Response output device.

12. 2. The response output device according to claim 1, the control unit provides a screen for the user to select an LMM to use in response to the instruction statement from the plurality of LLMs. Response output device.

13. 2. The response output device according to claim 1, The control unit selects, from the plurality of LLMs according to the category of the instruction sentence, a plurality of LLMs constituting an AI model to be used in answering the instruction sentence, and acquires the answer sentence synthesized based on responses generated by the plurality of LLMs constituting the AI ​​model. Response output device.

14. The response output device according to claim 13, The control unit provides a screen for the user to select the AI ​​model. Response output device.

15. The response output device according to claim 13, The control unit sets a ratio to be used in answering the instruction sentence in a plurality of LLMs constituting the AI ​​model, Response output device.

16. 2. The response output device according to claim 1, The control unit divides the instruction into parts and selects, for each part, an LLM to be used in response to the instruction. Response output device.

17. 2. The response output device according to claim 1, the control unit acquires a primary response from a first LLM selected from the plurality of LLMs and presents it to the user, and when the user requests an additional response to the primary response, acquires a secondary response from a second LLM selected from the plurality of LLMs and presents it to the user. Response output device.

18. 18. The response output device according to claim 17, The plurality of LLMs that can be used as the LLM include a specialized LLM that has been trained to specialize in a specific category and a general-purpose LLM that is not specialized in a specific category, the control unit obtains a primary response from the general-purpose LLM as the first LLM, and obtains a secondary response from the specialized LLM as the second LLM. Response output device.

19. 19. The response output device according to claim 18, the control unit performs the process of acquiring the primary response and the process of acquiring the secondary response in parallel, and when the user requests an additional response, presents the secondary response that has already been acquired to the user. Response output device.

20. A response output system comprising a large-scale language model (LLM) and a response output device, The response output device is a control unit that generates an instruction sentence based on a user's input, transmits the instruction sentence to the LLM, and receives a reply sentence from the LLM as a response generated by the LLM; a display unit that displays the instruction sentence and the answer sentence; Equipped with When there are multiple LLMs available as the LLM, the control unit acquires the answer sentence as a response generated by an LLM selected from the plurality of LLMs according to a category of the instruction sentence. Response output system.

Citation Information

Patent Citations

  • Structural unit for tank construction

    JP1977008512A