Response output device and system

The response output device and system address inefficiencies by selecting the most suitable large-scale language model for user inputs, enhancing response output effectiveness through improved versatility and accuracy.

WO2025220418A1PCT designated stage Publication Date: 2025-10-23MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/011101
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2025-03-21
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing response output technologies using artificial intelligence do not adequately consider configurations for providing suitable responses to users, leading to inefficiencies and suboptimal performance.

Method used

A response output device and system that includes a control unit to generate and transmit instructions to a large-scale language model, and a display unit to present responses, with the ability to select an appropriate language model based on the category of the instruction, using multiple large-scale language models (LLMs) for improved response generation.

Benefits of technology

Enhances the suitability and effectiveness of response output by selecting the most appropriate LLM for user inputs, balancing versatility and accuracy in response generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025011101_23102025_PF_FP_ABST
    Figure JP2025011101_23102025_PF_FP_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to provide a more suitable artificial intelligence response output technology. The present invention contributes to the Sustainable Development Goals (SDGs) of "9. Industry, innovation, and infrastructure" and "11. Sustainable cities and communities". This response output device comprises: a control unit that generates an instruction sentence on the basis of a user input, transmits the instruction sentence to a large-scale language model (LLM), and acquires, from the LLM, an answer sentence as a response generated by the LLM; and a display unit that displays the instruction sentence and the answer sentence. When there are a plurality of LLMs that can be used as the LLM, the control unit acquires an answer sentence as a response generated by an LLM which has been selected from among the plurality of LLMs according to the category of the instruction sentence.
Need to check novelty before this filing date? Find Prior Art

Description

Response output device and system

[0001] The present invention relates to a response output device and system.

[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1.

[0003] Special table 2019-528512 publication

[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.

[0005] An object of the present invention is to provide a more suitable response output technique.

[0006] To solve the above problem, for example, the configuration described in the claims is adopted. The present application includes multiple means for solving the above problem, and an example thereof may be configured as follows: A response output device, comprising: a control unit that generates an instruction sentence based on a user input, transmits the instruction sentence to a large-scale language model (LLM), and acquires an answer sentence from the LLM as a response generated by the LLM; and a display unit that displays the instruction sentence and the answer sentence, wherein if there are multiple LLMs that can be used as the LLM, the control unit acquires the answer sentence as a response generated by an LLM selected from the multiple LLMs according to a category of the instruction sentence.

[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments.

[0008] FIG. 1 is a diagram illustrating an example of an AI response output device and system according to an embodiment of the present invention. FIG. 1 is a diagram illustrating an example of an AI response output device and system according to an embodiment of the present invention. FIG. 2 is a diagram illustrating an example of an operation of an AI response output device and system according to an embodiment of the present invention. FIG. 3 is a diagram illustrating an example of a response output device and system according to an embodiment of the present invention. FIG. 4 is a diagram illustrating an example of field / category classification according to an embodiment. FIG. 5 is a diagram illustrating an overview of each embodiment. FIG. 6 is a diagram illustrating an example of processing by a control unit according to an embodiment. FIG. 7 is a diagram illustrating an example of processing for analyzing a directive sentence according to an embodiment. FIG. 8 is a diagram illustrating an example of presenting information on an LLM to be used in an answer according to an embodiment. FIG. 9 is a diagram illustrating an example of changing an LLM to be used in an answer according to an embodiment. FIG. 10 is a diagram illustrating an example of recommending an LLM to be used in an answer according to an embodiment. FIG. 11 is a diagram illustrating an example of a response output device and system according to an embodiment. FIG. 12 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. FIG. 13 is a diagram illustrating an example of a category classification process according to an embodiment. FIG. 14 is a diagram illustrating an example of an LLM selection process according to an embodiment. FIG. 15 is a diagram illustrating a specific processing example according to an embodiment. FIG. 16 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. FIG. 17 is a diagram illustrating an example of a sequence of a response output system according to an embodiment. FIG. 1 is a diagram showing a specific processing example according to an embodiment; FIG. 2 is a diagram showing an example sequence of a response output system according to an embodiment; FIG. 3 is a diagram showing an example screen according to an embodiment; FIG. 4 is a diagram showing an example screen according to an embodiment; FIG. 5 is a diagram showing an example sequence of a response output system according to an embodiment; FIG. 6 is a diagram showing a specific processing example according to an embodiment; FIG. 7 is a diagram showing an example sequence of a response output system according to an embodiment; FIG. 8 is a diagram showing an example sequence of a response output system according to an embodiment; FIG. 9 is a diagram showing an example screen according to an embodiment; FIG. 10 is a diagram showing an example screen according to an embodiment; FIG. 11 is a diagram showing an example screen according to an embodiment;

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.

[0010] Note that if the AI ​​response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI ​​response output device has an audio output function, it may be referred to as an audio output device. The AI ​​response output device may simply be referred to as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI ​​response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI ​​response output device or the display output of the AI ​​response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI ​​response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI ​​response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.

[0011] First Embodiment As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.

[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI ​​response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system including the AI ​​response output device 10010 and the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.

[0013] In the example of FIG. 1A, the AI ​​response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. The display unit 10011 may also be provided with a touch operation input sensor and configured as a touch panel.

[0014] 1A, the audio output unit 1140 of the AI ​​response output device 10010 is composed of a speaker. The AI ​​response output device 10010 also has a microphone 1139 that can pick up the user's voice. By audio input from the microphone 1139 or user operation input via an operation input unit (described later), the AI ​​response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.

[0015] The AI ​​response output device 10010 may be provided with a local large-scale language model within the AI ​​response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or as an audio output from the audio output unit 1140.

[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.

[0017] Alternatively, the AI ​​response output device 10010 may also include a local large-scale language model and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI ​​response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.

[0018] The configuration when the AI ​​response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A , the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include wired and wireless portions, or may go via a router or repeater. The AI ​​response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI ​​response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI ​​response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.

[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI ​​response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.

[0020] The example of Figure 1A shows an example in which the display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of Figure 1A, the prompt display area 10051 displays an icon 10052 indicating a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of Figure 1A, the artificial intelligence response display area 10061 displays an icon 10062 indicating an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI ​​response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI ​​response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.

[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment as well. Note that these large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called a Transformer. Reference 1, for example, has been published as an example of learning these models.

[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf

[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI ​​response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI ​​response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001, the AI ​​response output device 10010, etc. In this way, instead of performing pre-learning for each application or each terminal, replicating the large-scale language model that is the base model generated by large-scale pre-learning and using it on individual servers or terminals allows the resources used for learning to be shared, resulting in good resource efficiency.

[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.

[0025] Furthermore, large-scale language models can be pre-trained on natural languages ​​and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. In FIG. 1A , a server having a multimodal large-scale language model is shown as large-scale language model server 20001. For example, specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are AI models generated by large-scale pre-training on natural language and types of information other than natural language text information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. In addition, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.

[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf

[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.

[0028] The AI ​​response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, etc. The AI ​​response output device 10010 may have a large screen, such as a monitor or television.

[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be referred to as a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by a user's finger. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is artificial intelligence.

[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet (registered trademark) connection interface as hardware and communicate using a LAN communication method. This allows the AI ​​response output device 10010 to communicate with various servers connected to the Internet 19000.

[0031] The AI ​​response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.

[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI ​​response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.

[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller, and inputs a signal for an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. Separate from a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the AI ​​response output device 10010. By the user's operation input via the operation input unit 1107, the AI ​​response output device 10010 can acquire user input that serves as the basis for instructions (prompts) for the large-scale language model, which is the AI. A modified configuration is also possible in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107.

[0034] The video signal input unit 1131 connects to an external video output device and inputs video data. The video signal input unit 1131 may be configured with various digital video input interfaces. For example, the video signal input unit 1131 may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces.

[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces, etc. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.

[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.

[0037] The microphone 1039 is a microphone that picks up sounds around the AI ​​response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI ​​response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.

[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI ​​response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.

[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor device memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.

[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 performs, for example, video switching control, such as determining which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform control to perform image processing on the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling processing, which enlarges, reduces, or deforms an image; brightness adjustment processing, which changes the brightness; contrast adjustment processing, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.

[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI ​​response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.

[0042] The nonvolatile memory 1108 stores various data used by the AI ​​response output device 10010. The data stored in the nonvolatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI ​​response output device 10010, display icons, object data for user operations, layout information, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.

[0043] The local LLM processing unit 10028 includes a memory capable of storing a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI ​​response output device 10010.

[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI ​​response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028 or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, such as a speaker.

[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 provided in the AI ​​response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. All of these controls for obtaining a response from these large-scale language models can be performed by the control unit 1110.

[0046] The storage unit 1170 may also store a fixed response phrase database (which may also be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction statements from the AI ​​response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI ​​response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI ​​response output device 10010.

[0047] Furthermore, in the example of the standard response phrases in the database shown in FIG. 1C, if multiple standard response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the standard response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the standard response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.

[0048] Next, an example of condition number 5 shown in FIG. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains an obvious grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch what you said" or "I might not know about that." By responding in this manner, the user can be prompted to input again, and the system can wait for a corrected user input.

[0049] Next, an example of condition number 6 shown in Figure 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI ​​response output device 10010 shown in Figure 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the standard response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI ​​response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.

[0050] The AI ​​response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI ​​response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, the AI ​​response output device 10010 may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).

[0051] 1C described above may be stored in the storage unit 1170 and used by the control unit 1110 of the AI ​​response output device 10010. However, the fixed response database (fixed response DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response database (fixed response DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the fixed response database (fixed response DB) to the AI ​​response output device 10010, instead of a response generated using a large-scale language model stored in the respective server. In this way, even if the artificial intelligence response output device 10010 is not equipped with a standard response phrase database (standard response phrase DB), it is possible to generate a response using the standard response phrase database (standard response phrase DB).

[0052] In the above description, the AI ​​response output device 10010 has been described as having a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.

[0053] 1A and 1B, an example has been described in which the AI ​​response output device 10010 includes the display unit 10011. However, the AI ​​response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI ​​response output device 10010 may be configured to accept input from a user to the AI ​​via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the user input via the voice output unit 1140.

[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference by the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.

[0055] <Example 2> As Example 2 of the present invention, a response output device and system for outputting a response from a large-scale language model artificial intelligence will be described. The basic configuration of Example 2 is the same as that of Example 1, and the following mainly describes the components that are different from Example 1.

[0056] In Example 2, when there are multiple LLMs (e.g., corresponding LLM servers) each with its own LLM, the system, for example, controls the response output device to select an LLM to use in a response based on the content of a user's instruction and have the selected LLM generate a response. In particular, each unique LLM is an LLM specialized in learning a specific field / category (sometimes referred to as a specialized LLM). This specialized LLM can be, in other words, a non-general LLM, a field-specific LLM, a category-specific LLM, or a specialized LLM.

[0057] Issues Related to Example 2 Issues related to Example 2 will be described. LLMs include LLMs (general-purpose LLMs) that have a learning model trained without limiting the field / category / area, etc., and LLMs (specialized LLMs) that have a learning model trained (in other words, fine-tuned) specifically for a specific field / category / area, etc. In Example 2, a general-purpose LLM and a specialized LLM are used. Assume that multiple LLMs, including these, exist as candidates that can be used to respond (answer) to a directive. In this case, it is important to determine which LLM to use to obtain a more suitable answer. Domain specialization through fine-tuning of a specialized LLM enables improved accuracy in that domain, but can also lead to a decrease in versatility due to overfitting. Therefore, it is important to achieve both versatility and improved accuracy.

[0058] Currently, many of the response output devices using LLMs are based on general-purpose LLMs with general-purpose learning models. However, while general-purpose LLMs can generate answers to a variety of user instructions (questions), they have issues with the accuracy and precision of the answers. Hallucinations occur when the LLM places emphasis on events that it has not learned much about or on context compatibility.

[0059] On the other hand, when using a specialized LLM that has been trained (fine-tuned) for a specific field / category, the accuracy and precision of answers in that field / category can be improved, but conversely, the generality of the LLM is inferior. In a specialized LLM, over-training can result in a decrease in performance when instructing fields / categories other than the one it was trained for.

[0060] Increasing the reliability of answers in all fields using a general-purpose LLM, which is a general-purpose model, is practically limited in terms of physical resources (in other words, time, cost, energy, etc.). Therefore, it is conceivable to create a system that combines a general-purpose model with specialized LLMs, which are specialized models for each field / category, and configures the model (LLM) used for the answer to be selected and switched according to the user's instructions. This realizes a response output system that balances both versatility and reliability.

[0061] In Example 2, the system selects a suitable LLM from multiple candidate LLMs in response to a user instruction and uses it in a response. In other words, in Example 2, the system selects an LLM corresponding to a suitable field / category in accordance with the user instruction and controls to switch to a response using that LLM.

[0062] Specific solutions / methods in the second embodiment include the following:

[0063] (1) The response output device analyzes the user's instruction and selects the LLM to use for the response. The system classifies the instruction based on keywords and other information in the instruction to determine which field / category it belongs to and assigns flag information. The system associates each part (e.g., word, phrase, or sentence) with a field / category or an LLM corresponding to that field / category. A general-purpose LLM may be assigned to parts that cannot be classified. The system may also prioritize flags based on the results of keyword analysis. For example, for each part of the instruction, one or more fields / categories (or corresponding LLMs) with the highest number or ratio of keywords may be assigned, with priority assigned (as described below). A maximum number of flags (corresponding categories or LLMs) may be set for each instruction or part. While this example describes assigning flag information based on the results of keyword analysis, the instruction may also be classified based on vector analysis to determine which field / category it belongs to, or keyword analysis and vector analysis may be used in combination to perform classification.

[0064] The system selects an LLM to use based on the classified flag information in the instruction. For example, for each part of the instruction (e.g., a sentence), a specialized LLM corresponding to the category with the most flags is selected. When multiple extracted keywords in a instruction are assigned multiple flags from multiple categories, the system prioritizes the categories with the most flags, stores them, and associates the LLM to use. Furthermore, when a certain instruction / part is assigned multiple flags from multiple categories and cannot be identified as belonging to a single category, the system associates a generalized LLM as belonging to multiple categories. Alternatively, if the system does not use a generalized LLM, it may simply associate multiple specialized LLMs within the upper limit with multiple flags from multiple categories.

[0065] The system may present the user with information on the field / category to which the instruction applies and corresponding LLM information indicating which LLM was used in the answer for each instruction / part (as described below). By viewing the answer and the corresponding information, the user can recognize which LLM was used in the answer. The system may also allow the user to change the LLM used in the answer. If the answer and the corresponding information differ from the user's intended answer, the user can provide feedback on changing the LLM used in the answer (as described below).

[0066] (2) The response output device 10010 analyzes the user's instruction, weights it based on, for example, the ratio of flag information, and combines multiple LLMs to create and configure an AI model (in other words, a persona, etc.) to be used in the response. The system sends instructions to multiple LLMs that make up the AI ​​model and obtains responses from each LLM. The system then synthesizes / merges the multiple responses to create an answer to present to the user. For example, if the instruction relates to categories A, B, and C, the system combines specialized models A, B, and C corresponding to those categories with weights based on the flags (in other words, the ratios used in the answer) to create and configure an AI model. When configuring an AI model, only specialized models may be used, or general-purpose models and specialized models may be used together. Users can use an AI model tailored to their needs as their own personal assistant / advisor.

[0067] (3) Based on the user's instruction, the system may select a first LLM in the first stage to obtain and output a primary answer, and then, depending on the user's request, select another second LLM in the second stage to obtain and output a secondary answer. For example, the system may select a general model in the first stage to obtain a primary answer from the general model, and then select a specialized model related to the instruction in the second stage to obtain a secondary answer from the specialized model. The system may skip the second stage if the user is satisfied with the primary answer. If the user views the primary answer and requests more detailed information, the system may perform the second stage of interaction to provide a secondary answer. Based on an analysis of the user's instruction, the system may associate and flag the primary answer with a specialized LLM corresponding to the relevant field / category. Then, if the flag is set, the system applies a processing flow to provide a secondary answer corresponding to the detailed information.

[0068] (4) The user may select the LLM to use for the response to a given instruction. The user may select an LLM that corresponds to a similar category based on the content of the instruction. The user may also select multiple LLMs to use for a given instruction. The user may also set / select an AI model that combines multiple LLMs to use. The user may also set / select the ratio / weight of use among the multiple LLMs to use. The AI ​​model can be set as a personified character / persona. The user can set and use an AI model tailored to the user. The user can develop an appropriate AI model through dialogue with the LLM.

[0069] In addition, for an LLM or AI model selected by a user, the system may select and present a recommended LLM / AI model from the instruction text (described below).

[0070] Furthermore, the LLM available to a user may be turned on / off based on the user's billing information, etc. Furthermore, the LLM / AI model used by a user may be made available or shared not only by a single user but also between a user and others, or among multiple people. For example, a user may upload an AI model that he or she created and provide it to others. For example, a user may download and use an AI model created by others.

[0071] [Multiple LLM Servers] FIG. 2A shows an example of a system including a response output device according to the second embodiment. In the system of FIG. 2A, parallel servers are installed outside the response output device 10010, and each server has its own LLM (Large Scale Language Model). In FIG. 2A, multiple LLM servers 19001, such as LLM server 19001A, LLM server 19001B, ..., LLM server 19001Z, are connected to the Internet 19000. For example, as their own LLMs, LLM server 19001A has LLM-A, LLM server 19001B has LLM-B, and LLM server 19001Z has LLM-Z. While the Internet 19000 is shown here, it may be replaced by a local network such as an intranet.

[0072] FIG. 2B shows an example of a system including a response output device according to the second embodiment. The system in FIG. 2B is an example in which a parallel server is installed outside the response output device 10010 and is connected to other LLM servers via a general-purpose LLM server. In FIG. 2B, the parallel servers connected to the Internet 19000 include a general-purpose LLM server 19001G, to which other LLM servers, LLM server 19001A, LLM server 19001B, ..., LLM server 19001Z, are connected. The general-purpose LLM server 19001G has LLM-G as a general-purpose LLM. A, B, etc. are identifiers for explanatory purposes.

[0073] In Figure 2B, the LLMs on LLM server 19001A, LLM server 19001B, ..., LLM server 19001Z are specialized LLMs, which are LLMs specialized in a specific field, while the LLMs on general LLM server 19001G are general LLMs that are not specialized in a specific field. Specialized LLMs are those that have been studied in advance with specialization in a specific field / category. Furthermore, the study of specialized LLMs is not necessarily limited to one field / category; it may include multiple fields / categories as long as it is limited to a specific field / category.

[0074] Note that one LLM server in Fig. 2A, for example, LLM server 19001A, may be general-purpose LLM server 19001G in Fig. 2B. Also, each LLM server may be a multimodal LLM server that can handle multiple types of data, such as text, images, and audio.

[0075] [LLM Classification] FIG. 2C shows an example of LLM classification in a table format when multiple LLMs are provided, each with an LLM specialized for each field / category ("Specialized LLM"). In the example of FIG. 2C, a hierarchical classification is provided, with large, medium, and small classifications such as "Major Classification" 2C01, "Medium Classification" 2C02, and "Small Classification" 2C03. Examples of "Major Classification" 2C01 include "Natural Science," "Social Science," and "Humanities." Examples of "Medium Classification" 2C02 include Science, Engineering, Agriculture, and Medicine. Examples of "Small Classification" 2C03 include Mathematics, Physics, Chemistry, and Biology. Each of these classifications / categories has an associated specialized LLM. For example, a specialized LLM may be provided for each category of at least one of the large, medium, and small classifications. Furthermore, pre-trained LLMs may be limited to multiple categories in each field and used as specialized LLMs.

[0076] As each of the multiple LLMs (LLM server 19001) in FIG. 2A or 2B, an LLM (specialized LLM) specialized in a relatively large category, such as "major category" 2C01 or "medium category" 2C02, may be prepared. Alternatively, a relatively detailed LLM (specialized LLM) such as "minor category" 2C03 may be prepared. For example, an LLM specialized in "medicine" or an LLM specialized in "economics" may be prepared. Furthermore, the "major category" may be a general-purpose LLM, and the "medium category" and below may be specialized LLMs.

[0077] 2C , multiple LLMs for each category do not necessarily have to be provided comprehensively. Instead, only a specific field may be provided, with at least one general LLM and one or more specialized LLMs, for example, only a few LLMs in total. For fields / categories for which no specialized LLM exists, a general LLM (general LLM server 19001G) may be used to provide a response.

[0078] The system may also manage and store information, such as a data table, that associates classification information like that shown in Figure 2C with multiple candidate specialized LLMs that can be used (association information 2E05 in Figure 2E). When the system presents classification information like that shown in Figure 2C to the user, the user may set which level / hierarchy of classification to present.

[0079] [Summary of Several Examples] FIG. 2D is a table summarizing summaries of several examples belonging to Example 2.

[0080] In the illustrated embodiment 2A, the response output system including the response output device 10010 analyzes the user's instruction and selects one or more LLMs to use in the answer. Either a single LLM or multiple LLMs can be selected as the LLM to use in the answer. When multiple LLMs are selected, multiple answers from the multiple LLMs are used, for example, by combining the multiple answers to create a single answer to be presented to the user. The usage ratio of the multiple LLMs to be used, in other words, the weight to be reflected in the answer, may also be selected and set.

[0081] The illustrated example 2B is a variation of example 2A, in which a user selects and specifies an LLM to be used in a response using a predetermined GUI. The system provides a predetermined GUI for this purpose. For example, the user can select one or more specialized LLMs corresponding to a field / category that the user considers to be close / suitable for the instruction. The system then presents the user with an answer based on the selected specialized LLM or other LLMs.

[0082] In the illustrated example 2C, a response output system including a response output device 10010 analyzes a user's instruction and creates or selects an AI model (in other words, an AI persona, AI character, AI assistant, etc.) that combines multiple LLMs as the LLM to be used in the response. For example, the multiple LLMs to be used are weighted based on the ratio of keywords in each field / category in the instruction, and an AI model that combines the multiple LLMs is obtained. Furthermore, this system may pre-configure and prepare multiple such AI models and select and use them.

[0083] The illustrated Example 2D is a modified example of Example 2C, in which the user selects and specifies the AI ​​model to be used for the answer using a predetermined GUI. The system provides a predetermined GUI for this purpose. The user can select an AI model that they consider to be close / suitable for the instruction. The system presents the user with an answer based on the selected AI model. The user may also set and adjust the AI ​​model to their own preferences. The user selects multiple LLMs to use and sets the ratios (weights) to be reflected in the answers in the selected multiple LLMs, thereby creating an AI model that combines the multiple LLMs, and saving the setting information for the AI ​​model. The saved AI model can then be selected and used by the user.

[0084] The illustrated Examples 2E and 2F are examples and modifications of the above Examples 2A to 2D, which are related to other components and viewpoints. In Example 2E, the response output device 10010 (particularly the control unit 1110) analyzes the instruction sentence (e.g., extracts keywords), classifies / determines the field / category, and selects the LLM or AI model to be used.

[0085] In Example 2F, instead of the response output device 10010, an external LLM such as a general-purpose LLM analyzes and interprets the instruction sentence, for example classifying / determining the field / category, and selecting the LLM or AI model to use.

[0086] The illustrated example 2G is, in outline, realized by a two-stage instruction (question)-answer (response) exchange in response to a user instruction. In example 2G, an answer (i.e., a primary answer) is provided using a general-purpose model in the first stage, and if the user requests, an answer (i.e., a secondary answer) is provided using a specialized LLM in the second stage. The primary and secondary answers may be provided sequentially or in parallel.

[0087] For example, the system analyzes a user's instruction text and associates it with a specialized model to be used for the answer, and then in a first stage of interaction, obtains a primary answer using a general model and provides it to the user. If a specialized model is associated with the instruction text, the system flags the instruction text in advance for a secondary answer using the specialized model. Based on the flagging, the system applies a flow to the user for requesting a secondary answer. The system checks whether the user requires more detailed information (i.e., a secondary answer) for the primary answer. If the system receives an input requesting more detailed information, in a second stage of interaction, the system obtains a secondary answer using the flagged specialized model associated with the request and provides it to the user.

[0088] Details of each of the above embodiments are provided below.

[0089] [Control Unit] FIG. 2E illustrates an exemplary configuration of a control unit 1110 according to the second embodiment. In the control unit 1110, a prompt processing unit 2E02 creates a prompt 2E03 based on user input information 2E01 obtained through an input / output device. A prompt analysis unit 2E04 analyzes the prompt 2E03. This analysis includes keyword analysis, etc. (described below). The prompt analysis unit 2E04 performs this analysis while referencing information 2E05 that associates fields / categories with LLMs. As a result of the analysis, a prompt 2E06 is obtained that is categorized and flagged for each part, for example. An LLM selection unit 2E07 selects an LLM to be used in a response, for example, for each part, based on the association information 2E05 and the flag information of the prompt 2E06. The LLM selection unit 2E07 creates a prompt 2E08 for the selected LLM and transmits the prompt to the selected LLM. The response processing unit 2E09 waits for a response from the selected LLM, acquires the response, and outputs an answer (user output information) 2E10 corresponding to the acquired response to the user.

[0090] [Instruction Statement Analysis Processing] FIG. 2F shows an example of instruction statement analysis processing by the control unit 1110 (instruction statement analysis unit 2E04 in FIG. 2E). Instruction statement 2F01 is an example of instruction statement 2E03 (FIG. 2E) created based on user input. Here, the specific sentence content is illustrated in an abstract form. Instruction statement 2F01 has three sentences, for example, sentence 1 "AAAA BBBB ...", sentence 2 "HHHH III ...", and sentence 3 "OOOO PPPP ...". "AAAA" and the like are words or phrases. The control unit 1110 divides instruction statement 2F01 into parts, for example, three sentences. The control unit 1110 extracts field / category keywords and the like for each sentence, classifies the fields / categories, and assigns flags corresponding to the fields / categories.

[0091] For example, in sentence 1, the word "CCCC" is extracted as a keyword corresponding to category A. Similarly, the word "EEEE" is extracted as a keyword corresponding to category B. In sentence 2, the words "JJJJ" and "LLLL" are extracted as keywords corresponding to category C. The word "NNNN" is extracted as a keyword corresponding to category D. In sentence 3, the words "PPPP" and "QQQQ" are extracted as keywords corresponding to category E. The word "TTTT" is extracted as a keyword corresponding to category F.

[0092] The control unit 1110 associates the instruction sentence 2F02, to which a category classification flag has been assigned, with the LLM to be used for the answer. The instruction sentence and association information 2F03 are an example of such association. For example, for Part 1 corresponding to Sentence 1, there are flags for Categories A and B, but the category cannot be identified as a single category, so the control unit 1110 associates it as using a general-purpose LLM. For Part 2 corresponding to Sentence 2, there are two flags for Category C and one flag for Category D, with Category C having the most flags. Therefore, the control unit 1110 identifies Part 2 as Category C and associates it with the specialized LLM-C, which is associated with Category C, as the LLM with the first priority. For Part 3 corresponding to Sentence 3, there are two flags for Category E and one flag for Category F, with Category E having the most flags. Therefore, the control unit 1110 specifies that Part 3 is in Category E, and associates the specialized LLM-E associated with Category E as the LLM with the first priority.

[0093] The system may set an upper limit on the number of LLMs (particularly specialized LLMs) associated with each instruction or each part. For example, if the upper limit on specialized LLMs to be used for each instruction is set to two, the result is as follows. As shown in the figure, the control unit 1110 selects a general LLM for Part 1, a specialized LLM-C with priority 1 for Part 2, and a specialized LLM-E with priority 1 for Part 3. This instruction uses two specialized LLMs. As another example, if the upper limit on specialized LLMs to be used for each part of a sentence is set to two, the result is as follows. For Part 2, in addition to the specialized LLM-C with priority 1, the control unit 1110 selects a specialized LLM-D with priority 2, which corresponds to Category D. Two specialized LLMs are used in Part 2. For Part 3, the control unit 1110 selects, in addition to the specialized LLM-E with priority 1, specialized LLM-F corresponding to Category F as a specialized LLM with priority 2. Two specialized LLMs are used in Part 2.

[0094] In addition, if a directive or part has multiple category flags and cannot be identified as belonging to a single category, the system may treat it as spanning multiple categories and associate it with a general-purpose LLM as the LLM to be used. Sentence 1 is such an example.

[0095] In the above example, the upper limit for the number of specialized LLMs used is set to two, but this is not limiting and the upper limit may be set according to, for example, the number of sentences or paragraphs that make up the instruction. The upper limit may also be set according to the user's billing information, or according to the number of times, duration, or frequency of use by the user.

[0096] Furthermore, if two responses (answers) are generated for the same sentence using multiple LLMs, for example, two LLMs, it is possible to present the two responses (answers) to the user in parallel, but it is preferable to select only one response (answer) or to combine or merge them into one response (answer) so that it is easier for the user to understand.

[0097] [Screen Example] Figure 2G shows an example of a screen displayed when presenting and outputting instructions and answers to a user, particularly displaying information about the category of the instruction and information about the LLM selected and used to generate the answer (corresponding LLM information). This screen includes a "Category and LLM Used in Answer" column 2G03 in addition to columns for instruction 2G01 and answer 2G02. In column 2G03, category classification result information and information about the LLM selected by the system as the LLM used in the answer are displayed in association with each other for each part of the instruction, for example. On this screen, the user can view and confirm the instruction 2G01, answer 2G02, and the corresponding LLM information in column 2G03.

[0098] In this example, a "Change LLM to be used in answer" button 2G04 is also provided. The user can view the instruction statement 2G01, answer statement 2G02, and the corresponding LLM information in 2G column 03 on this screen, and if they do not match their own expectations, they can change the LLM to be used in the answer. When the user operates button 2G04, a GUI such as a list box is used to display options for the LLM to be used, for example, for each part, and the user can select from the options. If the user then issues another instruction (i.e., updates the answer) after the change, an answer using the changed LLM is obtained and displayed in the answer statement 02 column.

[0099] [Screen Example] Figure 2H shows another example screen in which a user selects an LLM / category to be used in an answer. The screen may be configured to select a category or an LLM. When a category is selected, an LLM associated with the selected category is automatically selected. The screen example of Figure 2H includes a "Select LLM / Category to Use in Answer" field 2H03 associated with instruction 2H01. Field 2H03 includes at least one of a GUI for selecting a category to be used in an answer and a GUI for selecting an LLM to be used in an answer. For example, the GUI for selecting an LLM may include a list box, and the user can use a cursor or the like to select an LLM to be used in an answer from the general LLMs and specialized LLMs displayed as options.

[0100] [Screen Example] Figure 2I is another example of a screen in which the system presents the user with recommended LLMs to use in their answers. The screen example in Figure 2I includes a column for instruction statement 2I01, a "Category and Recommended LLM" column 2I02, and a "Select LLM to Use in Answer" column 2I03. The "Category and Recommended LLM" column 2I02 displays, for each part, information on the system's category classification results and information on the LLM recommended for use in the answer, corresponding to instruction statement 2I01 (e.g., similar to column 2G03 in Figure 2G). The "Select LLM to Use in Answer" column 2I03 also displays a GUI that allows the user to select an LLM to use in their answer for each part, corresponding to the "Category and Recommended LLM" column 2I02. For example, for each part, a list box displays the same information as column 2I02 as the default display, and the user can use a cursor or other device to select an LLM to use in their answer from the general-purpose LLMs and specialized LLMs displayed as options.

[0101] [Local LLM] Figure 2J shows a modified example of the system in which an LLM is provided within the response output device 10010. Regarding the LLM used in each embodiment, Figures 2A and 2B illustrate the use of an LLM server 19001 external to the response output device 10010, but this is not limited to this. As shown in Figure 2J, one or more LLMs may be provided as local LLMs within the response output device 10010 and used as appropriate. In the example of Figure 2J, an LLM server 2J01 is provided as a local LLM that can be referenced by the control unit 1110. Examples of the LLM server 2J01 include an LLM server 2J01α that includes LLM-α and an LLM server 2J01β that includes LLM-β. The control unit 1110 can select a local LLM (LLM server 2J01) in addition to an external LLM (LLM server 19001) to generate a response. If the external LLM and the local LLM have a specialized LLM that includes the same field / category, the local LLM will be used.

[0102] [Sequence (1): Examples 2A and 2E] Figure 3A is a sequence diagram showing an example of a process for generating an answer by selecting an LLM to use based on the content of a directive (keywords, etc.) in Example 2 (particularly Examples 2A and 2E). The right side of this sequence diagram shows the processing and operation of the control unit 1110 (Figure 1B) of the response output device 10010, and the left side shows the processing and operation of, for example, the multiple LLM servers 19001 in Figure 2B, including a general-purpose LLM server 19001G, an LLM server 19001A (specialized LLM-A), an LLM server 19001B (specialized LLM-B), ..., and an LLM server 19001Z (specialized LLM-Z).

[0103] In step 3A01, the control unit 1110 of the response output device 10010 generates an instruction (in other words, a question / request) from the user input. In step 3A02, the control unit 1110 performs a process for classifying the content of the instruction by field / category (see FIG. 2F above). For example, keywords for each field / category are extracted from the instruction. The field / category to which the instruction belongs is determined based on the number and ratio of keywords. In step 3A03, the control unit 1110 performs a process for selecting an LLM (here, a specialized LLM) corresponding to the classification by category from multiple LLM servers 19001. In step 3A04, the control unit 1110 performs a process for transmitting the instruction to the selected LLM. In the example of FIG. 3A, the selected LLM is specialized LLM-A (LLM server 19001A).

[0104] In step 3A05, the LLM server 19001 of the selected LLM (e.g., LLM server 19001A) receives the instruction sentence and performs processing to generate a response (an LLM-generated response including natural language). In step 3A06, the LLM server 19001 performs processing to transmit the response to the response output device 10010. In step 3A07, the response output device 10010 receives the response, and the control unit 1110 performs processing corresponding to the response, such as outputting and presenting the response (in other words, an answer sentence) to the user.

[0105] When the system transmits the response of step 3A06 from the selected LLM server 19001 to the response output device 10010, it may transmit information about the corresponding LLM (specialized LLM-A in the example of FIG. 3A ), i.e., information indicating the used / selected LLM that generated the response, along with the response text. The corresponding LLM information may be, for example, the name, ID, or summary of the specialized LLM-A.

[0106] In addition, it may take a certain amount of time for the response output device 10010 to obtain a response to the instruction statement from the selected LLM server 19001. The response generation process in the LLM may take a long time. Therefore, while the selected LLM is generating a response, information about the LLM used to generate the response may be displayed and presented to the user first. For example, the response output device 10010 may output information about the selected LLM (corresponding LLM information) used to generate the response to the user first using the display unit 10011 of FIG. 2B . In this case ("prior output"), the process is performed at a timing such as step 3A08 shown in the figure (before step 3A07).

[0107] [Category Classification Process] The categorization process of step 3A02 is, in other words, a pre-processing performed on the user's instruction sentence. This categorization process can be realized using, for example, the following technology. The above-mentioned FIG. 2F is an example of this categorization process. This categorization process may be realized using any known technology.

[0108] FIG. 3B is an explanatory diagram summarizing, in a table format, examples of category classification processing that can be applied in this embodiment.

[0109] Example 1: The process of Example 1 involves identifying a corresponding category and LLM from keywords and phrases in an input instruction based on keywords and phrases predefined for each LLM (corresponding field / category), and then flagging the LLM. The flag is flag / control information that indicates the corresponding field / category or the selected LLM (i.e., an LLM in a specialized field / category) to which the instruction is to be sent in accordance with the field / category.

[0110] (Example 2) The processing in Example 2 is a process in which, based on the keywords and phrases in the input instruction sentence that are set in advance for each LLM (corresponding field / category), one LLM with the most number of relevant keywords, etc. is identified as the associated LLM (LLM to be used for the answer) and a flag is set for that LLM.

[0111] (Example 3) As with Example 2, the process of Example 3 first involves identifying the LLM with the largest number of keywords and phrases from the keywords and phrases of the input instruction, based on keywords and phrases previously set for each LLM (corresponding field / category). The process of Example 3 then involves identifying up to a predefined number (one or more) of LLMs, starting from the most numerous LLM, as associated LLMs (LLMs to be used in the answer). The process of Example 3 then involves flagging the identified LLMs.

[0112] (Example 4) In Example 4, if the keywords or phrases in the input instruction do not correspond to the keywords or phrases predefined for each LLM (corresponding field / category), in other words, if the instruction does not include keywords for the LLM, a general-purpose LLM is selected without selecting a specialized LLM and a flag is set. The flag is flag / control information indicating that the selected LLM to which the instruction is sent is a general-purpose LLM. Furthermore, if a general-purpose LLM does not exist, a specialized LLM with the most fields / categories may be selected and used.

[0113] (Example 5) The process of Example 5 first includes a process of identifying, as associated LLMs (LLMs to be used for the answer), up to a specified number of LLMs, starting from the LLM with the largest number of applicable keywords, as in Example 3. Then, if there is no significant difference in the number of applicable keywords, etc., between the LLMs, the process of Example 5 includes a process of selecting a general-purpose LLM and flagging the general-purpose LLM, as it is not possible to select a specific specialized LLM.

[0114] (Example 6) The process of Example 6 involves identifying, for each part of an input instruction, an associated LLM (an LLM to be used in the answer) based on keywords or phrases in the part of the instruction, and flagging the LLM for each part, based on keywords or the like previously set for each LLM (corresponding field / category). A part is a division of the instruction, such as a paragraph, sentence, or phrase.

[0115] Example 7 The process of Example 7 includes a process of setting a flag in the LLM (LLM to be used for the answer) set by the user for the input instruction sentence.

[0116] Example 8 The process of Example 8 includes a process of setting a flag in the LLM (LLM to be used for the answer) set by the user for each part in the input instruction sentence.

[0117] (Example 9) The processing of Example 9 includes a process of calculating the keyword relevance of an input instruction sentence with keywords, etc., previously set for each LLM (corresponding field / category), and recording the keyword relevance information. For example, there may be cases where there are no registered keywords / phrases that match the keywords / phrases detected from the input instruction sentence. In such cases, the relevance is calculated and recorded based on the similarity between the detected keywords / phrases and the registered keywords / phrases. Even when classification is performed using vector analysis, the relevance is calculated and recorded based on the similarity with the registered vectors. Furthermore, there may be cases where a keyword / phrase is registered in multiple categories. In such cases, the relevance of each keyword category may be calculated and recorded, taking into account the categories to which other keywords / phrases belong.

[0118] (Example 10) The processing of Example 10 includes flagging multiple LLMs (LLMs to be used in the answer) set by the user for the input instruction sentence, setting the usage ratio (ratio / weight to be reflected in the answer) for each of the multiple LLMs, and recording the usage ratio information.

[0119] Example 11: The process of Example 11 involves using an LLM (e.g., a general-purpose LLM) to classify an input instruction statement to determine to which field / category the instruction statement or keywords therein correspond / belong. In other words, the process involves querying an LLM with the instruction statement to determine which category the instruction statement corresponds to and obtaining a response.

[0120] FIG. 3G is an explanatory diagram showing an example of keyword flagging in the above-described categorization process, corresponding to Example 1. The system, including the response output device 10010, maintains a database (DB) in which keywords, etc., for each field / category are registered in advance. This DB is referred to as the keyword DB 10081. During the categorization process of step 3A02 in FIG. 3A, the control unit 1110 analyzes and extracts keywords and phrases from the user's instruction 3G01, compares the extracted keywords, etc. 3G02 with registered keywords, etc. 3G03 registered in the keyword DB 10081, and extracts and determines corresponding keywords, etc. 3G04. The registered keywords, etc. 3G03 are set in association with fields / categories. The control unit 1110 flags or quantifies the field / category to which each extracted corresponding keyword, etc. 3G04 relates. For example, a flag 3G05 indicating the field / category is assigned to each corresponding keyword, etc. 3G04. The flags shown in FIG. 2F are such flags.

[0121] 3H is an explanatory diagram of a case where an instruction for category classification is queried from an LLM, corresponding to Example 11. This system including the response output device 10010 realizes the category classification process by exchanging instructions and answers with an LLM inside or outside the response output device 10010. For example, the response output device 10010 creates an instruction regarding to which field / category the user's instruction or keywords therein relate, sends it to an LLM inside or outside the response output device 10010, and obtains a response from the LLM.

[0122] The example of FIG. 3H illustrates a case where an inquiry is made to an external general-purpose LLM (general-purpose LLM server 19001G). In step 3H01 of the categorization process, the control unit 1110 creates an instruction (referred to as a "category instruction") 3H02 from the user's instruction, indicating which field / category the instruction relates to, and transmits the instruction to the general-purpose LLM server 19001G. This category instruction 3H02 may be created as a category instruction for each part (multiple category instructions) based on sentences, keywords, etc. extracted from the user's instruction. In step 3H03, the general-purpose LLM of the general-purpose LLM server 19001G generates a response (referred to as a "category response") 3H04 indicating which field / category the user's instruction corresponds to in response to the input of the category instruction 3H02. The general-purpose LLM server 19001G transmits the response 3H04 to the response output device 10010. In step 3H05, the control unit 1110 obtains, from the response 3H04, the result of categorization indicating to which field / category the user's instruction sentence falls.

[0123] [LLM Selection Process] FIG. 3C is an explanatory diagram summarizing in table form examples of processes that can be applied as the LLM selection process in step 3A03 of FIG. 3A.

[0124] (Example 1) The process of Example 1 is a process of selecting a flagged LLM. The flag of the above-mentioned category classification result (e.g., flag 3G05 in FIG. 3G) represents a classification / category and also represents an LLM associated with the classification / category based on the correspondence between the classification / category and the LLM (FIGS. 2E and 2F). In this way, the LLM indicated by the flag may be selected.

[0125] (Example 2) In Example 2, a general-purpose LLM is selected when no flag is attached. As described above, if the field / category indicated by the flag is not specified / cannot be specified, a general-purpose LLM may be selected.

[0126] (Example 3) In the process of Example 3, when there are multiple flagged LLMs, all of the flagged LLMs are selected. For example, if a directive is flagged for categories A, B, and C, specialized LLMs A, B, and C corresponding to categories A, B, and C are selected.

[0127] (Example 4) In the process of Example 4, when there are multiple flagged LLMs, a predetermined number of LLMs (e.g., the upper limit value mentioned above) are selected from the multiple flagged LLMs. For example, if a directive is flagged for categories A, B, and C and the upper limit value is 2, two LLMs are selected from specialized LLMs-A, B, and C corresponding to categories A, B, and C.

[0128] (Example 5) The process of Example 5 is a process of selecting one or more LLMs based on keyword relevance information of an input instruction (Example 9 in FIG. 3B ). For example, a corresponding LLM is selected based on relevance information (matching degree) between keywords detected from the instruction and registered keywords. Alternatively, if the keywords of the input instruction exist in multiple categories, an LLM is selected based on keyword relevance, taking into account category information to which other keywords in the instruction belong.

[0129] Example 6 The process of Example 6 is a process of selecting an LLM to be used for each part of an input instruction sentence. For example, as shown in FIG. 2F, an LLM may be selected for each sentence.

[0130] (Example 7) The process of Example 7 corresponds to Example 7 in Fig. 3B and is a process for selecting an LLM set by the user. As will be described later, the user can select and specify the LLM to be used in the answer on the screen.

[0131] (Example 8) The process of Example 8 corresponds to Example 10 in Fig. 3B and is a process of selecting an LLM based on usage ratio information set by the user. As will be described later, the user can select and specify multiple LLMs to use in an answer and their usage ratios (weights) on the screen.

[0132] (Example 9) The process of Example 9 corresponds to Example 8 in Fig. 3B, and is a process of selecting an LLM set by the user for each part of a directive. As will be described later, the user can select and specify the LLM to be used for each part on the screen.

[0133] (Example 10) The process of Example 10 is a process of selecting an LLM based on usage ratio information set by another user (another user). A user can share multiple LLMs (AI models described below) configured with usage ratio information set by another user.

[0134] (Example 11) The process of Example 11 is a process of selecting an LLM based on usage ratio information recommended by the present system including the response output device 10010. The user can use multiple LLMs (AI models described below) configured with usage ratio information recommended by the present system.

[0135] [Processing Example (1)] Figure 3D shows a processing example based on Figure 3A. In this example, the control unit 1110 or LLM of the response output device 10010 searches for keywords, etc. from the sentences input by the user (sentences constituting the question / instruction sentence) and extracts the keywords, etc. This search may be performed word by word, phrase by phrase, sentence by sentence, etc. The control unit 1110 or LLM performs a classification process (step 3A02) of the instruction sentence based on the extracted keywords, etc.

[0136] In FIG. 3D , the vertical axis represents the time axis, and the horizontal axis represents, from left to right, user instruction input and response output, the display content of the user interface (GUI), and processing by the control unit 1110 or LLM. The GUI is displayed, for example, on the display screen of the display unit 10011 in FIG. 2A . For example, a user instruction input is, "Since waking up this morning, I've had a runny nose and a slight fever. What are these symptoms?" The control unit 1110 or LLM receives such an instruction and searches for keywords, etc. from the instruction. The control unit 1110 or LLM extracts parts such as "Since waking up this morning," "I've had a runny nose," and "I have a slight fever" as the [explanation portion] from the instruction, and extracts the part "What are these symptoms?" as the [instruction portion]. The control unit 1110 or LLM extracts the keyword "this morning" from the part "Since I woke up this morning," the keyword "runny nose" from the part "I have a runny nose," and the keyword "slight fever" from the part "I have a slight fever." The control unit 1110 or LLM also extracts the keyword "symptoms" from the part "What are these symptoms?" These extracted keywords correspond to, for example, registered keywords, etc. 3G03 in keyword DB 10081 in FIG. 3G.

[0137] Each keyword is associated with a predetermined field / category in advance (FIG. 3G). The control unit 1110 or the LLM classifies the instruction sentence into a field / category based on the relevant keyword, etc. (3G04) and assigns a flag (3G05).

[0138] 3D, three keywords related to illness ("runny nose," "slight fever," and "symptoms") and one keyword related to time ("this morning") are matched. The control unit 1110 or the LLM infers from the instruction that the most keywords related to illness are matched, and therefore, that the instruction is related to "illness."

[0139] The control unit 1110 or the LLM selects an LLM to use based on the classification of the instruction (step 3A03 in FIG. 3A). In the example of FIG. 3D, an LLM-X that includes medical care is selected as a specialized LLM (here, LLM-X) from the instruction related to "illness." The LLM-X has a predetermined name (e.g., "Medical-LLM"). The control unit 1110 or the LLM transmits and conveys the instruction to the selected LLM-X.

[0140] If there is no LLM available as a specialized LLM that can be associated with a command category, such as "illness," a general-purpose LLM may be selected, as described above. Examples of such cases include when there is no physically applicable LLM, when there is an applicable LLM but the load is too high to communicate, or when the bill for the applicable LLM is insufficient and the LLM cannot be used.

[0141] The selected LLM (LLM-X) generates a response (user response output 1) in response to the instruction and returns the response. The response output device 10010 receives the response from the used LLM and transmits and outputs the response to the user via a specified user interface. The display content of this response might be, for example, "This is a question about your symptoms. Medical-LLM will respond to you." (Notification of corresponding LLM information), and the response output device 10010 receives this response (answer text). The next response (user response output 2) might be, for example, "Your symptoms are likely the early symptoms of a cold. It is recommended that you see a doctor as soon as possible or take over-the-counter medicine and get plenty of rest. Nearby hospitals with reception are as follows: 1) XX Hospital, 2) XX Hospital, ..." (response text).

[0142] The above-described categorization process may be performed as follows: If the sentence entered by the user is a compound question, classification may be performed for each question sentence. Also, if the question spans multiple parts within a single sentence, new instruction sentences may be generated to separate the sentence, and each may be classified separately.

[0143] [Sequence (2): Example 2F] Figure 3E shows the sequence of another processing example. This example corresponds to Example 2F. In this example, a general-purpose LLM (general-purpose LLM server 19001G) selects an LLM (particularly a specialized LLM) to use for the answer based on the content of the user's instruction (keywords, etc.), and the selected LLM generates the answer. This example also shows a case in which the answer response is sent directly (in other words, without going through the general-purpose LLM) from the selected LLM to the response output device 10010.

[0144] In step 3E01, the control unit 1110 generates an instruction statement based on user input. In step 3E02, the control unit 1110 transmits the instruction statement to the general-purpose LLM server 19001G. In step 3E03, the general-purpose LLM server 19001G receives the instruction statement and performs category classification processing. The general-purpose LLM identifies the field / category to which the instruction statement is associated by inference. In step 3E04, the general-purpose LLM server 19001G selects an LLM (particularly a specialized LLM) to use for the field / category of the instruction statement. In step 3E05, the general-purpose LLM server 19001G transmits the instruction statement to the selected LLM, for example, specialized LLM-A (LLM server 19001A). Note that this instruction statement includes information indicating that the question source is the response output device 10010.

[0145] In step 3E05, the selected LLM, for example, specialized LLM-A (LLM server 19001A), performs processing to generate a response to the received instruction. In step 3E06, the specialized LLM-A (LLM server 19001A) transmits the generated response to the response output device 10010, which is the source of the query, without going through the general-purpose LLM server 19001G. Note that this response may include information about the selected LLM (for example, specialized LLM-A) that generated the response as corresponding LLM information. In step 3E07, the control unit 1110 receives the response and presents or outputs it to the user. The control unit 1110 may present or output corresponding LLM information to the user along with the response. It may also take a certain amount of time for the response output device 10010 to receive an answer (response) after transmitting the instruction. Therefore, the response output device 10010 may provide the user with information about the LLM used to generate the answer while the selected LLM is generating the response sentence. In this case, the information about the selected LLM may be sent directly from step 3E04 to step 3E07 as corresponding LLM information.

[0146] [Sequence (3): Example 2F] Figure 3F shows the sequence of another processing example. This example corresponds to Example 2F. In this example, a general-purpose LLM (general-purpose LLM server 19001G) selects an LLM (particularly a specialized LLM) to use for the answer based on the content of the instruction (keywords, etc.), and the selected LLM generates the answer. Also shown is a case where the selected LLM transmits a response to the answer to the response output device 10010 via the general-purpose LLM.

[0147] Steps 3F01 to 3F05 are the same as those shown in FIG. 3E. The selected LLM, for example, specialized LLM-A (LLM server 19001A), transmits the response generated in step 3F05 to the general-purpose LLM server 19001G. In step 3F06, the general-purpose LLM server 19001G receives the response and performs a response preparation process. This response preparation process is a process for providing the response generated by the selected specialized LLM-A and information corresponding to the specialized LLM-A (corresponding LLM information) to the response output device 10010. In step 3F07, the general-purpose LLM server 19001G transmits the response generated by the specialized LLM-A and the corresponding LLM information associated with it, which represents the selected LLM (specialized LLM-A) used in the response (answer), to the response output device 10010. In this case, the general-purpose LLM server 19001G may first transmit the corresponding LLM information to the response output device 10010 while waiting for a response from the specialized LLM-A. In step 3F08, the control unit 1110 receives the response generated by the specialized LLM-A and the corresponding LLM information from the general-purpose LLM server 19001G, and first presents and outputs the corresponding LLM information to the user. In step 3F09, the control unit 1110 presents and outputs the answer based on the response generated by the specialized LLM-A to the user. Note that the information on the selected LLM may be transmitted directly from step 3F04 to step 3F08 as the corresponding LLM information, or the transmission and output of the corresponding LLM information as in step 3F08 may be omitted.

[0148] [Sequence (4): Example 2F: Selection of Multiple LLMs] Figure 4A is a sequence diagram showing another processing example. Figure 4A corresponds to Example 2F. This example shows a case where multiple LLMs are selected according to the context of the instruction statement, and multiple responses (answers) from the multiple LLMs are synthesized and presented to the user. This example also shows a case where a request and response are performed via a general-purpose LLM (general-purpose LLM server 19001G). Steps 4A01 to 4A03 are the same as those in Figure 3F.

[0149] In step 4A04, the general-purpose LLM of the general-purpose LLM server 19001G selects multiple LLMs to use for the answer based on the category classification of the instruction. For example, assume that the category classification of the instruction is Classification A, Classification B, and Classification C in descending order of applicability (Classification indicates a field / category / area, etc.). The general-purpose LLM selects a specialized LLM associated with Classification A, a specialized LLM associated with Classification B, and a specialized LLM associated with Classification C. In this example, assume that specialized LLM-A (LLM server 19001A) and specialized LLM-B (LLM server 19001B) are selected. For example, the instruction may be divided into multiple parts by category classification, and a specialized LLM to be used may be selected and associated with each part.

[0150] The general-purpose LLM server 19001G transmits each instruction sentence to each selected specialized LLM (e.g., specialized LLM server 19001A and specialized LLM server 19001B). For example, in steps 4A05-1, 4A05-2, ..., and 4A05-N, each specialized LLM (e.g., specialized LLM-A and specialized LLM-B) generates a response and transmits the generated response to the general-purpose LLM server 19001G. In step 4A06, the general-purpose LLM server 19001G receives the responses from each LLM and performs a response preparation process. This response preparation process includes a process of synthesizing the responses from each specialized LLM to create a synthesized response, i.e., a response to be provided to the response output device 10010. Synthesis can involve, for example, constructing a single sentence from multiple answer sentences. A synthesis of responses may be inferred by using the responses from each specialized LLM and the information obtained in step 4A03 when the corresponding field / category was identified by inference from the instruction sentence.

[0151] In step 4A07, general-purpose LLM server 19001G transmits the synthesized response to response output device 10010. The synthesized response may include corresponding LLM information representing information on the multiple LLMs used (e.g., specialized LLM-A, specialized LLM-B). In step 4A08, control unit 1110 receives the synthesized response and first presents and outputs the corresponding LLM information to the user. In step 4A09, control unit 1110 presents and outputs the answer sentence based on the synthesized response to the user. The transmission and output of the corresponding LLM information as in step 4A08 may be omitted.

[0152] In addition, in this example, when presenting a synthesized response (answer sentence) in response to a user's instruction sentence, the system may present corresponding LLM information indicating which part of the answer sentence corresponds to which LLM. As an example, the synthesized response may be displayed in a different color for each corresponding LLM.

[0153] [Processing Example (4)] Figure 4B illustrates a processing example related to Figure 4A. In this example, the content of the user-input instruction (designated instruction 1) is "Please tell me about the value improvement in companies through SDG initiatives, including examples of environmental and market evaluation." The general-purpose LLM categorizes the instruction. For example, the instruction is classified into two categories: (1) the value improvement of SDG initiatives from an environmental perspective, and (2) the value improvement of SDG initiatives from a market evaluation perspective. The first category is associated with, for example, "environment," and the second category is associated with, for example, "market." The general-purpose LLM selects a first specialized LLM (e.g., "Environment LLM") corresponding to the first category "Environment" for (1) the environmental aspect, and selects a second specialized LLM (e.g., "Marketing LLM") corresponding to the second category "Market" for (2) the market evaluation.

[0154] The general LLM separates Instruction 1 into content corresponding to the selected specialized LLM according to the classification. In other words, the general LLM creates instructions corresponding to the selected specialized LLM from Instruction 1. In this example, Instruction 2 for the first specialized LLM and Instruction 3 for the second specialized LLM are created. Instruction 2, for example, could be, "Please tell me about the environmental aspects of how companies can improve their value through their efforts toward the SDGs, including examples." Instruction 3, for example, could be, "Please tell me about the market valuation aspects of how companies can improve their value through their efforts toward the SDGs, including examples."

[0155] The general-purpose LLM sends instruction statement 2 to the first specialized LLM and receives answer statement 2 as a response. The general-purpose LLM sends instruction statement 3 to the second specialized LLM and receives answer statement 3 as a response. The general-purpose LLM combines answer statements 2 and 3 to generate a single answer statement. The general-purpose LLM transmits the combined answer statement corresponding to instruction statement 1 to the response output device 10010 as a response. The content of the response might be, for example, "We will provide an answer, including examples, regarding the effect that efforts toward the SDGs have on improving corporate value. (1) Environmental effects: In terms of the environment, reducing carbon emissions can... (2) Market valuation effects: In terms of market valuation, the following effects can be achieved in the stock market..." The (1) environmental aspect portion of the answer statement is created from the response from the first specialized LLM, and the (2) market valuation portion is created from the response from the second specialized LLM.

[0156] Furthermore, when presenting information about the LLM used in the answer sentence, for example, "(1) Environmental effects: (answer based on environmental LLM)" may be output.

[0157] In this example, directives 2 and 3 were created by separating directive 1. However, this is not limiting; the same directive 1 may be sent to each specialized LLM to obtain a response. Even if the same directive is used, because the specialized LLMs learned different categories, it is expected that each specialized LLM will provide a different response according to the category. Note that sending the same directive to each LLM in this manner may result in overlapping responses for the directive in each LLM. Therefore, the general-purpose LLM performs a synthesis process to prioritize the response of the LLM selected by the classification process (or the LLM with the highest ratio / weight) for overlapping response portions. For example, for a section related to the category "environment," the response from the first specialized LLM corresponding to "environment" is prioritized.

[0158] [Sequence (5): Examples 2E and 2C: Selection of Multiple LLMs] Figure 4C is a modified example of Figure 4A and corresponds to Examples 2E and 2C. This example shows a sequence in which the response output device 10010 selects multiple LLMs according to the context of the instruction statement, synthesizes multiple responses (answers) from the multiple LLMs, and presents them to the user. This example also shows a case in which an AI model (persona) is configured using multiple LLMs. This example shows a case in which a request-response is performed without going through a general-purpose LLM (general-purpose LLM server 19001G).

[0159] In step 4C01, the control unit 1110 of the response output device 10010 generates an instruction statement based on the user input. In step 4C02, the control unit 1110 performs a category classification process on the instruction statement. In step 4C03, the control unit 1110 performs a process of selecting multiple LLMs (specialized LLMs) to use based on the classification of the instruction statement. The control unit 1110 creates an instruction statement for each selected LLM. In step 4C04, the control unit 1110 transmits the instruction statement to each selected LLM. The instruction statement to be transmitted may be the same instruction statement transmitted to each LLM, or the instruction statement may be divided into multiple parts according to category classification, and each part may be transmitted to the specialized LLM to be used.

[0160] In this example, if multiple LLMs are selected corresponding to multiple classifications through the category classification process in step 4C02, the control unit 1110 selects the adoption ratio (in other words, the ratio / weight) for the multiple LLMs from the instruction text in step 4C03. This corresponds to the persona composition described below (FIG. 4D). The control unit 1110 creates an instruction text for each LLM according to the adoption ratio.

[0161] In steps 4C05-1 to 4C05-N, each selected LLM (LLM server 19001) performs processing to generate a response to the instruction statement. In step 4C06, each LLM transmits the generated response to response output device 10010. The response may include corresponding LLM information indicating the LLM used. In step 4C07, control unit 1110 first presents and outputs the corresponding LLM information to the user. In step 4C08, control unit 1110 combines the responses from each LLM to obtain a single answer statement. In step 4C09, control unit 1110 presents and outputs the combined response (answer statement) to the user.

[0162] In this example, in the response synthesis process of step 4C08, the control unit 1110 synthesizes the responses (answer sentences) from each LLM based on the adoption rate of step 4C03 to create a single answer sentence by the persona. The answer sentence does not necessarily have to be created using all of the configured LLMs; it may be created by selecting from among the configured available LLMs. This persona is an AI model constructed by combining multiple selected LLMs. Furthermore, in step 4C09, the control unit 1110 may present the answer sentence to the user on a GUI screen (described below) so that the answer sentence is created by the persona. Furthermore, as described above, when presenting the synthesized response (answer), information regarding which parts of the answer sentence are based on which specialized LLM may be displayed.

[0163] [Persona Screen Example (1)] Figure 4D shows an example screen for creating and setting a persona (i.e., an anthropomorphic personality) as an AI model composed of multiple candidate LLMs (e.g., specialized LLMs in Figures 2A and 2B) that can be used for a response, as an example of Figure 4C, and generating and outputting an answer using the persona. The system creates and sets a persona corresponding to a single LLM, or a persona composed of multiple LLMs combined with a given ratio / weight. For example, the system creates an answer by the persona by exchanging instructions and responses with the multiple LLMs that make up the persona, as shown in Figure 4C, and synthesizing the multiple responses.

[0164] The screen example of FIG. 4D illustrates an example in which an AI model corresponding to a persona constructed based on multiple selected LLMs is set as a personal advisor / assistant that responds to a user's instructions. A screen such as that shown in FIG. 4D is displayed as a GUI screen on the display unit 10011 of the response output device 10010. The control unit 1110 creates various GUI screens. The screen of FIG. 4D includes an image, video, icon, etc. representing the persona / advisor / assistant, as shown in the "Dedicated Advisor" column 4D02. The image of the persona may be a human figure or character image, or may be set arbitrarily by the user. Furthermore, this GUI is not limited to display only, and may also include audio output. The persona may provide answers, guidance, etc., not only in text but also in audio.

[0165] The number and ratio of LLMs to be used as personas among multiple LLMs (for example, the general LLM, specialized LLM-A, ..., specialized LLM-Z in Figure 2B) are selected and set by the system in consideration of the content of the instruction. The system may also use past history information to call and use previously set personas.

[0166] In addition, in the present system, each persona configured by combining each LLM may be pre-configured. The control unit 1110 may select and call the persona to be used in response to an instruction. Alternatively, the user may select the persona to be used in response to an instruction. Furthermore, the user may create and configure a persona by combining each LLM in advance, save it as their own personal advisor / assistant, and call it up as desired. Furthermore, the system may display the configuration information of a persona prepared in advance as a default, etc., on a GUI screen, and the user may confirm the persona on the GUI screen and adjust and update the configuration information of the persona, thereby saving and using it as a personalized, customized persona. The various configured personas are stored in the system's resources (e.g., the memory of the response output device 10010) and can be called up as desired for use. The user can configure and use their own customized persona, which can be used as an expert / advisor / assistant.

[0167] Examples of the composition of LLMs that make up a persona include the following: One persona is created with 70% generic model and 30% medical model. This persona is a composite model that evokes the image of a generalist with medical experience. The medical model is an LLM specialized in the medical field, and corresponds to "medicine" in the classification of Figure 2C. In another example, a persona is created with 80% medical model and 20% generic model. In another example, a persona is created with 20% generic model, 40% legal model, and 40% economic model.

[0168] The example screen of FIG. 4D includes a "Instruction Statement Dedicated AI" field 4D01. The "Instruction Statement Dedicated AI" field 4D01 includes a "Dedicated Advisor" field 4D02, a "Used LLM Settings" field 4D03, a LOAD button, a SAVE button, and the like. The "Dedicated Advisor" field 4D02 displays an image representing a persona, which is a dedicated advisor. In addition to the persona image, the persona's name / ID, a description of its personality, characteristics, and the like may also be displayed. The "Used LLM Settings" field 4D03 displays one or more LLMs used by the persona and their usage ratios (ratios / weights). For example, as shown in the figure, the LLM usage ratios may be displayed in the form of a graph such as a radar chart. In this example, the usage ratios of five types of LLMs are displayed, but the number is not limited to five and may be increased or decreased as needed.

[0169] In the example screen of FIG. 4D , the persona corresponding to the "Dedicated Advisor" column 4D02 is set as follows: General LLM 50%, Specialized LLM-A 5%, Specialized LLM-B 10%, Specialized LLM-C 20%, and Specialized LLM-D 15%. The GUI for the user to set the persona may include, for example, a GUI for selecting an LLM and a GUI for selecting the adoption rate of the selected LLM. The LOAD button is a GUI for calling up a saved persona. The SAVE button is a GUI for saving the persona on the screen.

[0170] 4C, the control unit 1110 sets the adoption rate (ratio / weight) of each LLM to be used based on the result of categorizing the instruction sentence input by the user. Alternatively, the control unit 1110 may select and call the persona having the most suitable adoption rate from the personas stored in the system in accordance with the result of categorizing the instruction sentence.

[0171] Furthermore, when providing instructions and responses to each LLM constituting a persona, the adoption ratio (ratio / weight) of the LLMs used may be varied for each instruction sentence. For example, for part A of the entire instruction sentence, LLM-A and LLM-B may be used in a ratio of 7:3, and for part B of the instruction sentence, LLM-A and LLM-B may be used in a ratio of 2:8. In this case, if the answers obtained from each LLM for the same instruction sentence (part) differ between the LLMs, the control unit 1110 prioritizes the answer from the LLM with the higher adoption ratio (ratio / weight) during the synthesis process of step 4C08. Furthermore, during the synthesis process, the control unit 1110 creates or invokes a persona as shown in FIG. 4D, in which the answers from each LLM are adopted according to their adoption ratios.

[0172] The settings for the adoption rate of each LLM created by the control unit 1110 or the user of this system are saved as personas and can be recalled and used when issuing other instructions. The settings that provide the answers desired by the user can be used repeatedly as personas. This eliminates the need to reconfigure the settings when issuing subsequent or related instructions.

[0173] Furthermore, the personas used by a user may be shared and usable not only by the user himself but also by other users (other users of the system). For example, when the LOAD button is pressed on the screen of Fig. 4D, personas set by other users are listed among the options, and the user can select and use the personas. The number of available personas may increase depending on the user's billing status, etc., or depending on the number of times or period of use.

[0174] The control unit 1110 also presents to the user various personas with adoption rates set by the system as candidates. The control unit 1110 may select one or more suitable personas as candidates from the personas set by the system in response to the user or instruction, and present them as recommendations. The user can select a persona to use from the recommendations (described below).

[0175] [Persona Screen Example (2)] Figure 4E shows another example screen with a persona. The difference between Figure 4E and Figure 4D is that the "LLM Used" field 4E03 has a different GUI than field 4D03. Specifically, each LLM that constitutes a persona is displayed on a separate row. The LLM information 4E04 for each row includes an ON / OFF button, an LLM (LLM name), and a percentage (ratio / weight). The ON / OFF button allows the user to select whether or not to use each LLM row. The percentage value allows the user to set the adoption rate of the LLM. The percentage value may be entered directly or via GUI operations such as up / down buttons or a bar. In the illustrated example, a certain persona is configured with 50% general LLM, 30% specialized LLM-B, and 20% specialized LLM-E, for a total of 100%.

[0176] FIG. 4F shows another example of a screen in which a persona issues a response. In this example, the "AI Response" column 4F01 displays an image of the dedicated advisor, which is the persona being used, and the "Used LLM" column 4F02 displays information on multiple LLMs (e.g., category, LLM, and ratio) that make up the persona. In this example, a speech bubble GUI appears from the image of the dedicated advisor, which is the persona (e.g., a character image), and a response statement 4F03 from the persona is displayed in the speech bubble GUI. The user can receive a response to their instruction as a response from the persona. When the response is displayed, audio may be used in combination with the display, or audio alone may be used. Audio may be set for each dedicated advisor image, or the user may select from multiple options.

[0177] [Sequence (6): Two-Stage Response: Example 2G] Figure 5A shows a sequence diagram of an example corresponding to Example 2G. This example shows a two-stage request-response exchange using a general-purpose LLM and a specialized LLM. In this system, in the first stage, an instruction is sent from the response output device 10010 to the general-purpose LLM to receive a primary response. Then, if necessary, in the second stage, an instruction is sent from the response output device 10010 to the specialized LLM to receive a detailed response as a secondary response. Here, different LLMs are used for the primary and secondary responses. A secondary response may be available depending on the user's charges, number of uses, and usage period.

[0178] For the first stage of instruction-answer questions, the system generates an answer (primary answer) using a general-purpose LLM and presents and outputs the primary answer (i.e., the first answer) to the user. In parallel with this answer, for the second stage of instruction-answer questions, the system also generates an answer (secondary answer) using a specialized LLM corresponding to the category classification and presents and outputs the secondary answer (i.e., the second answer) to the user. While the first stage of instruction-answer questions and the second stage of instruction-answer questions are typically performed sequentially, this is not a limitation. For greater efficiency, the second stage may be processed in parallel while the first stage is being processed. When responding to the first stage of the primary answer, the system notifies the user that an additional response (secondary answer) can be requested. However, if a specialized LLM capable of providing an additional response (secondary answer) does not exist or cannot be used, the system does not notify the user that an additional response (secondary answer) can be requested. When the user confirms that they can request an additional response (secondary response) after viewing the primary response and inputs a request for the additional response (secondary response), the system presents and outputs the secondary response to the user using a specialized LLM that was previously generated. Furthermore, when the system requests an additional response (secondary response), the system may also notify the user in advance of the corresponding LLM information indicating which specialized LLM corresponds.

[0179] In FIG. 5A, in step 5A01, the control unit 1110 of the response output device 10010 generates a user-input instruction (hereinafter referred to as "first-stage instruction 1"). In step 5A02, the control unit 1110 transmits instruction 1 to the general-purpose LLM server 19001G. In step 5A03, the general-purpose LLM of the general-purpose LLM server 19001G performs a category classification process on the received instruction 1. In step 5A04, the general-purpose LLM performs a process to generate a response (hereinafter referred to as "first-stage response 1"; corresponding to the primary response) based on the category classification results. In step 5A05, the general-purpose LLM server 19001G transmits response 1 to the response output device 10010. Response 1 may include corresponding LLM information indicating the general-purpose LLM that generated response 1. In step 5A06, the control unit 1110 presents and outputs the response sentence of response 1 (primary response) to the user as a process corresponding to the received response 1.

[0180] In step 5A07, the control unit 1110 performs a request process for an additional response if an additional response (corresponding to a secondary response) is possible. The availability of an additional response is determined by determining whether a specialized LLM corresponding to the category classified in step 5A03 exists or whether the corresponding specialized LLM is available. The control unit 1110 performs a request process for an additional response if, for example, the user requests an additional response in response to a user input operation on the screen. The control unit 1110 creates a second-stage instruction statement (hereinafter referred to as instruction statement 2) for the additional response. In this embodiment, at this stage, a specialized LLM to be used for the secondary response has not yet been selected. In instruction statement 2, the LLM to be used has not yet been identified or specified. In step 5A08, the control unit 1110 sends instruction statement 2 to the general-purpose LLM server 19001G. In step 5A09, the general-purpose LLM of the general-purpose LLM server 19001G performs a process to select a specialized LLM to be used for the second-stage response (secondary answer) for the received instruction sentence 2. For example, assume that specialized LLM-B is selected based on the results of the previous category classification.

[0181] In this processing example, one specialized LLM is selected in the second stage, but multiple specialized LLMs to be used in the answer may be selected in the second stage, as in the previous example.

[0182] The general-purpose LLM server 19001G transmits Instruction 2 to the selected specialized LLM, for example, specialized LLM-B. The transmitted Instruction 2 has the same content as that received in step 5A08, but the requestor is the response output device 10010 and the destination is the specialized LLM-B. Instruction 1 and Response 1 may also be transmitted in step 5A08 along with Instruction 2. In step 5A10, the specialized LLM server 19001B of the selected specialized LLM-B performs processing to generate a second-stage response (response 2, corresponding to a secondary response) in response to the received Instruction 2. In step 5A11, the specialized LLM server 19001B transmits Response 2 to the response output device 10010. In step 5A12, the control unit 1110 presents and outputs the response text of Response 2 (secondary response) to the user as processing in response to the received Response 2.

[0183] 5A illustrates the first-stage processing (steps 5A01 to 5A06) and the second-stage processing (steps 5A07 to 5A12) as being performed sequentially, but as described above, the second-stage processing may be started in parallel midway through the first-stage processing. For example, before the user requests an additional response in step 5A07, control unit 1110 causes the general-purpose LLM to perform the LLM selection process in step 5A09, generates response 2 using the selected specialized LLM, and acquires this response 2. Then, when the user requests an additional response, control unit 1110 presents a reply sentence using the acquired response 2.

[0184] [Processing Example (6)] Figure 5B shows a processing example related to Figure 5A. This example is similar to Figure 3D up to the instruction sentence classification process. In Figure 5B, the control unit 1110 sends instruction sentence 1 to the general-purpose LLM, which performs instruction sentence keyword search and instruction sentence classification process to generate response 1. The control unit 1110 receives response 1 from the general-purpose LLM. The control unit 1110 presents the answer sentence based on response 1 to the user and also indicates that an additional answer is possible if a specialized LLM corresponding to the classified category exists and is available. If a specialized LLM corresponding to the classified category does not exist and the corresponding specialized LLM is unavailable, only the answer sentence based on response 1 is presented to the user. An example of the answer sentence in the user interface is, "Your symptoms are suspected to be a cold, influenza, or novel coronavirus infection. *You can add more detailed information to obtain a specialized answer."

[0185] The display of the answer also includes a GUI for requesting an additional response, such as a "Request Additional Response" button. The GUI for requesting an additional response allows the user to input additional information (i.e., detailed information / related information). This additional information is to be passed to the LLM along with the instruction. For example, additional information could be "body temperature 37.7 degrees, sore throat" or an attached image of the throat. In step 5A07 of FIG. 5A , the control unit 1110 uses the additional information entered by the user through the GUI to create Instruction 2 and sends it to the general LLM or the selected specialized LLM. The general LLM selects the LLM to use in step 5A09 in response to Instruction 2. For example, a specialized LLM-B (name: Medical-LLM) that includes a medical specialty is selected. The general LLM transmits Instruction 2 to the selected specialized LLM-B. The general-purpose LLM may also transmit instruction 1 and response 1 to the selected specialized LLM-B together with instruction 2.

[0186] In addition, when selecting an LLM, if there is no suitable specialized LLM corresponding to the classified category, a general LLM may be used without selecting a specialized LLM, and a second stage response 2 may be generated using additional information from the user.

[0187] Additionally, while waiting for response 2, the control unit 1110 first presents the corresponding LLM information to the user. For example, a message such as "Medical-LLM will respond to your questions regarding the details of your symptoms" is output on the user interface.

[0188] The selected specialized LLM-B generates response 2 in response to instruction sentence 2 and transmits response 2 to the response output device 10010. The control unit 1110 receives response 2 and presents it to the user. Response 2 may be displayed, for example, as follows: "Based on your symptoms of fever, sore throat, and runny nose, it is highly likely that you have influenza. You will need to be administered an anti-influenza drug. It is recommended that you undergo a viral test at the nearest hospital to identify the disease."

[0189] In this example, the system first obtains a first-stage primary answer from the general-purpose LLM. If the user determines that the answer is insufficient or that an additional answer is desired, the system obtains a second-stage secondary answer from the selected specialized LLM to elicit more specialized information. If the user is satisfied with the first-stage primary answer (no request for an additional answer is made), or if an additional answer is not possible, the second-stage request-response can be omitted.

[0190] Furthermore, if there is a specialized LLM in the field / category of the instruction, the system may indicate this in advance and ask the user whether to request an additional response / detailed response (a second-stage response using the specialized LLM). For example, at step 5A01 of FIG. 5A , this information may be displayed on the screen, allowing the user to input whether to request an additional response. If there is an input requesting an additional response, the control unit 1110 may create and transmit an additional instruction for requesting an additional response in addition to instruction 1. Alternatively, the additional instruction or information may be written in instruction 1, or the additional instruction or information may be written in metadata, etc. Alternatively, the response output device 10010 (control unit 1110) may request additional information from the user that is expected to be necessary for the additional response, and the user may input the additional information in response to the request. In this case, the control unit 1110 transmits the input additional information along with the additional instruction.

[0191] [Sequence (7): Example 2G] Figure 5C shows a modified example of Figure 5A, in which a response (primary response) is obtained from a general LLM in the first stage and a response (secondary response) is obtained in advance from a specialized LLM in the second stage, thereby reducing the time lag for the user. The processing example of Figure 5C, in other words, is a case in which the second stage processing is performed in parallel from the middle of the first stage processing. In Figure 5C, for the first stage answers, Answer 1 is generated using the general LLM and returned to the user of the response output device 10010. In parallel with the first stage processing, the system also generates Answer 2 in a specialized LLM corresponding to the category classification as the second stage processing, and obtains Answer 2 as soon as possible. When presenting Answer 1, the system notifies the user that an additional answer can be requested. If the user further inputs a request for an additional answer, the system presents Answer 2, which has been generated and obtained in advance, to the user of the response output device 10010. This reduces the time lag experienced by the user. When requesting an additional answer, the user may be informed in advance of which specialized LLM corresponds (corresponding LLM information).

[0192] In FIG. 5C, the first-stage processing from step 5C01 to step 5C06 is the same as in FIG. 5A. The general-purpose LLM of the general-purpose LLM server 19001G generates response 1 in step 5C04, transmits response 1 in step 5C05, and performs LLM selection processing to select a specialized LLM to use based on the category classification in step 5C09. Depending on the capabilities of the general-purpose LLM server 19001G, the processing in step 5C09 may be performed in parallel with the processing in step 5C04, or may be performed sequentially as shown. The general-purpose LLM server 19001G transmits instruction 1 (or instruction 2 created based on instruction 1) to the selected specialized LLM, for example, specialized LLM-B. In step 5C10, the selected specialized LLM-B generates response 2 in response to instruction 1 (or instruction 2) and transmits response 2 to the general-purpose LLM server 19001G. In step 5C11, the general-purpose LLM server 19001G receives and temporarily stores response 2.

[0193] In a modified example, response 2 may be transmitted from the specialized LLM to the response output device 10010.

[0194] Meanwhile, after processing corresponding to response 1 in step 5C06, the control unit 1110 of the response output device 10010 performs additional response request processing in step 5C07. If the user inputs a request for an additional response or detailed information, the control unit 1110 sends a corresponding additional response request to the general-purpose LLM server 19001G. If no such request is received, the process returns to step 5C01. Upon receiving the request, the general-purpose LLM server 19001G sends the temporarily stored response 2 (secondary response) together with the corresponding LLM information to the response output device 10010 in step 5C12. In step 5C13, the control unit 1110 receives response 2 and presents and outputs the corresponding LLM information and secondary response to the user.

[0195] [Sequence (8): Examples 2B and 2D: User Selection of LLM to Use] FIG. 6A shows a sequence diagram of an example corresponding to Examples 2B and 2D. This example illustrates a process example in which a user of the response output device 10010 selects an LLM to use for a response based on an instruction statement created by the user. In this process example, the user selects and specifies the LLM to use for outputting the response based on the content of the instruction statement created by the user. One or more LLMs may be selected as the LLM to use for the response. For example, one or more LLMs may be selected from the multiple available candidate LLM servers 19001 in FIG. 1A or FIG. 1B. FIG. 6A illustrates an example of selecting multiple specialized LLMs.

[0196] When multiple LLMs are specified, the following methods and processing examples can be applied.

[0197] (1) An LLM may be selected or designated based on the weight (i.e., ratio or adoption rate) or priority of each LLM in the multiple LLMs used. The weight here refers to the weight related to the reflection in the answer, and the priority refers to the priority of the LLM to be reflected in the answer. For example, if the category or keyword information of a directive is registered in multiple LLMs, the LLM is selected or designated based on the weight and priority information.

[0198] (2) You may select and specify the LLM to use for each part of the instruction.

[0199] (3) If the user does not make any particular selection or designation, the system (control unit 1110) may select the LLM to be used from among the corresponding LLM candidates based on the category and keyword information of the instruction statement.

[0200] (4) If the system determines that there is another LLM available that is more suitable / optimal than the LLM selected by the user, the system may present the other LLM to the user as a recommended LLM. The system may also present an answer using the recommended LLM to the user.

[0201] The information used to generate the answer, such as the selection of the LLM and the weight of the LLM, can be recorded in the system and read out for reuse or used by others. In the latter case, the system uploads the information to a server, for example, so that others can access the server and use the information.

[0202] In FIG. 6A , in step 6A01, the control unit 1110 of the response output device 10010 generates a user-input instruction sentence. In step 6A02, the control unit 1110 provides a GUI screen (described below) that allows the user to select an LLM to use in response to the instruction sentence, and the user selects and specifies the LLM to use as needed on the GUI screen. On the GUI screen, the user selects one or more LLMs from multiple available candidate LLMs. In this example, multiple specialized LLMs are selected. In step 6A03, the control unit 1110 transmits the instruction sentence, along with information on the LLM selected by the user (corresponding LLM information), to, for example, the general-purpose LLM server 19001G.

[0203] In step 6A04, the general-purpose LLM server 19001G receives the instruction and performs preprocessing. This preprocessing involves sending the instruction to one or more selected or designated LLMs (especially specialized LLMs). The instruction may be sent as a single instruction to each LLM, or the instruction may be divided into multiple parts based on category classification, with each part being sent to the specialized LLM to be used. In steps 6A05-1 through 6A05-N, each specialized LLM (LLM server 19001A-19001Z) that received the instruction performs a response generation process and sends the generated response to the general-purpose LLM server 19001G. In step 6A06, the general-purpose LLM server 19001G receives responses from each specialized LLM and performs a response preparation process. This response preparation process may involve a synthesis process that combines one or more responses from specialized LLMs into a single answer. Here, the general-purpose LLM server 19001G infers a synthesis of a response based on the weights and priority information set by the user for the response sentences obtained from each specialized LLM. In step 6A07, the general-purpose LLM server 19001G transmits the synthesized response to the response output device 10010 along with corresponding LLM information indicating the LLM actually used in the response. In step 6A08, the control unit 1110 receives the response and presents and outputs the corresponding LLM information and the response sentence to the user.

[0204] [Screen Example: Selecting the LLM to Use] Figure 6B shows an example of an input screen provided to the user for the processing example of Figure 6A. In step 6A02, control unit 1110 presents a screen such as that shown in Figure 6B. On this screen, the user can input a prompt and select and set the LLM to use for the answer. The user can also not select an LLM to use. When selecting multiple LLMs to use when setting the LLM to use, the ratio / weight / adoption rate and priority of each LLM to use can be selected and set.

[0205] The example screen of FIG. 6B includes a "User Input Instructions" column 6B01, a "Select LLM to Use" column 6B02, and an "AI-Generated Response" column 6B03. The "User Input Instructions" column 6B01 displays instructions entered by the user. For example, the instructions might read, "I'm thinking of launching an e-commerce website for the product shown in the photo. Please tell me information about the market size of this product and competing products. Please also provide a sample HTML code for a page introducing this product, including the following: (1) Product Features, (2) Differences from Similar Products, (3) Important Notes, etc." The "User Input Instructions" column 6B01 also allows for text input, as well as image, video, and audio input, with corresponding buttons provided for each. In the example of the instructions above, a photo of the product is attached using the "Image Input" button.

[0206] In the "LLM Selection" column 6B02, information on the LLMs available as options for the LLM is displayed row by row, and the user can turn the LLMs on (selected) or off (unselected) using the ON / OFF button. In addition, the "Percentage" field allows the user to set the weight / ratio / adoption rate to be reflected in the answer for the selected (ON) LLM. For example, a general LLM is specified as 50%, a specialized LLM-B as 30%, and a specialized LLM-E as 20%, totaling 100%. For example, a specialized LLM-B is an LLM in the field of economics, and a specialized LLM-E is an LLM in the field of IT.

[0207] The "AI-generated answer" field 6B03 displays the answer provided by the selected LLM to the instruction. An example of the answer might be, "The product in the photo is XXX. The annual sales amount for this product in 2023 will be XXX billion yen. The top-selling product in 2023 is from Company ABC. ... Sample code is below: <!DOCTYPE html> ..."

[0208] Furthermore, if a user's question is a compound question, in other words, a question having parts related to multiple fields / categories, it is desirable to use an LLM associated with each field / category corresponding to each part of the instruction in the answer to the question. In this case, the screen of FIG. 6B is provided with a GUI that allows the user to select an LLM to use for each instruction or for each part. For example, as a first GUI / method, the user enters a first sentence corresponding to a first category in field 6B01, taking into account the grouping of fields / categories, and selects an LLM to use in field 6B02. Next, the user enters a second sentence corresponding to a different second category in field 6B01, and selects an LLM to use in field 6B02. In this manner, the user may select a corresponding category and LLM for each sentence. As another second GUI / method, the user enters an instruction in field 6B01, and the user or control unit 1110 divides the instruction into parts, taking into account the grouping of fields / categories. For example, it is divided into two sentences or paragraphs. Then, in field 6B02, the user selects the LLM to be used for each part.

[0209] [Screen Example: Persona Selection] Figure 6C shows another example of an input screen related to the processing example of Figure 6A. This screen example shows a case where, when selecting an LLM to be used in an answer, instead of directly specifying one or more LLMs as in Figure 6B, a persona, which is an AI model composed of multiple LLMs, is specified. The system pre-configures persona settings information, including the LLM selected from candidate LLMs and the usage ratio of each LLM. In the screen example of Figure 6C, the "Answer AI Selection" field 6C02 displays information on multiple available candidate personas. Each of these personas generates an answer by combining the answers of the selected LLMs in accordance with the set LLM usage ratio. Selectable personas may also be added based on the user's billing information, etc., or based on the user's usage frequency, usage time, and usage frequency.

[0210] The system generates an answer for a persona by combining multiple answers generated by each LLM used in the persona, using a set ratio as a guideline. It is not necessary to generate an answer using all of the configured LLMs; an answer may be generated by selecting from the configured available LLMs. Alternatively, the system divides a directive into parts (e.g., sentences), assigns answers to LLMs taking into account the respective ratios of each part, and then combines multiple answers to generate an answer for a persona. Furthermore, these persona settings may be read and used by others, or users may upload their own settings and allow others to use them.

[0211] In FIG. 6C, examples of instructions and answers are similar to those in FIG. 6B. In the "Answer AI Selection" column 6C02, for example, personas 6Ca to 6Cf are displayed as candidate personas. Display information for each persona includes the persona's character image, name, selection button, and the like. The user can select the persona to use in the answer by turning the button ON / OFF. For example, persona 6Ca, whose name is "John," is selected. Detailed information for each persona can also be viewed. For example, when the user performs a predetermined operation (e.g., double-clicking) in the display area of ​​persona 6Ca, information about the used LLM that constitutes persona 6Ca is displayed as detailed information about the persona. The display of this detailed information may be similar to that in column 6B02 in FIG. 6B. Furthermore, unavailable personas may be hidden or their display state may be changed based on the user's billing information or usage information.

[0212] [Screen Example: Persona Setting] FIG. 6D shows another example of an input screen related to the processing example of FIG. 6A . This screen example allows a user to set a persona, which is an AI model to be used in a response. On this screen, the user can set the LLMs (type and name) that make up the persona and the adoption rate, etc. In another example, the rate setting may be omitted and only the used LLMs may be set. When only the used LLMs are set as the persona setting, the system generates the persona's response using the used LLMs of the selected persona in accordance with the instruction (a suitable LLM may be selected and used). When the persona setting includes the used LLMs and adoption rate (ratio / weight), the system generates the persona's response using the used LLMs of the selected persona at the set adoption rate in accordance with the instruction.

[0213] In another example, when the system uses a persona selected by the user for a command statement, the control unit 1110 may select a suitable LLM that matches the command statement from among the multiple LLMs that make up the persona, and generate a response. For example, assume that a persona set by the user uses specialized LLMs-A, B, and C that correspond to categories A, B, and C. Assume that the control unit 1110 determines, as a result of categorization, that categories A and B are highly relevant to a command statement. In this case, the control unit 1110 selects specialized LLMs-A and B as suitable LLMs from the composition of the user's persona, and generates a response.

[0214] 6D, the user may be allowed to set not only the LLMs and their ratios to be used, but also the response style (i.e., the characteristics of the response) required when creating a response. Examples of the response style (the characteristics of the response) include the tone / voice of the response, the length of the response, and whether or not examples are provided.

[0215] In the example screen of Figure 6D, an "AI setting screen" 6D01 has an "AI character" field 6D02 and a "LLM selection to use" field 6D03. The "AI character" field 6D02 displays an image of the AI ​​model / persona / character to be set, basic settings such as name, gender, and age, and other characteristics, and these settings can be changed in response to user input operations. Examples of other characteristics include "friendly tone of voice," "emphasis on short summaries in responses," and "not assertive tone for low-confidence responses."

[0216] In this way, in this embodiment, the user can tune the AI ​​model used for answering questions, and create an AI assistant that is tailored to the user.

[0217] [Screen Example: Persona Setting] FIG. 6E shows a screen example modified from FIG. 6D . The "Used LLM Selection" display may be displayed as a bar graph or radar chart. The "Used LLM Selection" field 6E03 in FIG. 6E is displayed as a bar graph, unlike the row display of the "Used LLM Selection" field 6D03 in FIG. 6D . In the bar graph 6E04, (1), (2), and (3) indicate the used LLMs and adoption rates, which are components of the AI ​​model, as displayed in the table below. The rates are arranged in descending order of (1), (2), and (3). The user may, for example, manipulate the table to change the used LLMs and rates. Alternatively, the user may, for example, manipulate the (1), (2), and (3) displays in the bar graph 6E04 with a cursor or the like to change the rates.

[0218] [Screen Example: Selecting an LLM to Use for Each Part] Figure 6F shows another example of an input screen for the process example of Figure 6A. This example includes a GUI that allows users to select an LLM to use for each part, e.g., sentence, along with inputting a directive. When inputting a directive, the user can set the LLM to use. At this time, the user can select and set which LLM to use for the answer for each part, e.g., sentence, of the directive. Furthermore, for parts of the directive not selected by the user (i.e., parts for which no LLM was specified), the system may automatically set the system to use a general-purpose LLM or a predetermined LLM. For answers to complex directives, it is desirable to use the LLM (especially a specialized LLM) corresponding to each part.

[0219] The example screen in FIG. 6F includes a "User-Input Instructions" column 6F01, a "Select LLM to Use" column 6F02, and an "AI-Generated Response" column 6F03. In the "User-Input Instructions" column 6F01, the instruction is divided into multiple parts. The system or the user divides the instruction into multiple parts. This division can be performed by a user entering a separator, or the control unit 1110 can analyze the instruction and automatically divide it into units such as sentences. For example, the first sentence might be "I'm thinking of launching an e-commerce site related to the product shown in the photo." The second sentence might be "Please tell me about the market size of this product and information about competing products." The third sentence might be "Please also provide three sample HTML code for a page introducing this product. The page should include the following: (1) Product Features, (2) Differences from Similar Products, (3) Notes, ...." In this example, the separators are by sentence, but they may also be by paragraphs or clauses.

[0220] In this example, the user can select the LLM to use for each delimiting sentence in the instruction sentence in the "Select LLM to Use" field 6F02. For example, a general-purpose LLM is selected for the first sentence, a specialized LLM-B is selected for the second sentence, and a specialized LLM-E is selected for the third sentence. Alternatively, to obtain a response that takes into account the content of the surrounding sentences, the surrounding sentences may be sent to the selected LLM.

[0221] An example of an overview of response generation using the LLM used for each sentence is as follows. First, the response generation for the first sentence is as follows. For the attached input image corresponding to the "product in the photo" in the first sentence, a general-purpose LLM is used to infer what kind of item it is through image recognition. For example, an inference result indicating what kind of item it is is obtained. Note that if the inference result does not indicate what category the item belongs to, the general-purpose LLM is used to classify the category. If the category, etc. is specified in advance in an instruction, etc., a specialized LLM matching that category, etc. may be selected to generate a response.

[0222] Next, the response generated for the second sentence is as follows. From the second sentence, a specialized LLM that specializes in market information for the competing product is selected. For example, a specialized LLM that has been educated with a focus on marketing theory is used as specialized LLM-B. This allows proposals to be made that utilize methods and concepts for seizing marketing opportunities, such as market research, segmentation, target market selection, and analysis of customer purchasing behavior, while suppressing the interference of other information.

[0223] The response generated for the third sentence is as follows. Because the third sentence is an instruction for creating an actual web page, for example, a specialized LLM-E that has been trained specifically in web design techniques is used. This allows for the creation of a sample web page that the user desires.

[0224] [Screen Example: Persona Settings] Figure 6G shows a variation of the persona settings screen example of Figure 6D. The "Select Used LLM" field 6G03 in Figure 6G is displayed differently from the "Select Used LLM" field 6D03 in Figure 6D. Specifically, in field 6G03, for each row of LLM information, those that can be used to generate answers are displayed normally, and those that cannot be used are displayed in gray. In the illustrated example, the available LLMs are "General LLM," "Specialized LLM-B," "Specialized LLM-D," and "Specialized LLM-E."

[0225] The LLM that can be used to generate a response to a command may be changed depending on the following information: (1) user billing information, (2) user contribution, and (3) user usage frequency.

[0226] The user's billing information (1) above is billing information when a user is charged for using the functions / services of this system. For example, the larger the billing amount, the more LLMs that can be used. It may also be possible to purchase usage rights for each LLM.

[0227] The user's contribution level in (2) above can be determined by entering into a contract to provide the system developer with information such as the user's usage status and setting data when using the system's functions / services, and the level of contribution can be determined based on the degree of provision. The period of use in the system usage status is also one example.

[0228] The frequency of use in (3) above is a perspective regarding the LLM that a user prefers to use. Even if there are multiple candidate LLMs across the entire service, the LLMs used may be limited and biased depending on the user. This system infers that some LLMs will be used more frequently and selects the LLM to use taking into account the user's preferences / inclinations.

[0229] The system determines which of the multiple candidate LLMs are usable and which are not, based on the above information according to the user. Unusable LLMs are displayed on the screen, for example, grayed out and unselectable, as shown in FIG. 6G . In the example of FIG. 6G , rows such as "Specialized LLM-A" and "Specialized LLM-C" are grayed out and unselectable depending on the user's status. This prevents the user from selecting the unusable LLMs and forces the user to select the LLM to use from the normally displayed usable LLMs. While the example of FIG. 6G illustrates a case in which the user selects an LLM to use in the composition of a persona, this is not limiting. Similarly, when selecting an LLM to use in response to a command, as shown in FIG. 6B , the display of usable and unusable LLMs can be controlled depending on the user's status.

[0230] The example in Figure 6G illustrates a case where a user selects an LLM to be used. However, even when the system automatically selects an LLM to be used, the available and unavailable LLMs may be changed based on the above information (factors). Furthermore, the system may present a recommended LLM based on the instruction text for the LLM selected by the user. In this case, the recommended LLM may be selected and presented from the available LLMs based on the available and unavailable LLMs based on the above information (factors).

[0231] The technology according to this embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into higher quality, more reliable infrastructure. The introduction of this technology into infrastructure will contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This will contribute to the achievement of the "9 Sustainable Development Goals (SDGs)" advocated by the United Nations, "Build resilient infrastructure, promote inclusive and sustainable industrialization, innovate and innovate."

[0232] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into public transportation facilities to improve access to transportation systems for vulnerable people. The introduction of this technology into public transportation can contribute to improving traffic safety through the expansion of public transportation and realizing access to a safe, affordable, and easily accessible sustainable transportation system for all people. This contributes to "Sustainable cities and communities," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.

[0233] Various embodiments have been described above in detail. However, the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments are detailed descriptions of the entire system in order to clearly explain the present invention, and the present invention is not necessarily limited to a system including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.

[0234] The configurations, functions, processes, etc. described in the above embodiments may be partially or entirely implemented in hardware, for example, by designing an integrated circuit, a general-purpose processor, or an application-specific processor. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. Furthermore, functions, processes, etc. may be implemented in software by the processor interpreting and executing a program that implements each function. Furthermore, the scope of software implementation is not limited, and hardware and software may be used together.

[0235] 10010...Response output device (artificial intelligence response output device), 10011...Display unit, 1110...Control unit, 19001...LLM server.

Claims

1. A response output device comprising: a control unit that generates an instruction sentence based on a user's input, sends the instruction sentence to a large-scale language model (LLM), and obtains an answer sentence from the LLM as a response generated by the LLM; and a display unit that displays the instruction sentence and the answer sentence, wherein when there are multiple LLMs that can be used as the LLM, the control unit obtains the answer sentence as a response generated by an LLM selected from the multiple LLMs according to the category of the instruction sentence.

2. The response output device according to claim 1, wherein the plurality of LLMs available as the LLM include specialized LLMs that have been trained to specialize in specific categories.

3. The response output device according to claim 2, wherein the plurality of LLMs available as the LLM include a general-purpose LLM that is not specialized in a particular category.

4. A response output device according to claim 1, wherein the control unit analyzes the category of the instruction sentence and selects an LLM associated with the category of the analysis result as the LLM to be used in response to the instruction sentence.

5. A response output device according to claim 3, wherein the control unit transmits the instruction to the general-purpose LLM, the general-purpose LLM analyzes a category from the instruction, selects an LLM associated with the category of the analysis results as the LLM to be used in response to the instruction, and transmits the instruction to the selected LLM.

6. A response output device according to claim 3, wherein the control unit sends a request to the general-purpose LLM to analyze the category of the instruction sentence, receives a response of the analysis results of the category by the general-purpose LLM, selects an LLM associated with the category of the analysis results as the LLM to be used in response to the instruction sentence, and sends the instruction sentence to the selected LLM.

7. A response output device according to claim 3, wherein the control unit analyzes the category of the instruction sentence, and if the instruction sentence cannot be classified into a specific category, selects the general-purpose LLM as the LLM to be used in response to the instruction sentence.

8. A response output device according to claim 1, wherein the control unit acquires the answer sentence as a response generated by a plurality of LLMs selected from the plurality of LLMs up to an upper limit value according to the category of the instruction sentence.

9. A response output device according to claim 1, wherein the control unit acquires the answer sentence as a response generated by a plurality of LLMs selected from the plurality of LLMs in accordance with a priority order according to a category of the instruction sentence.

10. The response output device according to claim 1, wherein the control unit provides a screen that displays information on the selected LLM to be used in response to the instruction sentence, or information on the LLM used in response to the instruction sentence.

11. The response output device according to claim 1, wherein the control unit provides a screen that displays the analysis results of the category of the instruction sentence or information on the LLM associated with the category.

12. The response output device according to claim 1, wherein the control unit provides a screen for the user to select an LMM to be used in response to the instruction sentence from the plurality of LLMs.

13. A response output device according to claim 1, wherein the control unit selects, from the plurality of LLMs, a plurality of LLMs constituting an AI model to be used in responding to the instruction sentence according to a category of the instruction sentence, and obtains the response sentence synthesized based on the responses generated by each of the plurality of LLMs constituting the AI ​​model.

14. The response output device according to claim 13, wherein the control unit provides a screen for the user to select the AI ​​model.

15. A response output device according to claim 13, wherein the control unit sets a ratio to be used in the response to the instruction sentence in a plurality of LLMs constituting the AI ​​model.

16. The response output device according to claim 1, wherein the control unit divides the instruction sentence into parts and selects, for each part, an LLM to be used in response to the instruction sentence.

17. A response output device according to claim 1, wherein the control unit obtains a primary response from a first LLM selected from the plurality of LLMs and presents it to the user, and when the user requests an additional response to the primary response, obtains a secondary response from a second LLM selected from the plurality of LLMs and presents it to the user.

18. A response output device as described in claim 17, wherein the multiple LLMs that can be used as the LLM include a specialized LLM that has been trained to specialize in a specific category and a general-purpose LLM that is not specialized in a specific category, and the control unit obtains a primary response from the general-purpose LLM as the first LLM and a secondary response from the specialized LLM as the second LLM.

19. A response output device according to claim 18, wherein the control unit performs the process of acquiring the primary response and the process of acquiring the secondary response in parallel, and when the user requests an additional response, presents the secondary response that has already been acquired to the user.

20. A response output system comprising a large-scale language model (LLM) and a response output device, wherein the response output device comprises: a control unit that generates an instruction sentence based on a user's input, sends the instruction sentence to the LLM, and obtains an answer sentence from the LLM as a response generated by the LLM; and a display unit that displays the instruction sentence and the answer sentence, wherein when there are multiple LLMs that can be used as the LLM, the control unit obtains the answer sentence as a response generated by an LLM selected from the multiple LLMs according to the category of the instruction sentence.

Citation Information

Patent Citations

  • Conversation control program, conversation control method, and conversation control device

    JP7334800B2