Response output device and response output system
Patent Information
- Application Number
- JP2024087410
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-12-11
AI Technical Summary
Existing response output technologies using artificial intelligence do not adequately consider the need for specialized information input to enhance user interaction.
A response output device that includes a control unit to determine the necessity of external information for a large-scale language model and outputs responses generated by either a local or external model, utilizing a first or second large-scale language model based on this determination.
Provides a more suitable response output technique by enhancing user interaction through tailored information input, improving the responsiveness and effectiveness of AI-driven interactions.
Smart Images

Figure 2025180231000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a response output device and a response output system. [Background technology]
[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special table 2019-528512 publication Summary of the Invention [Problem to be solved by the invention]
[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.
[0005] An object of the present invention is to provide a more suitable response output technique. [Means for solving the problem]
[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problem, and one example thereof may be a response output device including: a control unit that transmits an instruction sentence based on a user input to a large-scale language model and causes the large-scale language model to generate a response to the instruction sentence; and an output unit that outputs the response generated by the large-scale language model to a user, wherein the large-scale language models include a first large-scale language model and a second large-scale language model, and the control unit causes one of the first large-scale language model or the second large-scale language model to generate a response to the instruction sentence, depending on a result of a necessity determination that determines whether or not special information held by the response output device is required as input information from outside in order for the large-scale language model to generate a response, and causes the generated response to be output via the output unit. [Effects of the Invention]
[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments. [Brief explanation of the drawings]
[0008] [Figure 1A] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 1B] 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 1C] 1 is a diagram showing an example of the operation of an AI response output device and system according to an embodiment of the present invention; [Figure 2A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2C]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2I] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2J] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2K] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2L] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3C] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3D]1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3I] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4A] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 5A] FIG. 2 is an explanatory diagram of an example of the operation of the artificial intelligence response output device according to an embodiment of the present invention. [Figure 5B] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5C] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5D] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 6] FIG. 10 is an explanatory diagram of an example of a response generation process of the AI response output device according to an embodiment of the present invention. [Figure 7] 1 is a diagram showing an example of an artificial intelligence response output device constituting an artificial intelligence response system according to an embodiment of the present invention; [Figure 8]1 is a flowchart showing an example of a flow of response generation in an artificial intelligence response output system according to an embodiment of the present invention. [Figure 9A] 10A and 10B are diagrams illustrating a first flag of a special information filter according to one embodiment of the present invention. [Figure 9B] 10A and 10B are diagrams illustrating a second flag of a special information filter according to one embodiment of the present invention. [Figure 10A] A figure showing an example of a main message of an instruction sent to a first LLM in one embodiment of the present invention and an example of a main response generated by the first LLM. [Figure 10B] A figure showing an example of a main message of an instruction sent to a second LLM in one embodiment of the present invention and an example of a main response generated by the second LLM. [Figure 11] FIG. 2 is a diagram showing an example of a display state of a display unit serving as a user interface according to an embodiment of the present invention. [Figure 12] 1 is a flowchart showing an example of a flow of response generation in an artificial intelligence response output system according to an embodiment of the present invention. [Figure 13A] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 13B] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 14] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 15] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 16] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 17] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 18]FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 19] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 20] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 21] 1 is a flowchart showing an example of a flow of response generation in an artificial intelligence response output system according to an embodiment of the present invention. [Figure 22] 1 is a flowchart showing an example of a flow of response generation in an artificial intelligence response output system according to an embodiment of the present invention. [Figure 23] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 24] A figure showing an example of a main message of an instruction sent to a second LLM in one embodiment of the present invention and an example of a main response generated by the second LLM. [Figure 25] 1 is a flowchart showing an example of a flow of response generation in an artificial intelligence response output system according to an embodiment of the present invention. [Figure 26] FIG. 2 is a diagram showing an example of display content on a display unit as a user interface according to an embodiment of the present invention. [Figure 27] A figure showing an example of a main message of an instruction sent to a first LLM or a second LLM in one embodiment of the present invention, and an example of a main response generated by the first LLM or the second LLM. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.
[0010] Note that if the AI response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI response output device has an audio output function, it may be referred to as an audio output device. The AI response output device may simply be referred to as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI response output device or the display output of the AI response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.
[0011] Example 1 As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.
[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system in which the AI response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.
[0013] In the example of FIG. 1A, the AI response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Furthermore, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which the pixels emit light themselves. Furthermore, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.
[0014] 1A, the audio output unit 1140 provided in the AI response output device 10010 is configured with a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. By audio input from the microphone 1139 or operation input from the user via an operation input unit (described later), the AI response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.
[0015] The AI response output device 10010 may be provided with a local large-scale language model within the AI response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or an audio output from the audio output unit 1140.
[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.
[0017] Alternatively, the AI response output device 10010 may also include a local large-scale language model, and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.
[0018] The configuration when the AI response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions, or may go via a router or repeater. The AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as one system.
[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
[0020] 1A shows an example in which a display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of FIG. 1A, the prompt display area 10051 displays an icon 10052 representing a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of FIG. 1A, the artificial intelligence response display area 10061 displays an icon 10062 representing an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.
[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment. These large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in these artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called Transformer. Reference 1, for example, has been published as an example of the learning of these models.
[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI response output device 10010 itself. Any large-scale language model can be generated separately through large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001 and the AI response output device 10010. In this way, instead of performing pre-learning for each application or terminal, replicating the large-scale language model, which is the base model generated through large-scale pre-learning, and using it on individual servers and terminals allows the resources used for learning to be shared, resulting in good resource efficiency.
[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.
[0025] Furthermore, large-scale language models can pre-train natural languages and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. FIG. 1A illustrates a large-scale language model server 20001, which is a server having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are AI models generated by large-scale pre-training on natural language and types of information other than natural language text information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. In addition, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.
[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.
[0028] The AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a nonvolatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, and an imaging unit 1180. The AI response output device 10010 may have a large screen, such as a monitor or television.
[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire the user input that serves as the basis for a prompt to the large-scale language model, which is the artificial intelligence.
[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN-type communication method. This allows the AI response output device 10010 to communicate with various servers connected to the Internet 19000.
[0031] The AI response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.
[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.
[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller or the like, and inputs a signal regarding an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. In addition to a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may also be used by, for example, an administrator to operate the AI response output device 10010. By the user's operation input via the operation input unit 1107, the AI response output device 10010 can acquire a user input that serves as the basis for a command sentence (prompt) to a large-scale language model, which is an AI. Note that a modified configuration in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107 is also possible.
[0034] The video signal input unit 1131 is connected to an external video output device and inputs video data. The video signal input unit 1131 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces, etc.
[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.
[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
[0037] The microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals. The microphone may pick up a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI response output device 10010 can acquire user input that will be the basis for instructions (prompts) to the large-scale language model, which is the AI.
[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.
[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.
[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 controls video switching, such as which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also control image processing of the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling, which enlarges, reduces, or deforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.
[0042] The nonvolatile memory 1108 stores various data used by the AI response output device 10010. The data stored in the nonvolatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.
[0043] The local LLM processing unit 10028 has a memory capable of holding a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI response output device 10010.
[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, which is a speaker or the like.
[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 of the AI response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and responses are obtained from these large-scale language models. All of this control can be performed by the control unit 1110.
[0046] The storage unit 1170 may also store a fixed response phrase database (which may be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction sentences from the AI response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI response output device 10010.
[0047] Furthermore, in the example of the fixed response phrases in the database shown in FIG. 1C, if multiple fixed response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the fixed response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the fixed response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.
[0048] Next, an example of condition number 5 shown in Fig. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I might not know about that." By responding in this way, the user can be prompted to input again, and the corrected user input can be waited for.
[0049] Next, an example of condition number 6 shown in Fig. 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in Fig. 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the fixed response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.
[0050] The AI response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).
[0051] 1C described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the response template database (response template DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0052] In the above explanation, it has been explained that the AI response output device 10010 has a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.
[0053] 1A and 1B, an example has been described in which the AI response output device 10010 includes the display unit 10011. However, the AI response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI may be configured to accept input from a user to the AI via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the input from the user via the voice output unit 1140.
[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference of the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.
[0055] <Example 2> Next, as a second embodiment of the present invention, an example will be described in which the AI response output device 10010 described in the first embodiment is connected to the Internet, and is connected to a server equipped with a large-scale language model AI via the Internet to perform operation. In this embodiment, differences from the first embodiment will be described, and repeated explanations of the same configurations as those of these embodiments will be omitted.
[0056] An example of a connection state between an AI response output device 10010 and a large-scale language model server 19001 according to the second embodiment of the present invention will be described with reference to FIG. 2A. The AI response output device 10010 according to the second embodiment may be called a character conversation device. A system including the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment may be called a character conversation system. A video of a character 19051 is displayed on a display unit 10011 displayed by the AI response output device 10010. The video of the character 19051 is generated by rendering a 3D model of the character in a virtual space.
[0057] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can serve as an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this embodiment may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0058] In the example of FIG. 2A, the audio output unit 1140 provided in the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A, an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication is also acceptable. The communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions. The AI response output device 10010 can communicate with a large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with a second server 19002 different from the large-scale language model server 19001 via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 may be considered as one system.
[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2B. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that Figure 2B does not illustrate communication paths such as the Internet 19000 shown in Figure 2A. Figure 2B also illustrates a user 230 of the artificial intelligence response output device 10010.
[0060] Here, we will explain the sequence of operations of the AI response output device 10010. The AI response output device 10010 loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0061] First, the AI response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is in natural language. Note that the extraction of the text of the words spoken by the user 230 may be performed continuously for all words, or may start when the user utters a word within a predetermined period after a trigger keyword. For example, the trigger keyword may be when the user utters "hello" followed by the character's name. For example, if the name of the character 19051 is "Koto," then "Hello, Koto!" may be the trigger keyword.
[0062] The character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230 and transmits the prompt to the large-scale language model server 19001 using an API. Here, the prompt may be metadata containing information written in a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON. The prompt stores text information in a natural language as a main message. The prompts transmitted from the AI response output device 10010 to the large-scale language model server 19001 include setting prompts that store instructions such as initial settings, and user prompts that reflect instructions from the user. Type identification information identifying whether the prompt is a setting prompt or a user prompt may be stored in a portion of the prompt other than the main message. When the character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230, it creates the user prompt and sends it to the large-scale language model server 19001.
[0063] Next, the large-scale language model of the artificial intelligence of the large-scale language model server 19001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response including natural language text information based on the result. The large-scale language model server 19001 uses an API to send the response to the artificial intelligence response output device 10010. The response stores the natural language text information as a main message. Here, the response may be metadata storing information written in the same format as the instruction (e.g., a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON). If the response uses the same format as the instruction, type identification information may be stored outside the main message to indicate that the information is a different type from the initial setting instruction and the user instruction. For example, information indicating that the response is a response from the large-scale language model may be stored.
[0064] Next, the AI response output device 10010 receives a response from the large-scale language model server 19001 and extracts the natural language text information stored as the main message in the response. The character operation program of the AI response output device 10010 uses speech synthesis technology to generate a natural language voice as a response to the user based on the natural language text information extracted from the response, and outputs the voice from the speaker, i.e., the voice output unit 1140, so that it sounds as if it were the voice of the character 19051. This process may also be referred to as the character's "speaking."
[0065] Conversation examples 1 to 5 in Fig. 2C show specific examples of response voices from character 19051 in response to words from user 230, which are generated by the above-described processing of AI response output device 10010 and large-scale language model server 19001. In this way, user 230 can converse with character 19051 as if it were a real person.
[0066] 2B or a system including the AI response output device 10010, there is no need to install a large-scale language model, which requires a huge amount of data and computational resources for learning, in the AI response output device 10010. In addition, the advanced natural language processing capabilities of the large-scale language model can be utilized via an API, and when a user speaks to a character, a more appropriate response can be given to the user, enabling a more appropriate conversation to be held.
[0067] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2D. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2D shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0068] FIG. 2D also shows the exchange of instructions and responses in chronological order, from the display setting instruction, the first round of user instructions and their responses, to the fourth round of user instructions and their responses.
[0069] As shown in FIG. 2D , the setting directive can be used to initially instruct the large-scale language model of the artificial intelligence of the large-scale language model server 19001, such as the name of the large-scale language model itself, the role to be played, and conversation characteristics. The user's name can also be understood as the initial setting. This allows the large-scale language model to generate responses from the first round onward while respecting the role. Then, to a user who hears the voice of the character 19051 based on the responses from the first round onward, the character 19051 will seem to have the setting and personality of the person described in the setting directive. Furthermore, the large-scale language model server 19001 according to this embodiment is equipped with a memory that stores the content of the conversation until the end of the series of conversations, and is configured to store a series of user directives and their responses and then generate responses. This allows for a conversation such as that shown in FIG. 2D to be realized.
[0070] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2E shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.
[0071] 2E shows an example of a case where, after the series of conversations shown in FIG. 2D has ended, the user 230 speaks to the character 19051 again to start a new conversation. In FIG. 2E, the exchange of instructions and responses is shown in chronological order, from the first round of user instructions and their responses to the third round of user instructions and their responses.
[0072] Here, "termination" of the "continuation of a series of conversations" refers to a process in which, when a predetermined condition is met, the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversation that the large-scale language model server 19001 has maintained while the series of conversations was continuing. An example of the predetermined condition is, for example, when the AI response output device 10010 issues an instruction to the large-scale language model server 19001 to "terminate" the "continuation of a series of conversations." Another example of the predetermined condition is, for example, when a predetermined time or more has passed since the AI response output device 10010 stopped sending instruction statements to the large-scale language model server 19001 regarding the series of conversations (timeout). Another example of the predetermined condition is when, after authentication processing has been performed in the connection between the AI response output device 10010 and the large-scale language model server 19001, the authentication processing is terminated due to factors such as communication disconnection or the AI response output device 10010 being powered off while exchanging the instruction statements and responses.
[0073] Note that when the "continuation of a series of conversations" "ends," the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversations that it had maintained while the series of conversations was continuing. Therefore, even though the conversation shown in FIG. 2E takes place after the series of conversations shown in FIG. 2D, the server's response to the user's instruction is a response with content that does not include any memory of the character's name set in the large-scale language model, the role to be played, the characteristics of the conversation, or the user's name, which were included in the setting instruction shown in FIG. 2D. Similarly, the conversation shown in FIG. 2E is a response with content that does not include any memory of the series of conversations shown in FIG. 2D. In other words, with the "end" of the "continuation of a series of conversations" shown in FIG. 2D, the conversation in FIG. 2E starts from an initialized state of the large-scale language model of the artificial intelligence of the large-scale language model server 19001.
[0074] This causes the user 230 to feel as if the character 19051 has lost its memory or is a completely different person. From the user 230's perspective, the character's response feels very strange, resulting in a feeling of loneliness and disappointment. Such behavior poses a problem in that it is not possible to ensure the consistency of the settings and memories of the character 19051, such as its name, role, conversational characteristics, and personality, displayed on the AI response output device 10010.
[0075] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2F. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2F shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.
[0076] FIG. 2F illustrates an example of a case in which the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations illustrated in FIG. 2D has ended. Unlike the process illustrated in FIG. 2E, in the process illustrated in FIG. 2F, when starting a new conversation, the AI response output device 10010 transmits a setting instruction as the first instruction to the large-scale language model server 19001. The setting instruction stores the same natural language text as the initial setting instruction of FIG. 2D. This may be referred to as a reset text. The setting instruction is followed by natural language text describing the history of past conversations. This may be referred to as a conversation history text. The AI response output device 10010 may record the history of past conversations as natural language text information in the storage unit 1170 while the series of conversations illustrated in FIG. 2D is continuing, linking the history of the conversations to information on the date and time of the conversations. If there are conversations on different dates, each conversation may be recorded linked to date and time information, and the conversation history may be accumulated. When generating a setting instruction for the first instruction of a later conversation, such as that shown in Figure 2F, the natural language text information of the conversation recorded in storage unit 1170 and the information on the date and time when the conversation took place can be read out and used to generate the setting instruction.
[0077] When natural language text information of past conversation history is used to generate the setting instruction sentence, the format can be determined freely to a certain extent because it is data to be sent to a large-scale language model, but as shown in Figure 2F, it is sufficient to prepare prefixes and suffixes in natural language, such as "I talked about the following on ____ day of ____ month," or "You talked about the following on ____ day of ____ month," and combine these with the natural language text information of the recorded conversation to generate the text of the setting instruction sentence. Also, information on the date and time of the conversation read from storage unit 1170 may be combined with the above-mentioned "____ day of ____ month" portion to form part of the text of the setting instruction sentence.
[0078] Even if user 230 speaks to character 19051 again to start a new conversation after a series of conversations has ended, by performing the above-described generation process and transmission process of the setting instruction sentence in Fig. 2F, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation. This is preferable because it is perceived by the user as ensuring the consistency of the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation.
[0079] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2G shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0080] Figure 2G shows an example of a series of conversations shown in Figure 2F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 2G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 2F, so repeated description is omitted.
[0081] As shown in the natural language text of the server response in the table of Figure 2F, by using the setting directive shown in Figure 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 reflects the settings and conversation history of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation. This is more preferable because it allows the user to recognize that the settings and memories of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation are more closely matched. Note that this allows the characters to be viewed as the same from the user's perspective, and may therefore be referred to as a pseudo-identity of the characters from the user's perspective.
[0082] Furthermore, from the user's perspective, they can share memories with the character, providing a more enjoyable character conversation experience.
[0083] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, FIG. 2H shows an example of the operation in which a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. A character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input input to the operation input unit 1107 or an operation detected by a touch operation input sensor of the display unit 10011.
[0084] In the example of FIG. 2H, in addition to character 19051 (named "Koto") used in the description of FIGS. 2A to 2G, character 19052 (named "Tom") and character 19053 (named "Necco") are shown. Character 19051 (named "Koto") and character 19052 (named "Tom") are human characters, and character 19053 (named "Necco") is a cat character. The display of characters displayed on display unit 10011 can be switched by switching to and displaying on display unit 10011 an image generated by rendering a character in a different virtual 3D space for each character. The processing method for realizing the display of a rendered image of a 3D model of each character is not particularly limited. Furthermore, depending on the character, a dynamic 2D image may be displayed.
[0085] Furthermore, it is preferable that the character operation program executed by control unit 1110 also changes the synthetic voice used for the "speech" of each character when switching the display of the characters displayed on display unit 10011. This can be achieved by storing synthetic voice data of voice tones associated with each character in storage unit 1170 in advance, and performing synthetic voice change processing when switching the display of the character.
[0086] In the example of Fig. 2H, the AI response output device 10010 is configured so that the user 230 can converse with any of the characters. In the AI response output device 10010 of Fig. 2H, each of these characters is set with a different role, name, conversational characteristics, personality, etc. Also, the memories of each character based on the conversation history are managed as different for each character.
[0087] Therefore, the AI response output device 10010 constructs a database shown in FIG. 2I in the storage unit 1170, and manages the character settings and the character conversation history using this database.
[0088] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Fig. 2I. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.
[0089] A character operation program executed by the control unit 1110 of the AI response output device 10010 constructs the database 19200 in, for example, the storage unit 1170. The character ID is an identification number that identifies each of multiple characters that can be displayed on the AI response output device 10010, and may be a natural number or may use alphabets, etc. The name is data of the name of each of multiple characters that can be displayed on the AI response output device 10010.
[0090] The initial setting instruction is text information in a natural language that explains the settings such as the role, name, conversational features, or personality of each of multiple characters that can be displayed on the AI response output device 10010. The initial setting instruction is natural language text information that is the main data of the setting instruction sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the description can be read as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0091] The conversation histories, which continue as conversation histories 1, 2, ..., are records of conversations between each character and the user, and are recorded separately for each character. The conversation histories will be included in the natural language text information, which is the main data of the setting instruction statement sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the conversation histories be readable as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0092] When the character displayed on the display unit 10011 of the AI response output device 10010 is switched, the character operation program executed by the control unit 1110 of the AI response output device 10010 uses the database 19200 of Figure 2I to select and switch the initial setting instruction statement and conversation history used for the natural language text information that is the main data of the setting instruction statement transmitted from the AI response output device 10010 to the large-scale language model server 19001 so as to correspond to the character displayed on the display unit 10011 of the AI response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the conversation history in the conversation history area of the database 19200 of Figure 2I that corresponds to the character displayed on the display unit 10011.
[0093] By using the database 19200 in this manner, the character operation program executed by the control unit 1110 of the AI response output device 10010 establishes a conversation between the user 230 and the character using character utterances that utilize responses from the same AI large-scale language model of the same large-scale language model server 19001. From the user's perspective, the uniqueness of each character's personality and other settings is preserved, and it appears as if each character's unique conversational memories continue. This is more preferable because it appears as if the identity of each character's settings and memories, such as their role, name, conversational characteristics, or personality, from the time of the previous conversation has been more consistently maintained. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.
[0094] Therefore, even when the AI response output device 10010 is configured to switch the character displayed on the display unit 10011 from among multiple character candidates, the operation using the database 19200 described above reduces the sense of discomfort felt by the user when conversing with each character, and the user can share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.
[0095] Note that if the initial setting instructions for multiple characters cannot be edited by the user, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained close to the intentions of the provider of the AI response output device 10010 or the creator of the character content. Alternatively, the initial setting instructions for each character may be edited by the user in response to input from the operation input unit 1107 or the like. In this case, the character's role, name, conversational characteristics, or personality can be set to a preferred setting, allowing the user to converse with a character that they have individually set. In this case, the character's 3D model, its rendered image, and the type of synthesized voice for the character may be replaced accordingly.
[0096] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2J. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service at a lower cost using the character conversation device based on the artificial intelligence response output device 10010 and the character conversation system based on the artificial intelligence response output device 10010 and the large-scale language model server 19001 will be described.
[0097] As explained in Figure 2B, it is extremely resource inefficient to train a large-scale language model at this level of artificial intelligence by limiting it to a specific application. Therefore, it is more resource efficient to perform large-scale training to generate a foundation model that can be applied to a variety of applications, and then use it on various devices via an API (Application Programming Interface). In such cases, providers of large-scale language models often recover the costs incurred in training the large-scale language model from device users as API usage fees. In natural language models, API usage fees are often charged based on the number of tokens, which are units of words that separate sentences, processed.
[0098] Therefore, in the artificial intelligence response output device 10010 of the second embodiment of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010 and a character conversation service using a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.
[0099] For example, by using the processing and configuration of Examples 1 to 3 shown in the table of Figure 2J, it is technically possible to reduce the number of tokens in the natural language text information transmitted between the AI response output device 10010 and the large-scale language model server 19001 using an API.
[0100] Example 1 is an example of a method for reducing the number of tokens in the conversation history text stored in the API setting directive and transmitted, in which the conversation history text is shortened using a document summarization process to reduce the number of tokens. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. The text summarization may be performed at the start of the next conversation, but it is more time-efficient to perform it at the end of the "series of conversations."
[0101] Furthermore, the text summarization process may be requested from the large-scale language model itself of the large-scale language model server 19001. However, in this case, the effect of saving the number of tokens is small. Therefore, for example, if the second server 19002 provides text summarization process for natural language via an API at a lower cost than the large-scale language model of the large-scale language model server 19001, the text summarization process can be requested from the second server 19002 via the API, and the text summary of the conversation history can be stored in a setting instruction statement for the large-scale language model server 19001 and transmitted.
[0102] Furthermore, if it is only text summarization processing, it can be performed on the terminal side, and text summarization can be performed by the control unit 1110 executing a document summarization program loaded in the memory 1109 of the AI response output device 10010. In this case, the effect of saving the number of tokens is high. Furthermore, even if the conversation history becomes long, if an upper limit on the number of characters after summary is specified in the text summarization processing, the upper limit on the text length of the conversation history is determined, so that an upper limit on the number of tokens can be set and tokens can be saved.
[0103] In addition, since the text information of the character's initial settings, such as the character's role, name, conversational characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information of the character's initial setting instructions and reduce the number of tokens of the text information in the conversation history.
[0104] The processing described in Example 1 may be performed by the character movement program executed by the control unit 1110 controlling each unit.
[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored in the API setting instruction and transmitted. For example, the number of tokens is reduced by deleting the oldest conversation history among the conversation histories with characters recorded in the storage unit 1170. Specifying an upper limit on the number of characters in the conversation history determines the upper limit on the length of the conversation history text, thereby setting an upper limit on the number of tokens and enabling token conservation. Alternatively, a predetermined period of the conversation history may be specified and conversation history exceeding that period may be deleted. This also enables token conservation. Note that, in Example 2, the text information for the character's initial setting, such as the character's role, name, conversation characteristics, or personality, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the text information in the character's initial setting instruction and reduce the number of tokens in the text information for the conversation history.
[0106] The processing described in Example 2 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0107] Example 3 is a method of reducing the number of tokens by reducing the frequency of sending setting instruction sentences using an API. Specifically, even after the device is powered on, after the displayed character is switched, or after the video settings and synthetic voice settings of the displayed character are completed, setting instruction sentences are not sent in advance, and only when the control unit 1110 determines that natural language text information included in the user's speech picked up by the microphone 1139 is text information that should use a large-scale language model of artificial intelligence is the setting instruction sentence sent to the large-scale language model server 19001, thereby reducing the frequency of sending setting instruction sentences to the large-scale language model server 19001 and reducing the number of tokens.
[0108] Specifically, for example, after the device is powered on or after an operation input to switch the displayed character is made, character 19051 (named "Koto") is displayed on display unit 10011 as shown in FIG. 2H by display processing of display unit 10011 controlled by a character operation program executed by control unit 1110. At this time, for example, if a synthetic voice for the character 19051 to appear is stored and prepared in storage unit 1170 or the like, a synthetic voice for the character's appearance such as "Good morning, I'm Koto," "Hello, I'm Koto," or "Good evening, I'm Koto" may be output from the speaker, which is audio output unit 1140. At this time, the image of character 19051 has already been set as the image of the character to be displayed on display unit 10011, and the synthetic voice output from the speaker, which is audio output unit 1140, is set to the synthetic voice corresponding to character 19051.
[0109] Here, the inference process of the large-scale language model of the artificial intelligence in the large-scale language model server 19001 already described also takes time as the instruction becomes longer. In particular, if the setting instruction includes text information related to past conversation history, the number of tokens in the instruction increases, resulting in a particularly long inference process time. The setting instruction itself and its response are not output to the user 230. Based on the response to the user instruction following the setting instruction, a synthesized voice is output as the character's "utterance" from the speaker, which is the voice output unit 1140. In this case, it may seem preferable at first glance to send the setting instruction from the AI response output device 10010 to the large-scale language model server 19001 in advance to complete the inference process of the large-scale language model for the setting instruction in advance, as this would result in a faster response in the output of synthesized voice of the character's 19051's "utterance" after the user 230 speaks to the character 19051.
[0110] However, if the setting instruction is transmitted to the large-scale language model server 19001 before the user 230 speaks and the inference processing of the large-scale language model for the setting instruction is completed in advance, for example, the user 230 may turn off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or the user 230 may switch the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. In this case, the number of tokens processed in the inference processing of the large-scale language model by transmitting the setting instruction to the large-scale language model server 19001 in advance becomes the number of processed tokens for which the usage fee is wasted. This hinders the provision of a character conversation device using the AI response output device 10010 and a character conversation service using a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 at a lower cost to users.
[0111] Therefore, after the AI response output device 10010 is powered on (ON) or after an operation input to switch the displayed character is made, it is desirable that the AI response output device 10010, under the control of the character operation program executed by the control unit 1110, sets the image of character 19051 as the image of the character to be displayed on the display unit 10011, and sets the synthetic voice output from the speaker, which is the audio output unit 1140, to the synthetic voice corresponding to character 19051, continue not to send a setting instruction statement to the large-scale language model server 19001 until the user 230 recognizes that he or she is speaking to character 19051.
[0112] Here, the point in time at which it is recognized that user 230 is speaking to character 19051 may be, for example, the point in time at which the trigger keyword described in Fig. 2B is detected, or the point in time at which the text of the words spoken by user 230 is extracted. In this way, the number of processing tokens that waste usage fees can be reduced, and a character conversation service by a character conversation device using artificial intelligence response output device 10010 or a character conversation system using artificial intelligence response output device 10010 and large-scale language model server 19001 can be provided to users at a lower cost.
[0113] Furthermore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051 has passed, it is desirable to continue not sending setting instructions to the large-scale language model server 19001, for example, if the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Specifically, examples of preset keywords include keywords such as "try jumping" and "try dancing," which are keywords by which the user 230 requests the character 19051 to react, such as by animating the character 19051 to move or emitting synthetic voice. In this case, the character operation program executed by the control unit 1110 reads out motion data, animation video, and / or synthetic voice data corresponding to the reaction stored in the storage unit, corresponding to the character 19051, and / or synthetic voice data corresponding to the reaction, and uses these data to generate video to be displayed on the display unit 10011 and output synthetic voice from the speaker, which is the audio output unit 1140.
[0114] Such processing does not necessarily require inference processing of the large-scale language model of the large-scale language model server 19001. After the processing, if the user 230 turns off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or if the user 230 switches the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, if the setting instruction statement is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be the number of processing tokens that unnecessarily consumes the usage fee.
[0115] Therefore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051, it is desirable to continue not sending the setting instruction statement to the large-scale language model server 19001 until it is determined whether or not the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Only when it is determined that inference processing of a large-scale language model is required is it desirable to send the setting instruction statement to the large-scale language model server 19001 and proceed with the inference processing of the large-scale language model.
[0116] The processing described in Example 3 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0117] According to the method for reducing (saving) the number of processing tokens for a large-scale language model using the examples of Figure 2J described above, a character conversation device using the AI response output device 10010 and a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.
[0118] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2K. The example of FIG. 2K shows an example of displaying a response from a large-scale language model to an instruction statement from a user, which is described in each of FIGS. 2A to 2J, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063, which is a response from the large-scale language model, together with a video of a character 19051 on the display unit 10011. The text 10063, which is a response from the large-scale language model, may be displayed superimposed on top of the video of the character 19051, as shown in FIG. 2K. Alternatively, the text 10063, which is a response from the large-scale language model, may be displayed together with the video of the character 19051 without being superimposed on the video of the character 19051.
[0119] The display in Figure 2K is an example, but for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by audio.
[0120] Therefore, in this case, the control unit 1110 may perform control to start a display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, as shown in Fig. 2K. In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more suitably use the character conversation device (artificial intelligence response output device 10010). Note that the user 230 may be configured to manually switch ON / OFF the display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011.
[0121] Next, an example of a fixed response phrase database (fixed response phrase DB) in the character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in FIGS. 2H and 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition numbers and condition contents are the same as those in FIG. 1C. In the example of FIG. 2L, individual fixed response phrases are set for each of the multiple characters in response to these conditions. For example, fixed response phrases for each condition are stored for each of the three characters, Character 1: Koto, Character 2: Tom, and Character 3: Necco, as described in FIGS. 2H and 2I. The output control of fixed response phrases is the same as that in FIG. 1C, and therefore will not be repeated.
[0122] In the example of FIG. 2L, the control unit 1110 selects a corresponding fixed response phrase from a fixed response phrase database (fixed response phrase DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses the selected fixed response phrase for output control as a response uttered by the character. For example, in the example of the fixed response phrase database (fixed response phrase DB) in FIG. 2L, even under the same conditions, the fixed response phrases are changed to expressions or contents that correspond to the individuality of the characters. This allows the character conversation device (artificial intelligence response output device 10010) to provide the user with conversations that correspond to the individuality of the displayed characters. The user can feel that each character has a more consistent personality. This makes it possible to realize a character conversation device (artificial intelligence response output device 10010) that gives multiple characters a more realistic presence.
[0123] The fixed response phrase database (fixed response phrase DB) of FIG. 2L described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 2L may also be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in each server. In this way, even if the AI response output device 10010 does not have a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0124] The character conversation device and character conversation system according to the second embodiment described above can reduce the sense of discomfort felt by the user from the conversation with the character displayed on the AI response output device 10010. Furthermore, the character conversation device and character conversation system according to the second embodiment can provide the character conversation service to the user at a lower cost.
[0125] In the above description of the second embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 19001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B, and the large-scale language model held by the local LLM processing unit 10028 may be used instead of the large-scale language model held by the large-scale language model server 19001. In this case, in the above description of the second embodiment, the large-scale language model held by the large-scale language model server 19001 may be read as the large-scale language model held by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).
[0126] In this case, too, it is possible to further reduce the sense of discomfort felt by the user from conversations with characters displayed on the AI response output device 10010. Note that when a large-scale language model held by the local LLM processing unit 10028 is used instead of the large-scale language model held by the large-scale language model server 19001, there is less need to consider usage fees according to the number of processed tokens, but by reducing the number of processed tokens even for the large-scale language model held by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power required for inference. In this case, it is possible to provide users with a character conversation service that consumes less power.
[0127] In the above description of the second embodiment, an example has been described in which a conversation history with a character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). Alternatively, a conversation history with a character may be recorded and stored in a second server 19002 or another cloud server connected to the Internet 19000. In this case, when a user and a character start a new conversation, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or another cloud server, acquires (downloads) a past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create an instruction for a large-scale language model. Specific methods for using past conversation history to create an instruction for a large-scale language model are as described in the figures of the second embodiment, and therefore repeated description will be omitted.
[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character's conversation history up to a predetermined point in time, such as every time a conversation between the user and the character takes place or when the conversation between the user and the character ends, to the second server 19002 or another cloud server. That is, the character conversation device (artificial intelligence response output device 10010) may upload the conversation history with the character to the second server 19002 or another cloud server at a predetermined timing, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or another cloud server and use it to generate an instruction sentence for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) that the user used the previous day and the character conversation device (artificial intelligence response output device 10010) that the user is about to use are different individual devices and can display the same character, and when the user has multiple conversations with the same character between the different individual devices at different times, it is possible to realize a conversation in which the character's memories are pseudo-carried over from the previous conversation, which is more convenient for the user.
[0129] The process described above in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with a character to the second server 19002 or another cloud server, thereby pseudo-taking over the character's memory, is also effective when handling the database 19200 including the conversation histories of multiple characters described in Figures 2H and 2I. In other words, if the database 19200 described in Figure 2I is configured to be uploaded and downloaded to the second server 19002 or another cloud server, not only for one character but for multiple characters, and between different individual devices, when a user has multiple conversations with each of the multiple characters at different times, it is possible to realize a conversation in which the memory of each character is pseudo-taken over from the previous conversation, which is more convenient for the user.
[0130] Example 3 Next, the third embodiment of the present invention is an improvement of the character conversation device (artificial intelligence response output device 10010) and the character conversation system explained in the drawings of the second embodiment. In this embodiment, differences from the second embodiment will be explained, and repeated explanations of the same configurations as those of the second embodiment will be omitted.
[0131] As in Example 2, the character in Example 3 can provide the user with the service of a large-scale language model, which is an artificial intelligence, and can be of assistance to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this example may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0132] An example of a character conversation device and a character conversation system according to a third embodiment of the present invention will be described with reference to Fig. 3A. The character conversation system of the third embodiment is provided with a large-scale language model server 20001 instead of the large-scale language model server 19001 in Fig. 2A, and is connected to the Internet 19000.
[0133] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence, and is a multimodal large-scale language model artificial intelligence that can process not only the natural language text information that the large-scale language model server 19001 could process, but also types of information other than natural language text information.
[0134] Moreover, the AI response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (AI response output device 10010) of the second embodiment, as an example.
[0135] In the third embodiment as well, the AI response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the Internet 19000 using an API.
[0136] The character conversation system of the third embodiment includes a mobile information processing terminal 20010 used by a user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.
[0137] 3B, an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011, which is a touch operation input panel, a control unit 20012, an external power supply input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.
[0138] The display panel 20011 is equipped with a touch operation input sensor and can accept touch operation input by the finger of the user 230. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel. The display panel 20011 may also be referred to as a display unit.
[0139] The communication unit 20020 may be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and a memory, and the control unit controls the display panel 20011 and the communication unit 20020. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000. This allows the mobile information processing terminal 20010 to communicate with various servers connected to the Internet 19000.
[0140] The power supply 20014 converts AC current input from the outside via the external power supply input interface 20013 into DC current and supplies the required DC current to each unit of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Furthermore, the secondary battery 20015 supplies power to each unit requiring power via the external power supply input interface 20013 when power is not supplied from the outside.
[0141] The video signal input unit 20023 is connected to an external video output device and inputs video data. The video signal input unit 20023 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be various USB interfaces.
[0142] The audio signal input unit 20024 is connected to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 20024 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface in which a terminal and a cable are integrated.
[0143] The audio output unit 20021 can output audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 can also output audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured with a speaker. The audio output unit 20021 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard.
[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may pick up a person's voice, such as a user's voice, and the control unit 20012, which will be described later, may perform voice recognition processing on the generated audio signal to obtain text information from the audio signal.
[0145] The imaging unit 20025 is a camera having an image sensor. A camera may be provided on the front side of the mobile information processing terminal 20010, on the display panel 20011 side, or on the back side of the display panel 20011 side. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.
[0146] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 20016 before product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices, external servers, etc., via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may also be output to external devices, external servers, etc., via the communication unit 20020.
[0147] The video control unit 20017 performs various controls related to the video signal input to the display panel 20011. The video control unit 20017 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be referred to as a video processing unit or image processing unit. The video control unit 20017 performs video switching control, such as determining which video signal to input to the display panel 20011 between the video signal to be stored in the memory 20026 and the video signal (video data) input to the video signal input unit 20023. The video control unit 20017 may also perform image processing control on the video signal input from the video signal input unit 20023 and the video signal to be stored in the memory 20026. Examples of image processing include scaling processing that enlarges, reduces, or transforms an image, brightness adjustment processing that changes the brightness, contrast adjustment processing that changes the contrast curve of the image, and Retinex processing that decomposes an image into light components and changes the weighting of each component.
[0148] The attitude sensor 20018 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each unit connected thereto.
[0149] The nonvolatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the nonvolatile memory 20027 includes, for example, data for various operations to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 20026 stores video data to be displayed on the display panel 20011, data for controlling the device, etc. The control unit 20012 may read various software from the storage unit 20016 and expand and store it in the memory 20026.
[0150] The control unit 20012 controls the operation of each unit connected to it. The control unit 20012 may also work in conjunction with a program stored in the memory 20026 to perform arithmetic processing based on information acquired from each unit within the mobile information processing terminal 20010.
[0151] Next, an example of the operation of the character conversation device (AI response output device 10010) according to the third embodiment of the present invention will be described with reference to Figure 3C. This can also be said to be an example of the operation of a character conversation system including the AI response output device 10010 and the large-scale language model server 20001. In the third embodiment as well, the character conversation device (AI response output device 10010) loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0152] In the second embodiment, the actions performed by the user 230 with respect to the character conversation device (artificial intelligence response output device 10010) were mainly calls made by the user 230 using his / her voice. The character conversation device (artificial intelligence response output device 10010) of the second embodiment performed a series of operations starting from the process of collecting the voice of the user 230 with a microphone. In contrast, the character conversation device (artificial intelligence response output device 10010) of the third embodiment is also capable of performing the series of operations performed by the character conversation device (artificial intelligence response output device 10010) starting from the process of collecting the voice of the user 230 with a microphone, as described in the second embodiment. In addition, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the user 230 can perform an action with respect to the character conversation device (artificial intelligence response output device 10010) by user operation via the operation input unit 1107 of FIG. 1B. Here, examples of the operation input unit 1107 of FIG. 1B include a mouse, a keyboard, and a touch panel.
[0153] Furthermore, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by performing a touch operation by the user that can be detected by the touch operation input sensor of the display unit 10011 of Figure 1B.
[0154] In addition, the user 230 can operate the mobile information processing terminal 20010 and communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting the user's 230 operation input to the character conversation device (artificial intelligence response output device 10010).
[0155] Alternatively, an information storage image such as a two-dimensional code storing information that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as a two-dimensional code captured by the imaging unit 1180, and obtain the information. Alternatively, an image that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180, and obtain the results of the image recognition processing.
[0156] As described above, the character conversation device (artificial intelligence response output device 10010) of the third embodiment has a greater variety of actions that the user 230 can take toward the character conversation device (artificial intelligence response output device 10010) than the character conversation device (artificial intelligence response output device 10010) described in the second embodiment. As a result, the character conversation device (artificial intelligence response output device 10010) of the third embodiment can acquire the results of actions taken by the user 230 other than the user's voice, and generate an instruction sentence (prompt) to be sent to the large-scale language model server 20001 based on the results. As a result, the instruction sentence to be sent to the large-scale language model server 20001 can more preferably include information of a type other than text information in a natural language extracted from the user's voice. Examples of information of a type other than text information in a natural language extracted from the user's voice include images, videos, and sounds.
[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment uses an API to send an instruction to the large-scale language model server 20001. In this embodiment, the instruction may be metadata containing information written in a notation using tags, such as the markup format of a markup language, a notation using predetermined symbols, such as Markdown, or an object notation of a predetermined script, such as JSON. In this embodiment, the instruction may be classified into two types: a setting instruction that stores instructions, such as initial settings, and a user instruction that reflects instructions from a user. Type identification information identifying whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. In this case, the instruction includes text information in a natural language as the main message. Furthermore, in this embodiment, the main message of the instruction may include, in addition to the text information in natural language, a non-natural language information source, such as an image, video, or audio, as information of a type other than the natural language text information. A specific method for including a non-natural language information source in the instruction will be described later.
[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. The large-scale language model server 20001 receives an instruction from a character conversation device (artificial intelligence response output device 10010). Based on the instruction, the multimodal large-scale language model performs inference and generates a response including natural language text information that is the result of the inference. Here, because the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.
[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or sounds stored as the main message in the response. The character operation program of the character conversation device (artificial intelligence response output device 10010) may use a voice synthesis technology to generate a natural language voice as a reply to the user based on the natural language text information extracted from the response, and output the voice from the voice output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.
[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display characters in a natural language that serve as a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on text information in a natural language extracted from the above-mentioned response. At this time, the characters may be displayed together with the character 19051, may be displayed superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the information of the image of the non-natural language information source extracted from the above-mentioned response. At this time, the image may be displayed together with the character 19051, may be superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the video information of the non-natural language information source extracted from the above-mentioned response. At this time, the video may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed in place of the video of the character 19051. These specific processes may be executed by the video control unit 1160.
[0163] In addition, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output a voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140, which is a speaker.
[0164] As described above, the character conversation device (artificial intelligence response output device 10010) of FIG. 3C or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 does not require the large-scale language model itself, which requires vast amounts of data and computational resources for learning, to be installed in the character conversation device (artificial intelligence response output device 10010). Furthermore, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via an API. In response to a user's action toward a character, a response based on a non-natural language information source can be provided in addition to a response based on natural language text, enabling a more appropriate conversation to be held.
[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3D . This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3D shows examples of non-natural language information sources, such as natural language text and images, of the main message of an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001, and examples of non-natural language information sources, such as natural language text and images, of the main message of a server response that is the response. In this embodiment, the non-natural language information source can be an image, video, audio, or the like, but FIG. 3D shows an example of an image as the non-natural language information source.
[0166] 3D also shows an exchange of instructions and responses in chronological order, from a setting instruction, a first round of user instructions and their responses, to a second round of user instructions and their responses. The instructions and responses shown in FIG. 3D include non-natural language information source 20061 and non-natural language information source 20062, which were not shown in FIG. 2D of Example 2. In the example of FIG. 3D, both non-natural language information source 20061 and non-natural language information source 20062 are images.
[0167] Here, for ease of explanation, FIG. 3D shows an image of the non-natural language information source 20061 pasted into the instruction. However, there are multiple methods for transmitting or specifying data from the non-natural language information source 20061 in an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these multiple methods or switch between them. An example of each method will be described below.
[0168] The first method for transmitting or specifying non-natural language information source data in a directive is used, for example, when the non-natural language information source to be specified is a non-natural language information source located in a location such as a server connected to a network such as the Internet. A specific example of the first method is to use information such as tags and symbols in the directive to specify a non-natural language information source file located on a network such as the Internet by using location information (such as a URL) on the network such as the Internet and a file name.
[0169] For example, it is a tag that specifies an image in a markup language. <img src=""****”"> You can also specify an image on a network such as the Internet by writing the location information and file name information of the image file in the **** part using the tag. <video src=""****”">You can also specify a video file on a network such as the Internet by writing the location and file name information of the video file in the **** part using the tag. <audio src=""****”">By using the above and entering the location information and file name information of the audio file in the **** part, audio that exists on a network such as the Internet can be specified. Furthermore, if the notation is JSON, an image that exists on a network such as the Internet can be specified by preparing a key such as img_src and entering the location information and file name information of the image file as the value. For video files and audio files, it is sufficient to prepare the respective keys and values. This specific example of a format is just one example, and other unique formats may also be used. In either case, information that specifies the location information and file name information of the non-natural language information source file can be stored in the directive.
[0170] When information specifying the location information and file name information of a non-natural language information source file is stored in a directive, as in the first method, the directive itself does not need to store the data of the non-natural language information source file itself. Therefore, the amount of data in the directive can be reduced. In the first method, the large-scale language model server 20001 that receives a directive specifying non-natural language information source data simply uses the location information and file name information of the non-natural language information source file stored in the directive to obtain the non-natural language information source file located in a location such as a server connected to a network such as the Internet.
[0171] Here, how location information and file name information are input when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence using the first method will be described. In FIG. 3C, it has been explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are increased beyond the voice of the user 230 compared to the second embodiment. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, and the like, by user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in FIG. 1B.
[0172] Furthermore, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a web browser program and display a GUI of the web browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the web browser program may be accepted by a user operation via the operation input unit 1107 (e.g., a mouse, keyboard, or touch panel) or a user touch operation detectable by a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, video, or audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction. In this case, the web browser program may acquire location information and file name information of the non-natural language information source data and pass them to the character operation program.
[0173] Furthermore, user 230 may operate mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from mobile information processing terminal 20010, thereby inputting location information such as a URL for specifying non-natural language information source data to the character conversation device (artificial intelligence response output device 10010). Alternatively, location information such as a URL for specifying non-natural language information source data, file name information, and the like may be input in a manner such as displaying an information storage image such as a two-dimensional code on display panel 20011 of mobile information processing terminal 20010, performing image recognition processing on an image captured by imaging unit 1180 of character conversation device (artificial intelligence response output device 10010), and acquiring the results of the image recognition processing, as described in FIG.
[0174] Note that the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file is previously present in a location such as a server connected to a network such as the Internet. For example, if non-natural language information source data such as images, videos, or audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) is to be included in the instruction statement, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to the second server 19002 via the Internet 19000, and include in the instruction statement the location information on the Internet (such as a so-called URL) and file name of the non-natural language information source data on the uploaded second server 19002. In this case, the second server 19002 functions as a so-called intermediate server.
[0175] Similarly, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in an instruction statement, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the location information on the Internet (such as a URL) and the file name of the non-natural language information source data on the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include the acquired location information on the Internet (such as a URL) and the file name of the non-natural language information source data uploaded to the second server 19002 in the instruction statement.
[0176] Furthermore, a media server may be constructed within the character conversation device (artificial intelligence response output device 10010) so that the character operation program of the character conversation device (artificial intelligence response output device 10010) can cooperate with memory 1109 and storage unit 1170 to be accessible from other servers via the Internet 19000. In this case, when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction statement using the first method, the character conversation device (artificial intelligence response output device 10010) may store location information on the Internet (such as a URL) indicating the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data in the instruction statement.
[0177] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction sentence is, for example, a method in which the non-natural language information source data itself is simply stored (attached) to the instruction sentence (prompt) and transmitted. Generally, non-natural language information source data such as images, videos, and audio has a larger data volume than text information in natural language. Therefore, in this case, the data volume of the instruction sentence (prompt) itself is larger than that in the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) temporarily stores the non-natural language information source data to be stored (attached) to the instruction sentence (prompt) in the memory 1109, and when transmitting the instruction sentence (prompt), stores (attaches) the non-natural language information source data in the memory 1109 via the communication unit 1132 and outputs the data to the large-scale language model server 20001. The non-natural language information source data itself that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read from the storage unit 1170 and stored in memory 1109.
[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data using instruction sentences.
[0179] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. Therefore, in the first pass of the user instruction shown in the example of Figure 3D, it can obtain images of the swimming pool and poolside, which are non-natural language information source 20061, and text information in natural language, and output the text information in natural language as a response to the first pass of the user instruction as an inference result, as shown in the figure.
[0180] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, the large-scale language model server 20001 can include in the response a non-natural language information source 20062 generated by inference of the multimodal large-scale language model and send it to the character conversation device (artificial intelligence response output device 10010), as in the response to the second round of user instructions shown in the example of FIG. 3D. In FIG. 3D, the non-natural language information source 20062 is shown as an example of an image in which a circle image is added to the image of a swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in FIG. 3D, but may be video or audio.
[0181] When the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, a method similar to the first or second method in which the character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence can be used.
[0182] Specifically, as a method similar to the first method described above, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The non-natural language information source 20062, such as an image, video, or audio, itself may be stored in the large-scale language model server 20001, or the non-natural language information source 20062 may be transferred to and stored in a second server 19002 functioning as an intermediate server. In either case, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The character conversation device (artificial intelligence response output device 10010) that receives the response may access the large-scale language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction to acquire the non-natural language information source 20062.
[0183] Specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction sentence and use it for various outputs to the user 230.
[0184] According to the operation of the character conversation device (artificial intelligence response output device 10010) and the character conversation system of the third embodiment described above with reference to Fig. 3D, instructions and responses are sent and received between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230 to realize a conversation using non-natural language information such as images, videos, and sounds. This makes it possible to realize a more advanced and natural conversation as shown in each message in Fig. 3D.
[0185] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3E shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0186] 3E shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in FIG. 3D has ended. In the example of FIG. 3E, processing using the conversation history as described in FIGS. 2F, 2G, and 2I of Example 2 is not performed. Therefore, like FIG. 2E of Example 2, FIG. 3E shows a response with content that does not remember at all the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, the user's name, the conversation history, and the like that were included in the setting directive.
[0187] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3F. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3F shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.
[0188] Figure 3F shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in Figure 3D has ended. Here, Figure 3F shows an example in which the method of storing a message explaining the history of past conversations in the setting instruction sentence, which was explained in Figure 2F of the second embodiment, is also applied to the character conversation device (artificial intelligence response output device 10010) of the third embodiment. Specifically, in Figure 3F, the message that is the content of the setting instruction sentence in Figure 3D is stored as a reset message, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.
[0189] The large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. Therefore, non-natural language information source data may have been transmitted or specified in past instructions and responses. Therefore, in the example of Figure 3F, the conversation history message reflects not only the natural language text information in past instructions and responses, but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method for transmitting or specifying non-natural language information source data in the instruction in Figure 3F is similar to the transmission or specification of non-natural language information source data described in Figure 3D, and therefore a repeated description will be omitted.
[0190] In the example of Figure 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself in the instruction statement, and storing (attaching) the non-natural language information source data in the instruction statement without storing (attaching) the non-natural language information source data in the instruction statement. This also applies to the instruction statement of Figure 3F.
[0191] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3G shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0192] Figure 3G shows an example of a series of conversations shown in Figure 3F, from the first user instruction and its response following the first setting instruction to the third user instruction and its response. Figure 3G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 3F, so repeated description is omitted.
[0193] As described above, even when using the large-scale language model server 20001 having a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information according to the third embodiment, even when the user 230 speaks to the character 19051 again to start a new conversation after a series of conversations has ended, if the generation and transmission processes of the setting instruction sentence in Fig. 3F are performed, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation, as shown in Fig. 3G. This is preferable because it allows the user to recognize that the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation are more closely matched.
[0194] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, FIG. 3H uses the example described in FIG. 2H of the second embodiment for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.
[0195] Furthermore, database 20200 for managing character settings and character conversation history, shown in Figure 3H, has the same format as database 19200 shown in Figure 2I of Example 2, and in Figure 3H, only the differences from database 19200 shown in Figure 2I will be explained. Also, the content of the character "Koto" in the database will be explained, and the content of other characters will be omitted.
[0196] As described above, the large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information. Therefore, both the instruction statements from the character conversation device (artificial intelligence response output device 10010) and the responses from the large-scale language model server 20001 include not only natural language text information but also the transmission or specification of non-natural language information source data. Therefore, the database 20200 shown in FIG. 3H records, in the conversation history data, not only the natural language text information included in these instruction statements and responses, but also the transmission or specification of non-natural language information source data. The specific method for transmitting or specifying non-natural language information source data in the conversation history recording is the same as the transmission or specification of non-natural language information source data described in FIG. 3D, and therefore a repeated description will be omitted.
[0197] In the example of FIG. 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself to the instruction statement and not storing (attaching) the non-natural language information source data to the instruction statement. This is also true for the conversation history of FIG. 3H. However, in the conversation history of FIG. 3H, if the method of specifying non-natural language information source data is to specify location information and file name information of a non-natural language information source file on a server on a network such as the Internet (such as the second server 19002 functioning as an intermediate server or another cloud server), there is a possibility that the non-natural language information source file on the server may be deleted over a long period of time in the conversation history. In this case, even if the location information and file name information are used, the non-natural language information source file cannot be retrieved at a later date, and information in the conversation record may be lost.
[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts the instruction sentence and the response message into a conversation history and records them, it can use the location information and file name information to obtain the non-natural language information source file itself specified in the instruction sentence and the response from a server on the network and store it in the storage unit 1170. Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) can rewrite the location information and file name of the non-natural language information source file to Internet location information (such as a URL) indicating the location of the non-natural language information source file within the media server built within the character conversation device (artificial intelligence response output device 10010), and then record the non-natural language information source in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself deletes the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the conversation record information, which is more suitable for preserving the conversation record.
[0199] By using the database of Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between multiple character candidates to display on the display unit 10011, the user will feel less uncomfortable when conversing with each character, and will be able to share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.This effect can be achieved as shown in Figure 2I of Example 2, and can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information.
[0200] In addition, in the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, the large-scale language model server 20001 uses a multimodal large-scale language model artificial intelligence that can process not only natural language text information but also non-natural language information other than natural language text information.
[0201] Here, an API is used for communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001. In a multimodal large-scale language model, the API usage fee may be charged according to the number of processed natural language text information, which are units of words that separate sentences called tokens, as well as the amount of data from non-natural language information sources.
[0202] Therefore, in order to provide users with a character conversation service by the character conversation system according to this embodiment at a lower cost, the following modified example may be used.
[0203] In a first variation, the database conversation history record of FIG. 3H also records the transmission or specification of non-natural language information source data. However, the character and the user exchange conversations about the natural language information source data using text information in natural language, and the content of those conversations is recorded as text information in natural language. Even if the transmission or specification of the natural language information source data is omitted from the database conversation history record of FIG. 3H, the conversation itself about the natural language information source data will still be recorded as text information in natural language to some extent. Therefore, if a certain amount of information reduction is acceptable, the transmission or specification of the natural language information source data may be omitted from the database conversation history record of FIG. 3H. In this case, the transmission or specification of the natural language information source data is also omitted from the conversation history message of the setting instruction sentence of FIG. 3F. This reduces the amount of data from non-natural language information sources communicated using the API.
[0204] Next, as a second modification, in the recording of the conversation history of the database in FIG. 3H, natural language text information explaining the content of the non-natural language information source data is recorded instead of transmitting the non-natural language information source data or recording specified information. The natural language text information explaining the content of the non-natural language information source data may be obtained, for example, by starting a conversation between the large-scale language model of the large-scale language model server 20001 and a character conversation device (artificial intelligence response output device 10010) separately from the conversation as a character, and having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the natural language text information may be obtained by having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit through a conversation with another large-scale language model of another server that is available at a lower cost than the large-scale language model of the large-scale language model server 20001. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be used as natural language text information explaining the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is a tag in a markup language. <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">This is the text information written in the **** part of the above.
[0205] Furthermore, in the case of JSON notation, in an object that is stored in association with the location information and file name information of the non-natural language information source data, which are keys and values indicating the location information of the non-natural language information source data, a key corresponding to the alternative text can be further linked and stored as a value that is the alternative text data itself.
[0206] In this case, too, the transmission of the natural language information source data or the recording of the specified information can be omitted in the recording of the conversation history in the database of Fig. 3H, and the transmission of the natural language information source data or the specified information can also be omitted from the conversation history message of the setting instruction sentence of Fig. 3F, thereby reducing the amount of data from non-natural language information sources communicated using the API.
[0207] Next, a third variation is an example in which, at the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data is not stored in the user instructions, but is replaced with natural language text information explaining the content of the non-natural language information source data. For example, in the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source 20061 may be replaced with the user instructions, and a description such as "This image is of a swimming pool, poolside seats, and parasols. There is water in the swimming pool. There are drinks on the table next to the seats" may be stored as natural language text information. In this case, the description may be obtained by having another large-scale language model on another server, which is more inexpensive to use than the large-scale language model on large-scale language model server 20001, explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the description may be obtained from a server of various other services that can obtain summaries and descriptions of non-natural language information source data such as images, videos, and audio. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be natural language text information that explains the content of the non-natural language information source data.
[0208] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3I. The example of FIG. 3I shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which is described in each of FIGS. 3A to 3H, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063 of natural language information source data, image 10064 of non-natural language information source data, and / or video 10065 of non-natural language information source data, which are responses from the large-scale language model, together with a video of a character 19051, on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed superimposed on top of the video of the character 19051, as shown in FIG. 3I.
[0209] Furthermore, text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on the video of character 19051. The display in FIG. 3I is an example, but for example, if user 230 adjusts the volume of the audio output of audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating operation input unit 1107 or the touch operation input sensor of display unit 10011, user 230 will not be able to hear the response from the large-scale language model by audio. Therefore, in this case, control unit 1110 may control to start a display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in FIG. 3I.
[0210] In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more preferably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the image of the character 19051, may be manually switched on / off by the user 230 operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. According to the display example of FIG. 3I, the character conversation device (artificial intelligence response output device 10010) that supports multimodal display can more preferably output a response from the large-scale language model.
[0211] According to the character conversation device and character conversation system of Example 3 described above, in addition to the effects of the character conversation device and character conversation system of Example 2, it is possible to provide users with a more advanced conversation experience that includes information in non-natural languages in addition to information in natural languages by using a multimodal large-scale language model. Furthermore, according to the character conversation device and character conversation system of Example 3, it is possible to provide users with a character conversation service at a lower cost.
[0212] In the above description of the third embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 20001 is used as the large-scale language model. In contrast to this, the character conversation device (artificial intelligence response output device 10010) may be equipped with the local LLM processing unit 10028 shown in FIG. 1B and may use the multimodal large-scale language model held by the local LLM processing unit 10028. In this case, the multimodal large-scale language model held by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model held by the large-scale language model server 20001.
[0213] In this case, in the above description of the third embodiment, the multimodal large-scale language model held by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model held by the local LLM processing unit 10028 of the character conversation device (the AI response output device 10010). In this case, too, a more sophisticated conversation experience that includes non-natural language information in addition to natural language information can be provided to the user using the multimodal large-scale language model. Note that, when the multimodal large-scale language model held by the local LLM processing unit 10028 is used instead of the multimodal large-scale language model held by the large-scale language model server 20001, there is less need to consider usage fees based on the number of processed tokens and the amount of data in the non-natural language information source. However, even with the multimodal large-scale language model held by the local LLM processing unit 10028, the number of processed tokens and the amount of data in the non-natural language information source can be reduced, thereby reducing the consumption of resources such as power required for inference. In this case, a character conversation service that consumes less power can be provided to the user.
[0214] The configuration described in Example 2 for uploading and downloading the conversation history with a character or the database data including the conversation history with a character to the second server 19002 or another cloud server can also be used in the example using a multimodal large-scale language model described in Example 3. In this case, too, when a user has multiple conversations with each of the multiple characters at different times between different individual devices for one character or multiple characters, it is possible to realize a conversation in which the memory of each character is pseudo-carried over from the previous conversation, which is more convenient for the user.
[0215] Example 4 Next, Example 4 of the present invention is an improvement of the AI response output device 10010, the character conversation device, or these systems described in the drawings of Example 2 or Example 3. In this example, differences from Example 2 or Example 3 will be described, and repeated explanations of the same configurations as those examples will be omitted.
[0216] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0217] An example of operation using a database in a character conversation device (artificial intelligence response output device 10010) according to a fourth embodiment of the present invention will be described with reference to Fig. 4A. The database according to the fourth embodiment shown in Fig. 4A is an extension of the database described with reference to Fig. 2I or Fig. 3I. Specifically, the database shown in Fig. 4A assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores initial setting instructions and conversation histories corresponding to each user and character in the database.
[0218] In the example of Fig. 4A, for user 1 with a user ID of 1, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored.
[0219] These initial setting instruction sentences and conversation history data are stored as separate data in different areas for each combination of user and character. For the sake of explanation, in Fig. 4A, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (AI response output device 10010) uses the initial setting instruction sentences and conversation history stored in different areas for each combination of user and character based on the user currently using (logging in to) the character conversation device (AI response output device 10010) or its system, thereby making it possible to more appropriately maintain the consistency of the character's personality and the continuity of memory for each different user.
[0220] Specifically, consider a situation in which User 1 has previously had a conversation with character Tom using a character conversation device (AI response output device 10010), and User 2 is unaware of that conversation, and then User 2 subsequently has a conversation with character Tom. In this case, if AI response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify users, the response output from AI response output device 10010 will be based on a conversation history that User 2 does not remember, and the conversation between User 2 and the character of AI response output device 10010 may become inconsistent.
[0221] In contrast, even in a similar situation, if the database shown in Figure 4A is used, the control unit 1110 of the character conversation device (AI response output device 10010) identifies users by ID, stores the initial setting instructions and conversation history in a different area for each user, and uses the initial setting instructions and conversation history stored in a different area for each user to generate an AI response. As a result, the initial setting instructions and conversation history used to generate an AI response for each user are based on the operation or conversation history of that user and are managed separately from the operation or conversation history of other users. This makes it possible to more appropriately maintain consistency in the conversation history between each user and each character of the AI response output device 10010.
[0222] The database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in the storage unit 1170 of the AI response output device 10010 and used by the control unit 1110. Furthermore, without being limited to this, the database of initial setting directives and / or conversation histories may be stored in a server on the network. For example, if the AI response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 in generating an AI response, the database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in these servers themselves. In this way, it is possible to omit the process of transmitting the initial setting directives and conversation histories again from the AI response output device 10010 to these servers by including them in directives, thereby reducing the number of transmission tokens for use of the large-scale language model.
[0223] When storing the database of the initial setting instruction sentence and / or the conversation history described in FIG. 4A, the AI response output device 10010 can transmit the user ID, the character ID, and the user instruction sentence for the subsequent conversation to these servers. The large-scale language models on these servers can use the user ID and character ID obtained from the AI response output device 10010 to obtain the corresponding initial setting instruction sentence and the conversation history from the database of the initial setting instruction sentence and / or the conversation history of FIG. 4A. The large-scale language models on these servers can perform inference using the initial setting instruction sentence and the conversation history, and the user instruction sentence for the subsequent conversation transmitted from the AI response output device 10010, to generate an AI response and transmit it to the AI response output device 10010. In this way, the effect of more appropriately maintaining the consistency of the character's personality and the continuity of memory for each different user can be obtained while saving the number of transmitted tokens when using the large-scale language model.
[0224] Next, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) according to the fourth embodiment of the present invention will be described with reference to Fig. 4B. The database according to the fourth embodiment shown in Fig. 4B is an extension of the database described in Fig. 1C or Fig. 2L. Specifically, the database shown in Fig. 4B assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data on standard response phrases corresponding to each user and character in the database.
[0225] In the example of Fig. 4B, for user 1, whose user ID is 1, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3. In addition, for user 2, whose user ID is 2, and user 3, whose user ID is 3, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3.
[0226] These standard response phrase data are stored as separate data in different areas for each user-character combination. For ease of explanation, in FIG. 4B, the data stored in each area is represented as standard response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, standard response phrase data 101 is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 1: Koto shown in FIG. 2L. Data 201 in FIG. 4B is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 2: Tom shown in FIG. 2L.
[0227] Data 301 in FIG. 4B is stored as a database such as a table corresponding to the fixed response phrases corresponding to condition numbers 1 to 7 of character 3: Necco shown in FIG. 2L. Data 102, 202, and 302 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 2. Data 103, 203, and 303 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 3. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses fixed response phrase data stored in different areas for each combination of user and character, based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.
[0228] In this way, even the same character can respond with different template responses for each user. That is, even for the same character, it may be more appropriate to vary the content of the template responses depending on the relationship between the character and the user. For example, depending on the relationship between the character's set age and the age of the user registered in the AI response output device 10010 or the system, the user may be older, the same age, or younger than the character. In this case, different content for the template responses of the character to older users, the template responses of the character to users of the same age, and the template responses of the character to younger users can make the conversation between the user and the character more appropriate or natural. That is, by performing operations using the database of FIG. 4B and varying the content of the template responses depending on the relationship between the character and the user, it is possible to create a more appropriate or natural conversation.
[0229] 4B described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the response template database (response template DB) shown in FIG. 4B may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0230] According to the character conversation device and character conversation system of Example 4 described above, it is possible to produce more suitable or more natural conversation depending on the relationship between the character and the user, the conversation history, etc.
[0231] <Example 5> Next, a fifth embodiment of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of the first, second, and third embodiments. Specifically, this is an example in which the response generation process of the AI response output device 10010 is switched from response generation process using a large-scale language model on a network to response generation process using a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, or response generation process using a fixed response phrase database. In this embodiment, differences from these embodiments will be described, and repeated description of configurations similar to those of these embodiments will be omitted.
[0232] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0233] An example of response generation process switching processing in the AI response output device 10010 of the fifth embodiment of the present invention will be described using FIG. 5A. The table in FIG. 5A shows examples 1 to 9 of response generation process switching processing in the AI response output device 10010. In the table in FIG. 5A, the column "Switching Overview" shows an overview of the switching processing for each example. The column "State before switching of LLM on the network (API connected LLM)" shows the state before response generation processing by a large-scale language model on the network (large-scale language model connected using an API), such as the large-scale language model provided in the large-scale language model server 19001 in FIG. 1 and the multimodal large-scale language model provided in the large-scale language model server 20001, is switched to another response generation process. The column "Switching Occurrence Condition" shows the conditions under which switching processing of the response generation process occurs. The column "Switching destination from LLM on the network (API-connected LLM)" indicates the switching destination to which the response generation process of the AI response output device 10010 is switched from a large-scale language model on the network (a large-scale language model connected using an API), such as a large-scale language model provided in the large-scale language model server 19001 and a multimodal large-scale language model provided in the large-scale language model server 20001. When the condition indicated in "Switching occurrence condition" occurs in the state of "State before switching of LLM on the network (API-connected LLM)" shown in Figure 5A, the control unit 1110 of the AI response output device 10010 may perform control to switch to the large-scale language model, database, or correspondence indicated in "Switching destination from LLM on the network (API-connected LLM)."
[0234] Below, each example shown in the table of FIG. 5A will be described. Example 1, as shown in the "Switching Overview," is an example in which switching is performed depending on the network connection status of the AI response output device 10010. In Example 1, the "state before switching of the LLM (API-connected LLM) on the network" indicates that the network connection status of the AI response output device 10010 is connectable. Here, in Example 1, the "switching occurrence condition" indicates "when network connection becomes unavailable." That is, this means when the connection via the network between the AI response output device 10010 and a large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable. Specifically, the connection failure may be due to a communication failure state on the connection path from the AI response output device 10010 to the Internet 19000. Alternatively, the connection failure may be due to a communication failure state on the Internet 19000. Alternatively, the connection failure may be due to a situation in which the large-scale language model on the network (a large-scale language model connected using an API) itself cannot connect to the Internet 19000. Also, in Example 1, "local LLM" is shown as the "switching destination from LLM on the network (API-connected LLM)." This specifically means that switching processing is performed to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. That is, in Example 1, even if connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and response generation processing by the large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable, response generation processing is switched to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. This makes it possible to continue response generation processing using the large-scale language model, despite differences in performance as a large-scale language model.
[0235] Next, Example 2 in FIG. 5A will be described. In Example 2, the "switching destination from the networked LLM (API-connected LLM)" in Example 1 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 2, if a connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and the response generation process using the large-scale language model on the network (a large-scale language model connected using an API) cannot be used, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response through simpler processing and output the response to the user.
[0236] Next, Example 3 of FIG. 5A will be described. In Example 3, the "switching destination from LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "no-response handling." The "no-response handling" means that even if a user input requesting a response from a large-scale language model is received from the user via the touch panel, microphone 1139, or operation input unit 1107, no response to this input is generated, or even if a user input requesting a response from a large-scale language model is received, no response to this is output. In other words, Example 3 makes it easier to handle a situation where connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason, making it impossible to use the response generation process by the large-scale language model on the network (a large-scale language model connected using an API).
[0237] Next, Example 4 of FIG. 5A will be described. As shown in the "Switching Overview," Example 4 is an example in which switching is performed due to a response delay of an LLM on the network. In Example 4, the "state before switching an LLM on the network (API-connected LLM)" indicates a state in which a response from an LLM on the network is obtained within a predetermined time. Here, in Example 4, the "switching occurrence condition" indicates a case in which a response from an LLM on the network is not obtained within the predetermined time and exceeds the predetermined time. Also, in Example 4, the "switching destination from an LLM on the network (API-connected LLM)" indicates a "local LLM." The "local LLM" of the switching destination is the same as in Example 1, so repeated explanation will be omitted. That is, in Example 4, even if the response from an LLM on the network (a large-scale language model connected using an API) exceeds the predetermined time for some reason and the response generation process by the LLM on the network (a large-scale language model connected using an API) cannot be used smoothly, the response generation process is switched to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This makes it possible to continue the response generation process using a large-scale language model, even if there are differences in performance as a large-scale language model.
[0238] Next, Example 5 in FIG. 5A will be described. Example 2 is an example in which the "switching destination from the networked LLM (API-connected LLM)" in Example 4 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 5, if for some reason a response from the networked LLM (large-scale language model connected using an API) exceeds a predetermined time and the response generation process using the networked LLM (large-scale language model connected using an API) cannot be used smoothly, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response using simpler processing and output the response to the user.
[0239] Next, Examples 6 to 9 of FIG. 5A will be described. As shown in "Switching Overview," Examples 6 to 9 are examples in which switching is performed when the upper limit of API usage or usage fee is reached. As described in Example 2, providers of large-scale language models often collect the costs used to train the large-scale language model from terminal users as API usage fees for the terminal. In such cases, with natural language models, API usage fees are often charged based on the number of processing of units of words that separate sentences, called tokens. Here, various methods of charging and limiting API usage fees are conceivable. One possible example is to define the upper limit of the amount of large-scale language model usage services that a user can receive under normal circumstances using the number of tokens processed.
[0240] In this case, the user may be able to receive services using large-scale language models at a specified API usage fee until the usage volume (or corresponding usage fee) is reached, and once the upper limit of the usage volume (or corresponding usage fee) is reached, certain restrictions may be imposed, such as the user being unable to receive services using large-scale language models in their normal state (performance or frequency).
[0241] Examples 6 to 9 in FIG. 5A are examples of response generation process switching control by the control unit 1110 of the AI response output device 10010 when such restrictions occur in a service using a large-scale language model. Specifically, in Example 6, the "state before switching the on-network LLM (API-connected LLM)" is a state in which the API usage volume and API usage fee have not reached a predetermined upper limit. This means that the usage volume of the on-network LLM (API-connected LLM) has not reached a predetermined upper limit. In this case, the user can use the on-network LLM (API-connected LLM) in a normal state.
[0242] Here, in Example 6, the "switching trigger condition" is when the API usage volume or API usage fee reaches a predetermined upper limit. This means when the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit. Also, in Example 6, the "switching destination from the network-based LLM (API-connected LLM)" is a second LLM on the network that is different from the LLM (which may be referred to as the first LLM) used in normal operation. An example of a second LLM on the network is an LLM with a lower fee than the first LLM used in normal operation. Since it is a lower-fee service, the performance of the second LLM may be lower than that of the first LLM. Even in this case, there is still a significant advantage if large-scale language models can be used inexpensively even after the usage volume / fee limit of the first LLM is reached.
[0243] Next, Example 7 in Figure 5A will be described. In Example 7, the "switching destination from the network-based LLM (API-connected LLM)" in Example 6 is changed from a second LLM on a network different from the LLM (which may be referred to as the first LLM) used in the normal state to a "local LLM." In Example 7, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit, it is possible to continue performing response generation processing using a large-scale language model by switching to response generation processing using a local LLM that is not subject to restrictions such as the usage volume of the network-based LLM, the API usage volume, or the API usage fee.
[0244] Next, Example 8 in FIG. 5A will be described. In Example 8, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is similar to the process described in FIG. 1C, FIG. 2L, or FIG. 4B, and therefore will not be described again. In Example 8, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the networked LLM (API-connected LLM) reaches a predetermined upper limit, the process switches to response generation using a preparatory response phrase database, which is not subject to restrictions such as the usage volume of the networked LLM, the API usage volume, or the API usage fee. This makes it possible to generate a response using simpler processing and output the response to the user.
[0245] Next, Example 9 in FIG. 5A will be described. In Example 9, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "no-response handling." The "no-response handling" refers to a handling in which a response to the user is not generated or a response to the user is not output. Example 9 makes it easier to handle a situation in which response generation processing using a large-scale language model on the network (a large-scale language model connected using an API) is unavailable because the API usage or API usage fee has reached a predetermined upper limit, i.e., the usage of the networked LLM (API-connected LLM) has reached a predetermined upper limit.
[0246] According to the switching control of the response generation process of the AI response output device 10010 shown in Examples 1 to 9 of Figure 5A as described above, even in situations where the response generation process using an LLM (large-scale language model connected using an API) on the network cannot be used as usual, more appropriate switching or response can be made according to each situation.
[0247] 5A may be performed by combining a plurality of examples. For example, the switching control of Examples 1 to 3 may be combined with any of the controls of Examples 4 to 9. Similarly, the control of Example 4 or Example 5 may be combined with any of the controls of Examples 1 to 3 or Examples 6 to 9. Similarly, the control of Examples 6 to 9 may be combined with any of the controls of Examples 1 to 5.
[0248] Next, an example of display of an AI assistant or a character when the AI response output device 10010 of the fifth embodiment is configured as an AI assistant device or a character conversation device will be described with reference to FIGS. 5B to 5D.
[0249] First, Figure 5B is an example of the display of an AI assistant or character on the AI response output device 10010 when performing the switching control of Example 3 in Figure 5A. In the example of Figure 5B, the display state of the AI assistant or character is changed depending on whether the network connection status of the AI response output device 10010 is network connectable or network unconnectable. The network connectable and network unconnectable states of the AI response output device 10010 are as explained in Figure 5A, so a repeated explanation will not be given.
[0250] In the example of FIG. 5B, the AI response output device 10010 (1) displays the AI assistant or character in a normal, awake state when a network connection is possible, but (2) displays the AI assistant or character in a "sleeping" state when a network connection is not possible. In the switching control of example 3 of FIG. 5A, when the AI response output device 10010 cannot connect to the network, it does not generate or output a response even if a command is input from the user. In this case, if the AI assistant or character displayed by the AI response output device 10010 is in a normal, awake state, the user will feel uncomfortable. However, if the AI assistant or character displayed by the AI response output device 10010 is displayed in a sleeping state, the user will understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user.
[0251] In the case of Figure 5B(2), it is desirable that the user understand that "the reason the AI assistant or character is not responding is because it is asleep" before making a user input requesting a response using a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the start timing of the state in Figure 5B(2) where the AI assistant or character is displayed in a "sleeping" state when network connection is not possible is immediately after the control unit 1110 of the AI response output device 10010 determines that network connection is not possible, before the user makes a user input requesting a response using a large-scale language model.
[0252] Next, as another display example, the display example of FIG. 5C will be described. The display example of FIG. 5C is an example in which the display state of the AI assistant or character is changed depending on the state of the "switching destination from the LLM on the network (API-connected LLM)" in the table during the switching control of FIG. 5A. Specifically, FIG. 5C shows (1) a display example of the AI assistant or character in a state in which the AI response output device 10010 can connect to a large-scale language model on the network (a large-scale language model connected using an API) and is able to use a response generation process using the large-scale language model on the network (referred to as the normal state in this figure); (2) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to a response generation process using an LLM or a template response database with lower performance than the large-scale language model on the network (a large-scale language model connected using an API); and (3) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A.
[0253] In the example of FIG. 5C , for example, (1) when the AI response output device 10010 is in a "normal state," the AI response output device 10010 displays the AI assistant or character in a state where there are no particular problems. Note that the "normal state" in FIG. 5C may be considered a state other than states (2) and (3). Also, for example, (2) when the AI response output device 10010 has switched to a response generation process using an LLM or a response template database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), the AI response output device 10010 displays the AI assistant or character in a "sleepy" state. Note that "displaying the AI assistant or character in a "sleepy" state" may also be expressed as "a display indicating that the AI assistant or character is feeling drowsy."
[0254] The response generation process (2) has lower performance than the response generation process using a large-scale language model on the network (a large-scale language model connected using an API) in the normal state (1). Therefore, by displaying the AI assistant or character in a "sleepy" state, it is possible to implicitly convey to the user that the response performance of the AI assistant or character is low. This makes it possible to further reduce the sense of discomfort felt by the user due to a low-performance response. Note that the switching conditions under which the AI response output device 10010 switches to response generation processing using an LLM or a fixed response phrase database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), are as described in FIG. 5A, and therefore a repeated explanation will be omitted.
[0255] 5C(2), it is desirable to implicitly inform the user that the response performance of the AI assistant or character is low before the user makes a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the start timing of the state in FIG. 5C(2) where the AI assistant or character is displayed in a "sleepy" state be immediately after the AI response output device 10010 switches to a response generation process using an LLM or a template response database, which has lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user makes a user input requesting a response from a large-scale language model.
[0256] Also, for example, in (3) the state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A, the AI response output device 10010 displays the AI assistant or character in a "sleeping" state. As also described in FIG. 5B, by displaying the AI assistant or character displayed by the AI response output device 10010 in a "sleeping" state, the user can understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user. Note that the conditions under which the AI response output device 10010 switches to the no-response mode described in FIG. 5A are the same as those described in Example 3 or Example 9 of FIG. 5A, and therefore a repeated explanation will be omitted. Note that in the case of FIG. 5C(3), it is desirable for the user to understand that "the AI assistant or character is not responding because it is sleeping" before making a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the timing for starting the state (3) in Figure 5C in which the AI assistant or character is displayed in a "sleeping" state be immediately after the AI response output device 10010 switches to the no-response response described in Figure 5A, before the user input requesting a response from a large-scale language model.
[0257] 5C, the AI response output device 10010 displays a display that implicitly reflects a change in the state of the AI assistant or character without directly providing the user with a technical explanation of the state of the AI response output device 10010 related to the response generation process. This can reduce the sense of discomfort felt by the user compared to when a technical explanation of the state of the AI response output device 10010 related to the response generation process is directly provided to the user. Furthermore, this can reduce the sense of discomfort felt by the user compared to when the display state of the AI assistant or character remains the same as its normal state despite a change in the state of the AI response output device 10010 related to the response generation process.
[0258] However, some users may wish to know a more precise explanation of the technical state of each state. Therefore, a display example for such users will be described with reference to FIG. 5D. Among the rows of the table shown in FIG. 5D, the rows for explaining the device state and display state are identical to those in FIG. 5C, and therefore, repeated explanations will be omitted. Furthermore, the display example of the AI assistant or character shown in the row for the display example of the AI assistant or character is almost identical to that in FIG. 5C, except that a question mark (?) is displayed in the display example. The question mark (?) is a mark that the user operates when requesting an explanation from the AI response output device 10010, and may also be referred to as a help mark.
[0259] In the example of FIG. 5D , when a user selects the question mark (?) through a user operation, such as via the touch panel of the operation input unit 1107 or the display unit 10011 in FIG. 1B , the display of the AI assistant or character on the AI response output device 10010 changes to the display example shown in the row of the display example after user operation. Specifically, regardless of whether the device state is (1), (2), or (3), an explanation of the technical state of each state is displayed. For example, in the example of FIG. 5D , if the device state is (1) normal, a display explaining that the device is in a normal state with no particular technical limitations can be displayed, such as "Normal state." Furthermore, if the device state is (2) using a low-performance LLM or a standard response phrase database, a display technically explaining the low-performance state can be displayed, such as "Low-performance mode." This display can also be considered a display explaining the reason why the AI assistant or character is displaying a "sleepy" state.
[0260] In this case, a more detailed technical explanation may be provided. Specifically, a message such as "Low-performance LLM usage mode" or "Canned response mode" may be displayed. If the device status is (3) unresponsive, a message such as "Network connection unavailable" may be displayed, providing a technical explanation of the reason for switching to unresponsive mode. If the reason for switching to unresponsive mode is that a response from an LLM (a large-scale language model connected via an API) on the network exceeds a specified time, a message such as "Response from LLM is delayed" may be displayed. If the reason for switching to unresponsive mode is that the network LLM usage, API usage, or API usage fee has reached its limit, a message such as "LLM usage limit reached," "API usage limit reached," or "API usage fee has reached a specified amount" may be displayed. These messages may be considered to explain the reason why the AI assistant or character is displayed in a "sleeping" state.
[0261] According to the display example of FIG. 5D described above, even if there are technical constraints in the response generation process in the AI response output device 10010, first, instead of providing a direct explanation to the user, the state of the device is implicitly indicated by a change in the display state of the AI assistant or character, thereby further reducing the sense of discomfort felt by the user. This display is more suitable for users who do not need technical explanations. Furthermore, by displaying an operation mark to explain the technical state, a display is provided to users who operate the mark that technically explains the state of the response generation process in the AI response output device 10010 (normal state or state with technical constraints). This makes it possible to provide a more suitable display for users who want to know the technical state accurately.
[0262] In the examples of Figures 5B, 5C, and 5D, a "sleeping" state is shown as an example of the display state of the AI assistant or character when the AI assistant or character is "unresponsive," but this is only an example and the embodiment is not limited to this. Instead of the "sleeping" state, another display state that implies a situation where the AI assistant or character is unable to respond, such as "taking a break," may be used. In the examples of Figures 5C and 5D, a "sleepy" state is shown as an example of the display state of the AI assistant or character when a low-performance LLM or a standard response phrase database is being used, but this is only an example and the embodiment is not limited to this. Alternatively, another display state that implies that the AI assistant or character has low response performance, such as "hungry," may be used.
[0263] According to the AI response output device and the AI response output system according to the fifth embodiment described above, it is possible to more appropriately switch the response generation process used by the AI response output device depending on the connection state between the large-scale language model on the network and the AI response output device, the response delay state from the large-scale language model on the network, the usage amount of the large-scale language model on the network, etc. Furthermore, when the AI response output device according to the fifth embodiment is configured as an AI assistant device or a character conversation device, it is possible to perform a display that is less strange to the user.
[0264] Example 6 Next, Example 6 of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of Examples 1 to 5. Specifically, this is an example in which the response generation process of the AI response output device 10010 is more suitably combined with a response generation process using a large-scale language model on a network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, and a response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated explanations of configurations similar to those examples will be omitted.
[0265] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0266] An example of a response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 shows an example of a flowchart of the response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention. Specifically, a time axis in which time progresses from top to bottom, a processing flow, and a response output example are shown. The response shown in the response output example may be output via display on the display unit 10011 of the AI response output device 10010 or audio output by the audio output unit 1140.
[0267] In the example of FIG. 6, first, at time t0, a user input is made via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010, requesting a response based on a large-scale language model, and the control unit 1110 of the AI response output device 10010 acquires the user input (step 600). Next, at time t1, the control unit 1110 starts preparations for response output using a fixed response phrase database stored in the storage unit 1170, and starts response output using the fixed response phrase database (step 601). In the example of FIG. 6, response output using the fixed response phrase database starts at time t2, and as shown in the figure, the fixed response is being output but has not yet been completed. "Good morning" in the figure indicates the output of part of the sentence that continues "Good morning..."
[0268] At time t3, before the response output using the fixed response phrase database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and transmits the generated instruction sentence to a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, thereby starting a request for a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 starts acquiring a response from the large-scale language model (step 603).
[0269] An example of a response output in which the response output using the fixed response phrase database is completed at time t5 is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is the ____ day of the month, isn't it?" is completed at time t5 using the fixed phrases stored in the fixed response phrase database and date information stored in memory. Here, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 has already started acquiring a response from the large-scale language model. Therefore, at time t6, which follows time t5 when the response output display using the fixed response phrase database is completed, the control unit 1110 starts outputting a response from the large-scale language model following the response output using the fixed response phrase database (step 604). Thereafter, at time t7, a response from the large-scale language model is output following the response output using the fixed response phrase database. When the response output from the large-scale language model is completed, the response output according to the processing flow shown in FIG. 6 is completed (step 605).
[0270] Next, the effect of the processing flow shown in Fig. 6 of the present invention will be described. Processing a large-scale language model requires a large amount of computational resources. Generally, even if inference, which requires fewer computational resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several to several tens of seconds from when the control unit starts to request a response from the large-scale language model until it can obtain a response from the large-scale language model. This period corresponds to the period from time t3 to time t4 shown in Fig. 6. Furthermore, from time t0, when a user input is made, until time t4, the control unit 1110 is unable to obtain a response output from the large-scale language model, and therefore is unable to output a response from the large-scale language model to the user.
[0271] 6, there is no start of preparation for response output using the fixed response phrase database and no start of response output using the fixed response phrase database, the user may continue to wait for a period of several seconds to more than ten seconds from time t0 when the user input is made to time t4 without receiving a response from the AI response output device 10010. For example, when the AI response output device 10010 is configured as an AI assistant device or a character conversation device, the waiting time may give the user a sense of discomfort.
[0272] In contrast, in the processing flow according to the sixth embodiment of the present invention shown in Fig. 6, the control unit 1110 starts the process of outputting a response using a fixed response phrase database, which requires fewer computational resources than the process of a large-scale language model, before starting to acquire a response from the large-scale language model. This prevents the user from having to wait from time t0 to time t4 without receiving a response from the AI response output device 10010. From the user's perspective, whether the response output is using a fixed response phrase database or a response output from a large-scale language model, it is the same as receiving a response from the AI response output device 10010.
[0273] 6, by providing step 601 before step 603, the response of the AI response output device 10010 to the user can be artificially accelerated. This can further reduce the sense of discomfort felt by the user due to long waiting times. Furthermore, by outputting a response from a large-scale language model following a response using the template response database in step 604, the user can perceive these outputs as if they were a series of more natural outputs.
[0274] According to the AI response output device and AI response output system of Example 6 described above, the waiting time for a response from the AI response output device can be shortened, and the sense of discomfort felt by the user can be further reduced.
[0275] Example 7 The seventh embodiment of the present invention is an improvement of the AI response output system described in the drawings of the first to sixth embodiments, in particular the AI response output device 10010. Note that in the seventh embodiment, differences from the first to sixth embodiments will be explained, and repeated explanations of the same configurations as those embodiments will be omitted.
[0276] As in the above-mentioned embodiments, the AI response output system of Example 7 is configured to include an AI response output device 10010, a large-scale language model server 19001 connected to the AI response output device 10010 via the Internet 19000, a multimodal large-scale language model server 20001, and a second server 19002 (see Figure 1A).
[0277] As in the above-described embodiments, the AI response output device 10010 may be referred to as a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0278] However, when an instruction sentence based on user input is sent from the AI response output device 10010 to a large-scale language model (LLM) and a response to the instruction sentence is generated by the large-scale language model, special information such as confidential information held by the AI response output device 10010 may be required to generate this response.
[0279] Furthermore, the large-scale language models that generate responses to instructional statements may include a first large-scale language model (hereinafter also referred to as the first LLM) that uses external input information for additional learning, and a second large-scale language model (hereinafter also referred to as the second LLM) that does not use input information for additional learning. The first large-scale language model is more versatile than the second large-scale language model. Therefore, the response content to an instructional statement generated by the first large-scale language model is more likely to be appropriate for the instructional statement than the response content generated by the second large-scale language model.
[0280] On the other hand, from the viewpoint of protecting input information, the second large-scale language model, which does not use the input information for additional learning, is superior to the first large-scale language model. If the input information includes, for example, personal information or confidential information, there is a risk that the personal information or confidential information may be exposed to a third party in the first large-scale language model because the input information is used for additional learning or because security is not implemented for the input information (for example, data is not encrypted or concealed). In contrast, the second large-scale language model does not use the input information for additional learning or because security is implemented for the input information, so there is a low possibility that the personal information or confidential information may be exposed to a third party.
[0281] Therefore, in the response output system of Example 7, a necessity determination is made to determine whether or not special information including confidential information is required as input information for generating a response using large-scale language models (first large-scale language model and second large-scale language model), and depending on the result of this necessity determination, one of the first large-scale language model or the second large-scale language model generates a response to the instruction sentence. In other words, in the response output system of Example 7, the large-scale language model used for generating a response is switched based on the result of the necessity determination. That is, in the response output device and response output system of Example 7, a destination to which an instruction sentence is sent is selected from multiple large-scale language models based on the result of the necessity determination.
[0282] More specifically, if the necessity determination determines that acquisition of special information is necessary for generating a response using the large-scale language model, a response to the instruction sentence is generated using the special information by the second large-scale language model. On the other hand, if the necessity determination determines that acquisition of special information is not necessary for generating a response using the large-scale language model, a response to the instruction sentence is generated mainly by the first large-scale language model without using special information. However, if the necessity determination determines that acquisition of special information is not necessary for generating a response using the large-scale language model, the response to the instruction sentence does not necessarily have to be generated by the first large-scale language model, but may be generated by the second large-scale language model.
[0283] In this way, by determining the large-scale language model that generates a response to an instruction sentence based on the result of the necessity determination, it is possible to appropriately protect the information contained in the instruction sentence or the response to the instruction sentence according to the type of information, while allowing the large-scale language model to generate a highly accurate response to the instruction sentence that the user intends.
[0284] 7 is a diagram illustrating an example of an AI response output device according to a seventh embodiment. As illustrated in FIG. 7, the AI response output device 10010 according to the seventh embodiment includes, as in the above-described embodiments, a display unit 10011, a control unit 1110, a memory 1109, a nonvolatile memory 1108, a local LLM processing unit 10028, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, and the like. Furthermore, the AI response output device 10010 according to the seventh embodiment includes a positioning sensor 1114, a timer 1115, a special information processing unit 1116, and an LLM connection switching processing unit 1117.
[0285] The positioning sensor 1114 acquires position information of the AI response output device 10010. The positioning sensor 1114 acquires, for example, the current position of the AI response output device 10010 as the position information of the AI response output device 10010. As an example, the positioning sensor 1114 acquires the current position of the AI response output device 10010 using a GPS (Global Positioning System) signal.
[0286] The timer 1115 acquires time information, such as the current time, from the AI response output device 10010. The method for acquiring the time information is not particularly limited. For example, the timer 1115 may have a clock function. Alternatively, the timer 1115 may acquire the time information from an external device connected via the Internet 19000 using location information acquired by the positioning sensor 1114.
[0287] The special information processing unit 1116 determines whether or not it is necessary to acquire special information as the above-mentioned necessity determination. Furthermore, the LLM connection switching processing unit 1117 controls the connection and disconnection between the control unit 1110 and the first LLM or the second LLM based on the determination result of the necessity determination by the special information processing unit 1116. In other words, the LLM connection switching processing unit 1117 performs switching processing to switch the destination of the instruction statement sent by the control unit 1110 based on the determination result of the special information processing unit 1116.
[0288] In this example, when the special information processing unit 1116 determines that acquisition of special information is unnecessary, the control unit 1110 of the AI response output device 10010 and the first LLM are communicatively connected, and the control unit 1110 and the second LLM are disconnected. Also, when the special information processing unit 1116 determines that acquisition of special information is necessary, the control unit 1110 and the first LLM are disconnected, and the control unit 1110 and the second LLM are connected. Also, when the control unit 1110 and the second LLM are connected in this manner, the special information processing unit 1116 transmits special information to the second LLM, as described below.
[0289] Also in this example, the large-scale language model server 190001 is provided with the first LLM, and the local LLM processing unit 10028 of the artificial intelligence response output device 10010 is provided with the second LLM. However, the large-scale language model server 19001 does not necessarily have to have the first LLM, and the artificial intelligence response output device 10010 does not necessarily have to have the second LLM. For example, the first LLM may be provided by the multimodal large-scale language model server 20001 or the artificial intelligence response output device 10010, and the second LLM may be provided by the large-scale language model server 19001 or the multimodal large-scale language model server 20001.
[0290] 7 has been described as an example in which the special information processing unit 1116 and the LLM connection switching processing unit 1117 are provided separately from the local LLM processing unit 10028, but the special information processing unit 1116 and the LLM connection switching processing unit 1117 may form part of the local LLM processing unit 10028. Alternatively, the special information processing unit 1116 and the LLM connection switching processing unit 1117 may form part of the control unit 1110.
[0291] Next, an example of the flow of generating a response to a directive sentence using a large-scale language model in the AI response output system of the seventh embodiment will be described.
[0292] FIG. 8 is a flowchart showing an example of the flow of a response generation process in the AI response output system of Example 7. As shown in FIG. 8, first, in step S01, when a user input (question / request, etc.) is received by the AI response output device 10010, the control unit 1110 of the AI response output device 10010 generates instruction statement 1 for the large-scale language model based on the user input. Next, in step S02, a necessity determination is made to determine whether or not special information needs to be acquired when generating a response to instruction statement 1 using the large-scale language model. As an example, the special information processing unit 1116 makes a necessity determination to determine whether or not special information needs to be acquired based on the content of instruction statement 1 or the user input, and the control unit 1110 acquires the determination result by the special information processing unit 1116. Note that in this example, the special information processing unit 1116 makes the necessity determination, but the control unit 1110 may also make the necessity determination.
[0293] In this necessity determination, it is determined whether or not it is necessary to acquire special information, or in other words, it is determined whether or not special information is necessary for generating a response to instruction sentence 1 using a large-scale language model. In yet another way, in the necessity determination, it is determined whether or not filtering of the special information held by the AI response output device 10010 is necessary (whether or not a special information filter is necessary).
[0294] As an example, if it is determined in this necessity determination that a special information filter is unnecessary (no filter), the first flag of the special information filter is set to "0" as shown in FIG. 9A. If the first flag is set to "0" (step S02: No), the control unit 1110 is connected to either the large-scale language model server 190001 having the first LLM or the local LLM processing unit 10028 having the second LLM through switching processing by the LLM connection switching processing unit 1117. In other words, the control unit 1110 is ready to send instruction statement 1 to either the large-scale language model server 19001 or the local LLM processing unit 10028. In this example, if the first flag is set to "0", the control unit 1110 is connected to the large-scale language model server 19001. Thereafter, the process proceeds to step S03, where the control unit 1110 sends the generated instruction statement 1 to, for example, the large-scale language model server 19001.
[0295] In addition, if the necessity determination determines that acquisition of special information is unnecessary (a special information filter is unnecessary), it is preferable that the first LLM primarily generates the response. However, depending on the content of instruction statement 1 or the initial setting of the LLM that transmits instruction statement 1, the second LLM may also generate the response. That is, in step S03, the control unit 1110 may transmit the generated instruction statement 1 to the local LLM processing unit 10028 that has the second LLM.
[0296] Next, the first LLM of the large-scale language model server 19001 performs a response generation process based on the instruction statement 1 transmitted from the control unit 1110 of the artificial intelligence response output device 10010 (step S04). As the response generation process, the first LLM generates, for example, a response sentence to a question or request in the instruction statement 1. Thereafter, the response (answer sentence) generated by the first LLM is transmitted to the control unit 1110, and the control unit 1110 performs a response reflection process in step S05. Specifically, the control unit 1110 displays the response content, such as the answer sentence, generated by the first LLM on the display unit 10011 serving as a user interface. The display unit 10011 is an example of an output unit that outputs the response content generated by the large-scale language model to the user. Note that the output unit is not limited to the display unit 10011 and may output the response content by voice, for example.
[0297] On the other hand, if it is determined in the necessity determination of step S02 that a special information filter is necessary (filter present), the first flag of the special information filter is set to "1." If the first flag is set to "1" (step S02: Yes), the process proceeds to step S06, where the control unit 1110 transmits instruction statement 1 to the second LLM (in this example, the local LLM processing unit 10028) provided in the response output device 10010. Furthermore, if it is determined in the necessity determination that a special information filter is necessary, the type of special information (which can also be called the type of special information filter) is further identified.
[0298] Here, the special information refers to various types of information held by the AI response output device 10010, and is information that serves as a basis for determining whether to use the first LLM or the second LLM to generate a response to a command. The special information can also be said to be information that can be acquired independently by the AI response output device 10010 and cannot be acquired from an external database via an external network such as the Internet. Examples of the special information held by the AI response output device 10010 include personal information of the user input to the AI response output device 10010, real-time information measured by the positioning sensor 1114, the timer 1115, etc., and information that cannot be acquired from outside. The special information acquired by the AI response output device 10010 is stored in the storage unit 1170 or the memory 1109 as appropriate.
[0299] In this example, such special information is classified into a plurality of types. As an example, as shown in Fig. 9B, the special information is classified into 12 types: "location information," "time information," "weather information," "schedule information," "email information," "telephone information," "speed information," "direction information," "image information," "device built-in information," "response failure information," and "other." These 12 types of special information are assigned second flags 1 to 12.
[0300] If the special information processing unit 1116 determines that a special information filter is necessary in the necessity determination, it identifies the type of special information (which can also be called the type of special information filter) that it has determined is necessary from the above 12 types and sets the corresponding second flag.
[0301] The classification of special information shown in FIG. 9B is merely an example. The "telephone information" with a second flag of 6 includes not only incoming call information and outgoing call information, but also voice information. The "device-internal information" with a second flag of 10 may be information specific to the AI response output device 10010, or may be information stored in the device with a second flag of 1 to 9 or 11. The "response deficiency information" with a second flag of 11, as described below, includes, for example, information such as "Cannot access ~" or "Cannot identify ~," or information indicating that a response from the second LLM cannot be received. The "other" with a second flag of 12 includes, for example, internal information of a device connected to the response output device 10010 via a wired connection, or internal information of a device that has been granted access permission after a decision to grant access. While the special information is classified into the 12 types described above in this example, the types and number of categories of special information are not particularly limited.
[0302] Then, when the special information processing unit 1116 determines in step S02 that acquisition of special information is necessary (step S02: Yes), it determines which of these 12 types of special information the special information that needs to be acquired corresponds to. For example, if the special information processing unit 1116 determines that the special information that needs to be acquired is "time information," it sets the second flag of the special information filter to "2." Also, for example, if the special information processing unit 1116 determines that the special information that needs to be acquired is "speed information," it sets the second flag of the special information filter to "7." Thereafter, the result of the necessity determination, including the first flag and the second flag, is sent from the special information processing unit 1116 to the control unit 1110, and in step S06, the control unit 1110 sends instruction statement 1 together with the above determination result to the local LLM processing unit 10028 of the response output device 10010.
[0303] Next, the local LLM processing unit 10028, which is the second LLM, acquires the special information held by the AI response output device 10010 (step S07). As an example, when the local LLM processing unit 10028 receives instruction statement 1, it requests the special information processing unit 1116 to acquire predetermined special information. In response to this request, the special information processing unit 1116 transmits the special information identified by the second flag to the local LLM processing unit 10028. As a result, the local LLM processing unit 10028 acquires the predetermined special information.
[0304] The special information may be transmitted by the control unit 1110 to the local LLM processing unit 10028 together with the instruction statement 1. Of course, the local LLM processing unit 10028 may access the storage 1170 or the memory 1109 to acquire the special information stored in the storage 1170 or the memory 1109.
[0305] After acquiring the special information, the local LLM processing unit 10028 then performs a response generation process for instruction statement 1 in step S08. For example, the local LLM processing unit 10028 generates an answer sentence as a response to the question or request of instruction statement 1. The response (answer sentence) generated by the local LLM processing unit 10028 is then transmitted to the control unit 1110, and the control unit 1110 performs a response reflection process (step S05). Specifically, the control unit 1110 causes the display unit 10011 to display the response content, such as the answer sentence, generated by the local LLM processing unit 10028.
[0306] Here, an example of response generation in the AI response output system of Example 7 will be further described with reference to Figures 10A and 10B. Figure 10A is a diagram showing an example of a main message (natural language text and non-natural language information source) of an instruction sent from the AI response output device 10010 to the first LLM, and an example of a main response (natural language text and non-natural language information source) generated by the first LLM. Figure 10B is a diagram showing an example of a main message of an instruction sent from the AI response output device to the second LLM, and an example of a main response generated by the second LLM.
[0307] Each example shown in FIG. 10A is an example in which it is determined in step S02 that acquisition of special information is not necessary (step S02: No), and a response to an instruction statement (instruction statement 1) is generated in the first LLM. The instruction statement in example 1 is "Please tell me the nearest station to Tokyo Tower," and "Tokyo Tower" is clearly indicated as information (location information) necessary for generating a response. Therefore, in step S02, it is determined that acquisition of special information is not necessary (step S02: No). Therefore, for the instruction statement in example 1, the first LLM generates a response statement, for example, "The nearest station to Tokyo Tower is Akabanebashi Station on the Oedo Subway Line," based on information acquired via an external network without acquiring special information (step S04).
[0308] Example 2 in Figure 10A is an example in which location information and time information, which are information necessary for generating a response, are clearly indicated in the instruction, and the first LLM generates a response without acquiring specific information. Example 3 is an example in which date and time information (time information) and target information, which are information necessary for generating a response, are clearly indicated in the instruction, and the first LLM generates a response without acquiring specific information. Example 4 is an example in which location information, which is information necessary for generating a response, is clearly indicated in the instruction, and the first LLM generates a response without acquiring specific information. Example 5 is an example in which a target image, which is information necessary for generating a response, is clearly indicated in the instruction, and the first LLM generates a response without acquiring specific information. Example 6 is an example in which a target video, which is information necessary for generating a response, is clearly indicated in the instruction, and the first LLM generates a response without acquiring specific information.
[0309] In this way, if it is determined in the necessity determination that acquisition of special information is unnecessary, a response to the instruction statement can be generated without acquiring the special information. In addition, in this case, by generating the response mainly using the first LLM, it becomes easier to obtain a highly accurate response that the user intends in response to the instruction statement based on the user input.
[0310] Each example shown in FIG. 10B is an example in which it is determined in step S02 that acquisition of special information is necessary, and a response to the instruction sentence is generated by the local LLM processing unit 10028 having the second LLM. The instruction sentence in Example 1 is "Please tell me the nearest station from here," and the information (current location information) necessary to generate the response is not clearly indicated. Therefore, in step S02, it is determined that acquisition of special information (location information) is necessary for the large-scale language model to generate a response to the instruction sentence (step S02: Yes). Therefore, the local LLM processing unit 10028 acquires the special information (current location information) (step S07) and generates a response sentence such as "The nearest station from here is Tokyo Station" based on the acquired special information (step S08).
[0311] Example 2 in FIG. 10B is an example in which the instruction does not clearly indicate location information and date and time information (time information), which are information necessary for generating a response, and the local LLM processing unit 10028 acquires specific information (location information, time information, weather information, etc.) and generates a response based on the acquired specific information. Example 3 is an example in which the instruction does not clearly indicate time information, which is information necessary for generating a response, and the local LLM processing unit 10028 generates a response based on specific information (time information, schedule information, etc.). Example 4 is an example in which the instruction does not clearly indicate time information, which is information necessary for generating a response, and the local LLM processing unit 10028 generates a response based on specific information (time information, email information, etc.). Example 5 is an example in which the instruction does not clearly indicate time information necessary for generating a response, and the local LLM processing unit 10028 generates a response based on specific information (time information, telephone information, etc.).
[0312] Example 6 is an example in which the instruction does not clearly indicate the time information required for response generation, and the local LLM processing unit 10028 generates a response based on specific information (time information, speed information, etc.). Example 7 is an example in which the instruction does not clearly indicate the location information required for response generation, and the local LLM processing unit 10028 generates a response based on specific information (location information, direction information, etc.). Example 8 is an example in which the instruction does not clearly indicate the time information or image required for response generation, and the local LLM processing unit 10028 generates a response based on specific information (time information, image information, etc.). Example 9 is an example in which the instruction does not clearly indicate the time information or video required for response generation, and the local LLM processing unit 10028 generates a response based on specific information (time information, image information, etc.).
[0313] In this way, if it is determined in the necessity determination that acquisition of special information is necessary, the local LLM processing unit 10028 having the second LLM acquires the special information as appropriate and generates a response to the instruction statement based on the acquired special information. As a result, when generating a response to the instruction statement using the special information, it becomes easier to obtain a highly accurate response as intended by the user while protecting the acquired special information, etc.
[0314] Incidentally, when a response is generated by the first LLM or the second LLM in response to an instruction statement generated from a user input as described above, and the content of the user input (instruction statement) and the content of the response thereto are displayed on the display unit 10011 as a user interface, it is preferable to make the display format of the response generated by the first LLM different from that of the response generated by the second LLM. For example, as shown in Fig. 11, it is preferable to make the display format of the respondent's name, character, response format (line style, color, etc.) different between the response generated by the first LLM and the response generated by the second LLM.
[0315] In the example shown in Figure 11, the response generated by the first LLM shows a bear character along with the respondent's name "First LLM," and the response content including the respondent's name and character is surrounded by a roughly rectangular dotted line. On the other hand, the response generated by the second LLM shows a dog character along with the respondent's name "Second LLM," and the response content including the respondent's name and character is surrounded by a so-called cloud-shaped two-dot chain line.
[0316] By using different display formats for the responses generated by the first LLM and the second LLM in this way, the user can easily recognize whether the first LLM or the second LLM is responding. Furthermore, even if a user input (question / request) to the AI response output device 10010 and a response by the first LLM or the second LLM occur multiple times in succession, that is, even if the conversation becomes long, the user can easily distinguish between the response by the first LLM and the response by the second LLM.
[0317] Fig. 12 is a flowchart showing another example of the flow of response generation in the AI response output system of Example 7. Another example of the flow of response generation in the AI response output system of Example 7 will be described with reference to Fig. 12. In the flowchart of Fig. 12, the same steps as in the flowchart of Fig. 8 are given the same reference numerals, and duplicated explanations will be omitted.
[0318] 12 is an example in which the special information processing unit 1116 of the artificial intelligence response output device 10010 generates supplemental information related to the generation of a response using a large-scale language model as needed, and the control unit 1110 displays the supplemental information generated by the special information processing unit 1116 on the display unit 10011 along with the conversation between the user and the first LLM and second LLM. Note that, as an example, the special information processing unit 1116 generates the supplemental information, but the control unit 1110 may also generate the supplemental information, for example.
[0319] Specifically, if it is determined in step S02 that acquisition of special information is unnecessary (step S02: No), the process proceeds to step S011, where, for example, the special information processing unit 1116 generates supplemental information 1 related to the result of the necessity determination and the connection destination. The control unit 1110 acquires the supplemental information 1 generated by the special information processing unit 1116 and performs a reflection process. As shown in an example in FIG. 13A, when a message such as "Connecting to the first LLM" is generated as supplemental information 1, the control unit 1110 displays the generated message (supplemental information 1) on the display unit 10011 following a user input (question / request) as a reflection process. At this time, it is preferable to display the supplemental information 1 separately from the user input and the response content of the first LLM.
[0320] Thereafter, as described above, instruction statement 1 is transmitted from control unit 1110 to the first LLM (step S03), and the first LLM performs a response generation process for instruction statement 1 (step S04). Here, additional input may be provided by the user between the time when generated statement 1 is generated in step S01 and the time when instruction statement 1 is transmitted to the first LLM in step S03. In this case, control unit 1110 transmits the additional information input by the user (additional input information) together with instruction statement 1 to the first LLM (step S03), and the first LLM performs a response generation process for instruction statement 1 taking into account the additional input information (step S04).
[0321] The response (answer sentence) generated by the first LLM is transmitted to the control unit 1110, and in step S05, the control unit 1110 performs a response reflection process. Next, the process proceeds to step S012, where, for example, the special information processing unit 1116 generates supplemental information 2 related to the response reflection process in step S05. The control unit 1110 acquires the generated supplemental information 2 and performs the reflection process. For example, as shown in FIG. 13A, when a message stating "The first LLM has responded" is generated as supplemental information 2, the control unit 1110 causes the display unit 10011 to display the generated message (supplemental information 2).
[0322] On the other hand, if it is determined in the necessity determination in step S02 that acquisition of special information is necessary (step S02: Yes), the process proceeds to step S013, where the special information processing unit 1116 generates supplemental information 3 regarding the determination result of the necessity determination. The control unit 1110 performs a process of reflecting the generated supplemental information 3. The content of the supplemental information 3 includes the determination result of the necessity of acquiring special information, a message to confirm in advance with the user that a response will be generated using the special information, etc.
[0323] 13B, when the messages "Connecting to the second LLM" and "To respond, you must obtain ** information," are generated as supplemental information 3, the control unit 1110 causes the display unit 10011 to display the generated message (supplemental information 3). Furthermore, in this example, the special information processing unit 1116 generates a question, "Do you want to obtain ** information and generate a response?" as supplemental information 3. As a reflection process, the control unit 1110 causes the display unit 10011 to display the generated question (supplemental information 3) along with options of "Yes" and "No."
[0324] Incidentally, as described above, in the necessity determination of step S02, if the special information processing unit 1116 determines that a special information filter is necessary (filter present) (step S02: Yes), it sets the first flag of the special information filter to "1." Furthermore, if the special information processing unit 1116 determines that a special information filter is necessary in the necessity determination (step S02: Yes), it further identifies the type of special information (which can also be called the type of special information filter). In other words, the special information processing unit 1116 sets the second flag of the special information filter according to the type of special information.
[0325] In this example, the type of special information is identified based on the answer to the above question. Specifically, if the user selects "Yes" in response to the above question, the special information processing unit 1116 determines that the user has permission to acquire special information, and identifies the type of special information (type of special information filter) that is determined to be necessary to acquire based on the content of instruction statement 1 as described above, and sets the corresponding second flag. On the other hand, if the user selects "No" in response to the above question, the special information processing unit 1116 determines that the user has not permission to acquire special information, and identifies the type of special information (type of special information filter) as "response deficiency information" and sets the second flag to "11" regardless of the content of instruction statement 1.
[0326] Then, in step S06, the control unit 1110 transmits the result of the necessity determination including the first flag and the second flag together with the instruction statement 1 to the local LLM processing unit 10028 of the response output device 10010. Next, the local LLM processing unit 10028 acquires special information held by the AI response output device 10010 as needed (step S07).
[0327] Then, when the local LLM processing unit 10028 acquires the special information, it then performs a response generation process in step S08. At this time, if the second flag of the special information is set to anything other than "11," the local LLM processing unit 10028 generates a reply sentence as a response to the question or request of instruction statement 1. On the other hand, if the second flag of the special information is set to "11," the local LLM processing unit 10028 generates a message such as "Cannot acquire ** information, so cannot respond to instruction statement 1" as a response generation process.
[0328] Thereafter, a response (answer sentence) or message generated by the local LLM processing unit 10028 is transmitted to the control unit 1110, and the control unit 1110 performs a response reflection process (step S05). Specifically, the control unit 1110 causes the display unit 10011 to display the response content, such as the answer sentence, generated by the local LLM processing unit 10028, or the message described above. Also in this example, the process then proceeds to step S012, where, for example, the special information processing unit 1116 generates supplemental information 2 related to the response reflection process performed in step S05. The control unit 1110 acquires the generated supplemental information 2 and performs a reflection process. For example, as shown in FIG. 13B, when a message such as "A response was generated using ** information" is generated as supplemental information 2, the control unit 1110 causes the display unit 10011 to display the generated message (supplemental information 2). Thereafter, the process proceeds to the flow of FIG. 17, which will be described later, as necessary.
[0329] In this way, the user can easily understand the processing status for the instruction statement by generating supplementary information as appropriate and displaying it on the display unit 10011 in addition to the response to the instruction statement. Also, by displaying the response content for the instruction statement and the supplementary information separately, the user can more easily recognize the response content from the LLM.
[0330] 13A and 13B, the supplemental information is displayed separately from the response of the second LLM even after the control unit 1110 is connected to the second LLM (local LLM processing unit 10028), but the display format of the supplemental information is not limited to this. For example, as shown in FIG. 14, the supplemental information after connection to the second LLM may be displayed as a response generated by the second LLM. In this case, the local LLM processing unit 10028 having the second LLM may generate the supplemental information, rather than the control unit 1110.
[0331] Furthermore, in the example shown in FIG. 13B, the type of special information that needs to be acquired is displayed as supplemental information, but the content that is displayed as supplemental information is not particularly limited.
[0332] The supplemental information may, for example, display information indicating which part of the question (instruction) entered by the user causes the need to acquire special information. Specifically, as shown in an example in FIG. 15, a message such as "** information needs to be acquired in response to the instruction '****'" may be displayed as supplemental information. This allows the user to more clearly understand why the acquisition of special information is necessary, making it easier for them to take action such as correcting, adding, or deleting the instruction.
[0333] Furthermore, for example, if it is determined in the necessity determination that there are multiple types of special information that need to be acquired, a message may be displayed as supplemental information to ask the user which special information to acquire. For example, as shown in FIG. 16, along with the message "Please select the information to acquire," a selection option such as "XX information," "△△ information," or "All necessary information" may be displayed, and the user may select one of the options. This allows the user to generate a response to the instruction while appropriately adjusting the special information to be used in the response to the instruction and the content to be prioritized in the response.
[0334] 16, the user has selected individual pieces of information such as "XX information," and a response to the instruction is generated based on the selected individual pieces of specific information. In this case, a message such as "Further response requires XX information" may be displayed for the remaining information that is missing from the response to the instruction.
[0335] Furthermore, when a response to a command is generated using special information, details of the special information used and its storage location may be displayed as supplemental information. For example, if a personal image is used as the special information, details of the image "** information" and information about the storage location "Storage A" may be displayed, as shown in FIG. 17. Alternatively, an application or the like for the storage location may be launched to display information about the image, etc.
[0336] This allows the user to immediately check the content and storage location of the special information used for the response. Furthermore, the user can confirm whether or not the special information appropriate for the response has been acquired. Furthermore, by launching an application or the like in the storage location and displaying information such as images, the user can easily understand how to access the storage location. Another effect is that the user can easily search for information similar to or related to the acquired special information.
[0337] Furthermore, when a response to a command sentence is generated using special information, a message may be displayed as supplemental information to prompt the user to confirm the storage location of conversation information including this response. The storage location of the conversation information may also be displayed as supplemental information. In other words, the user may be allowed to specify the storage location of conversation information including the response generated using special information. For example, as shown in an example in FIG. 18, a message saying "Please select the storage location of this conversation information" may be displayed along with options such as "Storage A," "Storage B," and "Storage C," allowing the user to select one of the options.
[0338] Furthermore, when a response to a command sentence is generated using special information, when a predetermined time has elapsed since the response was generated or when a predetermined time has arrived, the content of the conversation information displayed on display unit 10011 may be automatically deleted, and a message to that effect may be displayed as supplemental information on display unit 10011. For example, as shown in Fig. 18, a message such as "This conversation information will be deleted in 24 hours" may be displayed as supplemental information.
[0339] This allows the user to store highly confidential conversation information, including responses generated using special information, separately from other conversation information, making it easier to manage the conversation information. Note that deleting conversation information at a predetermined time can further reduce the possibility of information leakage.
[0340] In the seventh embodiment, when it is necessary to acquire special information to generate a response to a command statement as described above, the response is generated by the second LLM. Therefore, the security of the acquired special information is ensured, but further measures may be taken to further enhance the security of the special information.
[0341] For example, when a response is generated using special information in the second LLM, as shown in an example in FIG. 19 , the display state of the display unit 10011 may be temporarily hidden so that the response generated by the second LLM cannot be read by the user, and the user may be prompted to perform an authentication procedure to cancel this hidden state. In this example, the display unit 10011 displays messages such as "Please complete the authentication procedure to display the response" and "Do you want to start the authentication procedure?" along with options of "Yes" and "No." If the user selects "Yes," the authentication procedure is executed. The authentication procedure may be performed by any method, including, but not limited to, facial recognition, voice recognition, or input of a password for acquiring special information. After the authentication procedure is completed, the display unit 10011 displays a message such as "The authentication procedure has been completed" as supplemental information, and then displays the response content from the second LLM, which was hidden. This enhances security for responses generated using special information and prevents leakage of special information contained in the response.
[0342] Furthermore, for example, when a response is generated using special information in the second LLM, as shown in an example in FIG. 20, the user may be required to perform an authentication procedure when acquiring the special information. In other words, acquisition of the special information may be permitted on the condition that the user has completed the authentication procedure. In the example of FIG. 13B described above, the user is prompted in advance to confirm that a response will be generated using special information. However, in the example of FIG. 20, following the advance confirmation, a message such as "Please complete the authentication procedure to acquire ** information" is displayed, prompting the user to perform the authentication procedure. This prevents the second LLM from generating a response before user authentication is complete. As a result, leakage of special information, etc., contained in the response content generated in the second LLM can be more reliably prevented.
[0343] In addition, in the example shown in Figure 12, the special information processing unit 1116 of the artificial intelligence response output device 10010 generated supplementary information as needed, but for example, as shown in an example in Figure 21, the large-scale language model server 19001 having a first LLM or the local LLM processing unit 10028 having a second LLM may generate supplementary information as needed.
[0344] 21, if it is determined in step S02 that acquisition of special information is unnecessary (step S02: No), the process proceeds to step S011, as in the above example, where, for example, the special information processing unit 1116 generates supplemental information 1 regarding the result of the necessity determination, and the control unit 1110 performs a process of reflecting the supplemental information 1 (see FIG. 13A). Next, the control unit 1110 transmits instruction statement 1 to the large-scale language model server 19001 having the first LLM (step S03), and the first LLM performs a process of generating a response to instruction statement 1 (step S04).
[0345] Thereafter, as in the above example, the response (answer sentence) generated by the first LLM is transmitted from the large-scale language model server 19001 to the control unit 1110, and in step S05, the control unit 1110 performs a response reflection process. Next, the process proceeds to step S012, where the control unit 1110 generates supplemental information 2 regarding the response reflection process in step S05, and performs a reflection process of the supplemental information 2.
[0346] On the other hand, if it is determined in step S02 that acquisition of special information is necessary (step S02: Yes), the process proceeds to step S06, where instruction statement 1 including the first flag and the second flag is transmitted from control unit 1110 to local LLM processing unit 10028 having a second LLM. Next, in this example, local LLM processing unit 10028 generates supplemental information 3 regarding the determination result of the necessity determination and performs processing to reflect the generated supplemental information 3. Examples of the content of supplemental information 3 include the determination result of the necessity of acquiring special information and a message to the user confirming in advance that a response will be generated using the special information (see FIG. 13B). In other words, the example of FIG. 21 can also be said to be an example in which local LLM processing unit 10028 is provided with special information processing unit 1116.
[0347] Thereafter, as in the above example, the local LLM processing unit 10028 acquires special information held by the AI response output device 10010 as needed (step S07). Once the special information is acquired, a response generation process is then performed in step S08. The response (answer sentence) generated by the response generation process is sent to the control unit 1110, and the control unit 1110 performs a response reflection process in step S05. Next, the process proceeds to step S012, where the control unit 1110 generates supplemental information 2 related to the response reflection process in step S05 and performs a reflection process of the supplemental information 2.
[0348] As described above, the generation and reflection processing of each piece of supplemental information does not necessarily have to be performed by the control unit 1110, but may instead be performed by, for example, the local LLM processing unit 10028 having the second LLM. That is, the local LLM processing unit 10028 may be provided with the special information processing unit 1116. Even with this configuration, when generating a response to a command statement using special information, it becomes easier to obtain a highly accurate response as intended by the user while protecting the special information, etc. Note that, although the example of FIG. 21 has been described as an example in which the second LLM generates and reflects the supplemental information, it is of course also possible for the first LLM to generate and reflect the supplemental information.
[0349] Fig. 22 is a flowchart showing another example of the flow of response generation in the artificial intelligence response output system of Example 7. In the above example, an example was described in which the control unit 1110 determines the need to acquire special information from the content of instruction statement 1, but in the example shown in Fig. 22, the need to acquire special information is determined from the response content generated by the first LLM. In Fig. 22, steps that are the same as those in the above flowchart are given the same reference numerals and descriptions thereof will be omitted.
[0350] In this example, when instruction statement 1 is generated in step S01, the generated instruction statement 1 is first transmitted to a large-scale language model server 19001 having a first LLM. More specifically, when instruction statement 1 is generated in step S01 (step S01), for example, supplemental information 1 regarding the connection destination is generated and the supplemental information 1 is reflected (step S011). As shown in an example in FIG. 23, the display unit 10011 of the AI response output device 10010 displays a message saying "Connecting to the first LLM" as supplemental information 1. Next, instruction statement 1 is transmitted to the large-scale language model server 19001 having the first LLM (step S03), and the first LLM generates a response (pre-response) to instruction statement 1 (step S04). The response (pre-response) generated by the first LLM is then transmitted from the large-scale language model server 19001 to the control unit 1110.
[0351] Next, the process proceeds to step S02, where it is determined whether or not it is necessary to acquire special information based on the content of the preliminary response generated by the first LLM. This determination of whether or not it is necessary may be made based on the content of instruction statement 1 as well as the content of the preliminary response generated by the first LLM. If it is determined in this determination of whether or not it is necessary to acquire special information (step S02: No), the process proceeds to step S05, where a response reflection process is performed. That is, the response (preliminary response) generated by the first LLM in step S04 is displayed on display unit 10011 as a response after the determination of whether or not it is necessary. Next, the process proceeds to step S012, where, for example, supplementary information 2 related to the response reflection process is generated, and a reflection process of supplementary information 2 is performed.
[0352] On the other hand, if it is determined in the necessity determination in step S02 that acquisition of special information is necessary (step S02: Yes), the process proceeds to step S031, where a reflection process is performed on the prior response generated by the first LLM.
[0353] Here, as described above, the first LLM cannot acquire the special information held by the AI response output device 10010. Therefore, if various information such as location information, date and time information, etc. is not clearly stated in instruction statement 1, the first LLM cannot generate an appropriate response to instruction statement 1 in step S04, and generates a response indicating that the necessary information cannot be acquired. Therefore, in the process of reflecting the prior response in step S031, information indicating that the special information cannot be acquired is displayed on the display unit 10011. Specifically, as shown in an example in FIG. 23, messages such as "Cannot access ~" and "Cannot identify ~" are displayed on the display unit 10011.
[0354] 24 is a diagram showing an example of a main message (natural language text and non-natural language information source) of an instruction sent from an AI response output device to a first LLM, and an example of a main response (natural language text and non-natural language information source) generated by the first LLM. Note that the response example in FIG. 24 corresponds to the above-mentioned preliminary response.
[0355] Example 1 in FIG. 24 is an example of a case where the content of location information is not clearly indicated in the instruction statement. Example 2 is an example of a case where the content of location / date and time information is not clearly indicated in the instruction statement. Examples 3 to 6 are examples of a case where the content of date and time information is not clearly indicated in the instruction statement. Example 7 is an example of a case where the content of location information is not clearly indicated in the instruction statement. Example 8 is an example of a case where the target image is not clearly indicated in the instruction statement. Example 9 is an example of a case where the target video is not clearly indicated in the instruction statement.
[0356] In the necessity determination in step S02, if the response (pre-response) generated by the first LLM indicates that the specified information cannot be acquired, as illustrated in FIG. 24, it is determined that acquisition of special information is required to generate a response to instruction statement 1. In other words, the first flag of the special information filter is set to "1." Furthermore, since the generated response indicates "Cannot access ~," the second flag is set to "11."
[0357] If the first flag is set to "1" in the necessity determination in step S02, that is, if it is determined that acquisition of special information is necessary (step S02: Yes), the process proceeds to step S031 as described above, and a reflection process is performed on the response (pre-response) generated by the first LLM in step S04. For example, as shown in an example in Fig. 23, a message such as "Cannot access *******.****" is displayed on the display unit 10011 as the response of the first LLM.
[0358] Next, the process proceeds to step S013, where supplemental information 3 regarding the result of the necessity determination is generated and the supplemental information 3 is reflected. If there is an error in the content of the preliminary response, this supplemental information 3 is information that clearly indicates the defective portion of the preliminary response or the user's instruction that caused the defective preliminary response by the first LLM. As shown in an example in FIG. 23, as information clearly indicating the defective portion of the preliminary response by the first LLM, messages such as "The response to the instruction '****' is unclear" and "'**** cannot be accessed' is not a clear response" are displayed. Furthermore, as messages prompting the acquisition of special information, messages such as "** information must be acquired" and "We recommend connecting to the second LLM and acquiring ** information" are displayed as necessary. Furthermore, in this example, a question such as "Do you want to connect to the second LLM?" is displayed to confirm with the user whether or not to connect to the second LLM.
[0359] Thereafter, the type of special information is identified based on the answers to the questions as described above, and in step S06, the control unit 1110 transmits the result of the necessity determination, including the first flag and the second flag, along with instruction statement 1, to the local LLM processing unit 10028 having the second LLM. Next, the local LLM processing unit 10028 acquires the special information held by the AI response output device 10010 as needed (step S07), and performs a response generation process using the special information (step S08). Once the response generated by the local LLM processing unit 10028 is transmitted to the control unit 1110, the process proceeds to step S05, where the control unit 1110 performs a response reflection process. Next, the process proceeds to step S012, where supplemental information 2 is generated and reflected.
[0360] In this example, it is possible to determine whether or not special information needs to be acquired in consideration of the results of the response generation process (pre-response generation process) of the first LLM. Therefore, it is possible to improve the accuracy of determining whether or not special information needs to be acquired. Furthermore, if the response (pre-response) generated by the first LLM is not the response intended by the user, it is possible to eliminate the need for the user to manually send an instruction to another LLM, thereby improving user usability. Furthermore, by appropriately displaying supplemental information 3, the user can grasp the defective and non-defective parts of the response (pre-response) by the first LLM, as well as the corresponding instruction parts, and this information can be used to obtain a highly accurate response as intended by the user.
[0361] In this example, the first LLM generates a response (pre-response) to the instruction, and then the second LLM generates a response to the instruction as needed. However, the order in which the responses are generated is not limited to this, and the second LLM may generate a response (pre-response) to the instruction without using special information prior to the first LLM.
[0362] Fig. 25 is, for example, a flowchart following the flowchart shown in Fig. 12, and is a flowchart when there is additional input from the user to the AI response output device 10010. With reference to Fig. 25, another example of the flow of the response generation process in the AI response output system of the seventh embodiment will be described.
[0363] After the generation and reflection of supplemental information related to the response reflection process is performed in step S012 of the flowchart shown in Fig. 12, if the user provides additional input to the AI response output device 10010, for example, as shown in Fig. 25, the control unit 1110 generates instruction statement 2 for the first LLM or the second LLM based on the additional input (step S041). Next, it is determined whether the control unit 1110 is currently connected to the second LLM (step S042). That is, it is determined whether the control unit 1110 is connected to the large-scale language model server 19001 having the first LLM or the local LLM processing unit 10028 having the second LLM.
[0364] If it is determined that the control unit 1110 is connected to the large-scale language model server 19001, that is, if it is determined that the control unit 1110 is not connected to the local LLM processing unit 10028 (step S042: No), the process proceeds to step S043, where a necessity determination is made to determine whether acquisition of special information is necessary to generate a response to instruction statement 2. If it is determined in this necessity determination that acquisition of special information is not necessary (step S043: No), the process proceeds to step S044, where supplemental information is generated and reflected, as in the example described above. Next, in step S045, instruction statement 2 is sent from the control unit 1110 to the first LLM.
[0365] Thereafter, similar to step S04 described above, the first LLM performs a response generation process for the instruction sentence (step S046), and the response (answer sentence) generated by the first LLM is transmitted to the control unit 1110. Also, similar to step S05 described above, the control unit 1110 performs a response reflection process (step S047). Furthermore, similar to step S012 described above, supplementary information related to the response reflection process of step S047 is generated and reflected (step S048).
[0366] On the other hand, if it is determined in step S042 that a connection to the second LLM is in progress (step S042: Yes), and if it is determined in the necessity determination in step S043 that acquisition of special information is necessary (step S043: Yes), the process proceeds to step S049, where instruction statement 2 is transmitted from control unit 1110 to local LLM processing unit 10028 having the second LLM. Note that the transmission of instruction statement 2 in steps S045 and S049 is not limited to the process of transmitting newly generated instruction statement 2, but also includes the process of adding additional input content by the user to the above-mentioned instruction statement 1 and transmitting it.
[0367] In other words, in steps S042 and S043, when multiple exchanges (conversations) are conducted between the user and the first and second LLMs, it is determined whether special information has been acquired when generating a response to a past instruction. If special information has been acquired in the past, the control unit 1110 is connected to the local LLM processing unit 10028 having the second LLM (step S042: Yes). Therefore, even if special information does not need to be acquired to generate a response to the current instruction, instruction 2 is sent to the local LLM processing unit 10028 (step S050).
[0368] That is, when determining whether or not it is necessary in step S043, if the conversation information including past instruction sentences and past responses generated in response to the instruction sentences contains special information, it is determined that obtaining special information is also necessary to generate a response to the current instruction sentence (step S043: Yes), and the process proceeds to step S050.
[0369] On the other hand, if special information has not been acquired in the past, the control unit 1110 is connected to the large-scale language model server 19001 (step S042: No), and if special information needs to be acquired to generate a response to the current instruction statement (step S043: Yes), instruction statement 2 is sent to the local LLM processing unit 10028, which has the second LLM (step S050). The local LLM processing unit 10028 acquires special information held by the AI response output device 10010 as needed (step S051), and performs response generation processing using the acquired special information (step S052).
[0370] In this example, when instruction statement 2 is sent to the local LLM processing unit 10028, the user is asked whether or not to acquire special information and generate a response. As shown in an example in FIG. 26, messages (supplemental information) such as "Connecting to the second LLM" and "You must acquire ** information" are displayed on the display unit 10011. Furthermore, the display unit 10011 displays the question (supplemental information) "Do you want to acquire ** information and generate a response?" along with options of "Yes" and "No." If the user selects "Yes," instruction statement 2 is sent to the local LLM processing unit 10028.
[0371] 25, when multiple exchanges (conversations) occur between a user and a large-scale language model, the large-scale language model to which an instruction sentence is sent is selected depending on whether acquisition of special information was necessary to generate a response to a past instruction sentence. That is, in this example, if it is determined that acquisition of special information is necessary when generating a response to a past instruction sentence in a series of conversations, and the control unit 1110 is connected to the local LLM processing unit 10028 having the second LLM, the second LLM will generate a response to the instruction sentence thereafter, regardless of the need to acquire special information. In other words, if special information is included in past conversation information including instruction sentences and responses to the instruction sentences, it is determined in the necessity determination that acquisition of special information is necessary, regardless of whether acquisition of special information is necessary to generate a response to the current instruction sentence.
[0372] FIG. 27 is a diagram illustrating an example of a conversation between a user and a large-scale language model. In Example 1 of FIG. 27, it was determined that acquisition of specific information was not necessary to generate a response to the instruction sentence in the first round, and the first LLM generated the response. However, it was determined that acquisition of specific information was necessary to generate a response to the instruction sentence in the second round in the same conversation, and the second LLM generated the response. Therefore, although acquisition of specific information was not necessary to generate a response to the instruction sentence in the third round, the second LLM generated the response. Also, in Example 2 of FIG. 27, it was determined that acquisition of specific information was necessary to generate a response to the instruction sentence in the first round in the series of conversation, and the second LLM generated the response. As a result, the second LLM generated responses to all instruction sentences.
[0373] In this way, when special information is acquired during a series of conversations, the second LLM generates all subsequent responses, making it possible to more reliably protect all special information contained in the series of conversations while generating the response intended by the user for each instruction.
[0374] After the local LLM processing unit 10028 generates a response in step S052, the process proceeds to step S047, where the control unit 1110 reflects the response. Furthermore, in step S048, supplemental information is generated and reflected.
[0375] For example, as shown in FIG. 26, when the response content from the second LLM is displayed on the display unit 10011, the message "A response was generated using ** information" is displayed as supplementary information, along with the reason why the response was generated in the second LLM. In this example, the message "This conversation cannot switch to the first LLM. To switch to the first LLM, start a new conversation" is displayed as the reason why the response was generated in the second LLM. Furthermore, if the response generated by the second LLM is incomplete, i.e., if an appropriate response to instruction statement 2 is not generated, information recommending that a response be generated in the first LLM is displayed. For example, the message "There is an inadequacy in the response. We recommend switching to the first LLM and starting a new conversation. Do you want to do this?" is displayed along with the options "Yes" and "No." If the user selects "Yes," a new conversation is started and a response is generated in the first LLM.
[0376] Although the response output system of Example 7 has been described above, the configuration of the response output system is not limited to the above. For example, in Example 7, the response output system includes a large-scale language model including a first LLM and a second LLM, and one of the first LLM and the second LLM generates a response to a command statement depending on whether acquisition of special information is required to generate a response to the command statement. However, the response output system does not necessarily include the first LLM and the second LLM. For example, the large-scale model may include a first mode that generates a response without using special information and a second mode that generates a response using special information, and switch between the first mode corresponding to the first LLM and the second mode corresponding to the second LLM depending on whether acquisition of special information is required to generate a response to the command statement.
[0377] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable AI response output technology. For example, such AI response output technology is expected to be introduced into higher quality, more reliable infrastructure. The introduction of this technology into infrastructure can contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This will contribute to the achievement of "Build resilient infrastructure, promote inclusive and sustainable industrialization, inclusive and sustainable technological development," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0378] Furthermore, the technology according to the present embodiment makes it possible to provide a more suitable AI response output technology. For example, such AI response output technology is expected to be introduced into public transportation facilities to improve access to transportation systems for vulnerable people. The introduction of this technology into public transportation can contribute to improving traffic safety through the expansion of public transportation and realizing access to a safe, affordable, and easily usable sustainable transportation system for all people. This contributes to "Sustainable cities and communities," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0379] Although various embodiments have been described above in detail, the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments are detailed descriptions of the entire system to clearly explain the present invention, and the present invention is not necessarily limited to a system including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0380] Furthermore, the numerical values, messages, etc. appearing in the text and figures are merely examples, and the effects of the present invention will not be impaired even if different ones are used.
[0381] Furthermore, some or all of the above-described configurations, functions, processing units, processing means, etc. may be implemented in hardware, for example, by designing them as integrated circuits. Furthermore, the above-described configurations, functions, etc. may be implemented in software by a processor interpreting and executing a program that implements each function. The processor may include transistors and other circuits and be considered circuitry or processing circuitry. Information such as programs, tables, and files that implement each function may be stored in memory, a recording device such as a hard disk or solid-state drive (SSD), or a recording medium such as an IC card, SD card, or DVD, or may be stored in a device on a communication network. Furthermore, the control lines and information lines shown are those considered necessary for explanation, and do not necessarily represent all control lines and information lines in the product. In reality, it may be assumed that almost all components are interconnected. [Explanation of symbols]
[0382] 10010... Artificial intelligence response output device, 10011... Display unit, 10028... Local LLM processing unit, 1107... Operation input unit, 1110... Control unit, 1132... Communication unit, 1140... Audio output unit, 1139... Microphone, 1160... Video control unit, 1170... Storage unit, 1180... Imaging unit< / audio> < / video> < / audio> < / video>
Claims
1. a control unit that transmits an instruction sentence based on a user input to a large-scale language model and causes the large-scale language model to generate a response to the instruction sentence; an output unit that outputs a response generated by the large-scale language model to a user, the large-scale language models include a first large-scale language model and a second large-scale language model; The control unit causing one of the first large-scale language model or the second large-scale language model to generate a response to the instruction sentence, and outputting the generated response via the output unit, according to a result of a necessity determination that determines whether acquisition of special information held by the response output device is necessary as input information from outside in order to generate a response by the large-scale language model; Response output device.
2. 2. The response output device according to claim 1, the large-scale language models include the first large-scale language model that uses external input information for additional learning, and the second large-scale language model that does not use the input information for additional learning. Response output device.
3. 3. The response output device according to claim 2, If it is determined in the necessity determination that acquisition of the special information is necessary for generating a response by the large-scale language model, The control unit causing the second large-scale language model to generate a response to the instruction using the special information; Response output device.
4. 3. The response output device according to claim 2, If it is determined in the necessity determination that acquisition of the specific information is not necessary for generating a response using the large-scale language model, The control unit causing the first large-scale language model to generate a response to the instruction; Response output device.
5. 3. The response output device according to claim 2, the output unit is a display unit that displays a response generated by the large-scale language model; The control unit displaying the response generated by the first large-scale language model and the response generated by the second large-scale language model in different formats on the display unit; Response output device.
6. 3. The response output device according to claim 2, the output unit is a display unit that displays a response generated by the large-scale language model; The control unit displaying supplemental information regarding the generation of a response by the large-scale language model together with the response on the display unit; Response output device.
7. 7. The response output device according to claim 6, the supplemental information includes content that confirms to the user in advance that a response will be generated by the large-scale language model using the specific information; Response output device.
8. 7. The response output device according to claim 6, the supplemental information includes a reason why the special information is needed to generate a response by the large-scale language model; Response output device.
9. 7. The response output device according to claim 6, If it is determined in the necessity determination that acquisition of multiple types of special information is necessary to generate the response, The control unit before transmitting the instruction sentence to the large-scale language model, displaying the plurality of types of special information determined to need to be acquired as the supplemental information on the display unit, and requesting a user to select from the plurality of types of special information an item to be permitted to be used in generating the response. Response output device.
10. 7. The response output device according to claim 6, The supplemental information includes information on a storage location where the special information is stored. Response output device.
11. 7. The response output device according to claim 6, The supplemental information includes information on a storage location where conversation information including responses generated by the large-scale language model is stored. Response output system.
12. 3. The response output device according to claim 2, the output unit is a display unit that displays a response generated by the large-scale language model; If a response is generated by the second large-scale language model using the special information, The control unit When a predetermined time has elapsed since the response was displayed on the display unit, the content of the response displayed on the display unit is deleted. Response output device.
13. 3. The response output device according to claim 2, the output unit is a display unit that displays a response generated by the large-scale language model; if the response is generated using the specific information by the second large-scale language model, The control unit the display of the display unit is in a non-display state in which the response cannot be read, and when the authentication procedure by the user is completed, the non-display state is canceled and the response is displayed. Response output device.
14. 3. The response output device according to claim 2, The control unit prior to the necessity determination, causing the first large-scale language model or the second large-scale language model to generate a preliminary response to the instruction sentence without using the special information; In the necessity determination, it is determined whether or not acquisition of the special information is necessary based on the content of the prior response. Response output device.
15. 15. The response output device according to claim 14, the output unit is a display unit that displays a response generated by the large-scale language model; The control unit If there is a defect in the content of the preliminary response, the defect in the preliminary response is displayed on the display unit together with the response as supplementary information regarding the generation of the response by the large-scale language model. Response output device.
16. 3. The response output device according to claim 2, In the determination of necessity, if the special information is included in conversation information including a past instruction sentence and a past response generated by the large-scale language model, it is determined that acquisition of the special information is also necessary for generating a response to the current instruction sentence. Response output device.
17. Large-scale language models and a response output device that causes the large-scale language model to generate a response to an instruction sentence based on a user input, and outputs the generated response to the user, the large-scale language models include a first large-scale language model and a second large-scale language model; one of the first large-scale language model and the second large-scale language model generates a response to the instruction sentence according to a result of a necessity determination that determines whether acquisition of special information held by the response output device is necessary as input information from outside in order to generate a response by the large-scale language model; Response output system.
Citation Information
Patent Citations
Structural unit for tank construction
JP1977008512A