Response output system and response output device

The response output device improves user interaction by integrating a control unit and storage for large-scale language models, addressing inadequate configurations in existing AI response technologies.

JP2026077224APending Publication Date: 2026-05-13MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MAXELL LTD
Filing Date
2024-10-25
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing response output technologies using artificial intelligence do not adequately consider user interactions, leading to unsuitable configurations.

Method used

A response output device comprising an input interface, control unit, storage unit, and output interface, capable of interacting with a large-scale language model to generate and output responses based on user input, while storing settings for character conversations.

Benefits of technology

Provides a more suitable response output technology by enhancing user interaction through a comprehensive system design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026077224000001_ABST
    Figure 2026077224000001_ABST
Patent Text Reader

Abstract

To provide a more suitable artificial intelligence response output technology. According to this invention, it will contribute to Sustainable Development Goals (SDGs) "9. Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation" and "11. Make cities and human settlements inclusive, safe, resilient and sustainable." [Solution] A response output system comprising: a large-scale language model; an input unit that receives user input; a control unit that generates instruction sentences for the large-scale language model based on the user input and acquires the response generated by the large-scale language model to the instruction sentences; and an output unit that outputs the response acquired by the control unit to the user. The control unit transmits information regarding emotion types that classify user emotions acquired from user input to the large-scale language model, the large-scale language model generates a response according to a response mode selected according to the emotion type, and the output unit outputs the response to the user according to output conditions corresponding to the response mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a response output system and a response output device.

Background Art

[0002] Regarding response output technologies using artificial intelligence such as language models, for example, they are disclosed in Patent Document 1.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the disclosure of Patent Document 1, considerations regarding configurations for more suitably providing response output technologies using artificial intelligence to users were not sufficient.

[0005] An object of the present invention is to provide a more suitable response output technology.

Means for Solving the Problems

[0006] To solve the above problems, for example, the configuration described in the claims may be adopted. The present application includes multiple means for solving the above problems, but to give one example, a response output device comprising an input interface for receiving user input, a control unit, a storage unit, and an output interface for outputting a response to the user, wherein the control unit is capable of executing a client application that can send and receive information with a server outside the response output device or a large-scale language model application that controls a large-scale language model stored inside the response output device, the client application is capable of generating instruction sentences for the large-scale language model based on user input received via the input interface, sending control information different from the instruction sentences to the large-scale language model application, sending the instruction sentences to the large-scale language model application, receiving response sentences from the large-scale language model application which are the results of inferences performed by the large-scale language model, and outputting a response based on the response sentences to the user via the output interface, and the storage unit stores settings related to the characteristics of character conversations. [Effects of the Invention]

[0007] According to the present invention, a more suitable response output technology can be provided. Other problems, configurations, and effects will be clarified in the following description of embodiments. [Brief explanation of the drawing]

[0008] [Figure 1A] This figure shows an example of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 1B] This figure shows an example of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 1C] This figure shows an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 2A] This is an explanatory diagram of an example of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2B]This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2C] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2D] This is an explanatory diagram illustrating an example of a conversation in a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2E] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2F] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2G] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2H] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2I] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2J] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2K] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 2L] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 3A] This is an explanatory diagram of an example of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 3B] This is an explanatory diagram illustrating an example of the operation of a character conversation device and character conversation system according to one embodiment of the present invention. [Figure 3C]It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3D] It is an explanatory diagram of an example of conversation in a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3E] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3F] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3G] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3H] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3I] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4A] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4B] It is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 5A] It is an explanatory diagram of an example of the operation of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5B] It is an explanatory diagram of an example of a display example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5C] It is an explanatory diagram of an example of a display example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5D] It is an explanatory diagram of an example of a display example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 6] It is an explanatory diagram of an example of a response generation process of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 7]This is an explanatory diagram of an example of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 8A] This is an explanatory diagram illustrating an example of the configuration of an artificial intelligence response output system according to one embodiment of the present invention. [Figure 8B] This is an explanatory diagram illustrating an example of the configuration of an artificial intelligence response output system according to one embodiment of the present invention. [Figure 8C] This is an explanatory diagram illustrating an example of the configuration of an artificial intelligence response output system according to one embodiment of the present invention. [Figure 9A] This is an explanatory diagram of an example of a conversation used to describe one embodiment of the present invention. [Figure 9B] This is an explanatory diagram of an example of a conversation used to describe one embodiment of the present invention. [Figure 10] This flowchart shows an example of processing in an artificial intelligence response output system according to one embodiment of the present invention. [Figure 11] This is an explanatory diagram of an example of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 12A] This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 12B] This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 13A] This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 13B] This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 14A] This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 14B] This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 15] This figure shows an example of information about the response mode according to one embodiment of the present invention. [Figure 16] This figure shows an example of information regarding output conditions according to one embodiment of the present invention. [Figure 17] This figure shows an example of information about the response mode according to one embodiment of the present invention. [Figure 18]This figure shows an example of emotion type information according to one embodiment of the present invention. [Figure 19] This is an explanatory diagram of an example of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 20] This flowchart shows an example of processing in an artificial intelligence response output system according to one embodiment of the present invention. [Figure 21] This is an explanatory diagram illustrating an example of a conversation in an artificial intelligence system according to one embodiment of the present invention. [Figure 22] This is an explanatory diagram illustrating an example of a conversation in an artificial intelligence system according to one embodiment of the present invention. [Figure 23] This is an explanatory diagram illustrating an example of a conversation in an artificial intelligence system according to one embodiment of the present invention. [Figure 24] This flowchart shows an example of processing in an artificial intelligence response output system according to one embodiment of the present invention. [Figure 25] This is an explanatory diagram illustrating an example of a conversation in an artificial intelligence system according to one embodiment of the present invention. [Figure 26] This flowchart shows an example of processing in an artificial intelligence response output system according to one embodiment of the present invention. [Figure 27] This flowchart shows an example of processing in an artificial intelligence response output system according to one embodiment of the present invention. [Figure 28] This figure shows an example of a threshold used in the processing of an artificial intelligence response output system according to one embodiment of the present invention. [Figure 29] This figure illustrates an example of a response mode transition according to one embodiment of the present invention. [Figure 30] This is an explanatory diagram of an example of an information table according to one embodiment of the present invention. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described in detail below with reference to the drawings. However, the present invention is not limited to the examples described herein, and various modifications and alterations are possible by those skilled in the art within the scope of the technical ideas disclosed herein. Furthermore, in all the figures used to illustrate the present invention, components having the same function are given the same reference numerals, and repeated descriptions may be omitted.

[0010] Furthermore, if the artificial intelligence response output device according to each embodiment of the present invention has a display screen, it may be called a display device. If the artificial intelligence response output device has a voice output function, it may be called a voice output device. The artificial intelligence response output device may simply be called an information processing device. A system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence response output system. Also, if the artificial intelligence response output device provides a response service of a large-scale language model, which is artificial intelligence, to the user and assists the user, the artificial intelligence response output device or the display output of the artificial intelligence response output device can become an artificial intelligence (AI) assistant for the user. Therefore, in this case, the artificial intelligence response output device may be called an AI assistant device or an AI assistant display device. Similarly, in this case, a system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an AI assistant system or an AI assistant display system. Also, in this case, since the artificial intelligence response output device becomes an interface between the user and artificial intelligence, it may be called an artificial intelligence interface device. In this case, a system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence interface system.

[0011] <Example 1> As Embodiment 1 of the present invention, an artificial intelligence response output device and system that outputs a response from a large-scale language model artificial intelligence will be described.

[0012] An example of the artificial intelligence response output device 10010 of the present invention will be described using Figure 1A. Furthermore, an example of a system in which the artificial intelligence response output device 10010 cooperates with a large-scale language model server 19001 through communication or other means will be described, including the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001.

[0013] In the example shown in Figure 1A, the artificial intelligence response output device 10010 has a display unit 10011. In the example shown in Figure 1A, the display unit 10011 may be a flat panel display, a screen that projects images from the back, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which pixels emit light themselves. Furthermore, the display unit 10011 may be equipped with a touch operation input sensor and configured as a touch panel.

[0014] In the example shown in Figure 1A, the voice output unit 1140 of the artificial intelligence response output device 10010 is composed of a speaker. The artificial intelligence response output device 10010 is also equipped with a microphone 1139 that can pick up the user's voice. Through voice input from the microphone 1139 and user operation input via the operation input unit described later, the artificial intelligence response output device 10010 can acquire user input that forms the basis of instructions (prompts) for the large-scale language model, which is an artificial intelligence.

[0015] The artificial intelligence response output device 10010 may be equipped with a local large-scale language model. In this case, the response of the large-scale language model may be output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.

[0016] Furthermore, the artificial intelligence response output device 10010 may not have a local large-scale language model, but instead communicate with an external large-scale language model server 19001 and output the response received from the large-scale language model server 19001 as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.

[0017] Alternatively, the artificial intelligence response output device 10010 may also include a local large-scale language model and be configured to communicate with an external large-scale language model server 19001 having a large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be switched and output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140. Alternatively, the response generated based on both the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.

[0018] The configuration when the artificial intelligence response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the internet 19000 via a communication unit 1132. In the example in Figure 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as a wireless example, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may have both wired and wireless sections, or it may pass through routers or repeaters. Similarly, the communication path from the communication unit 1132 to the internet 19000 may also have both wired and wireless sections, or it may pass through routers or repeaters. The artificial intelligence response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the internet 19000. Furthermore, the artificial intelligence response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 that is different from these servers, via the communication device 19011 and the internet 19000. The configuration including the artificial intelligence response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.

[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" should be understood as encompassing the local large-scale language model provided by the artificial intelligence response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.

[0020] In the example in Figure 1A, the display unit 10011 shows an example where each element is displayed in two display areas: an instruction display area 10051 where the user inputs an instruction (prompt) to a large-scale language model which is artificial intelligence, and an artificial intelligence response display area 10061 which displays the response from the large-scale language model. In the example in Figure 1A, the instruction display area 10051 shows an example where an icon 10052 representing the user, text such as natural language or software code 10053 as a component of the instruction, an image 10054 as a component of the instruction, a video 10055 as a component of the instruction, etc. In the example in Figure 1A, the artificial intelligence response display area 10061 shows an example where an icon 10062 representing artificial intelligence or an artificial intelligence assistant, text such as natural language or software code 10063 as a component of the response from artificial intelligence, an image 10064 as a component of the response from artificial intelligence, a video 10065 as a component of the response from artificial intelligence, etc. Note that the display example of the display unit 10011 of the artificial intelligence response output device 10010 shown in Figure 1A is merely an example. Depending on the implementation example in which the artificial intelligence response output device 10010 is used, a different display from the example shown in Figure 1A may be used.

[0021] Here, we will explain large-scale language models. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT have been made publicly available. These technologies can be used in this embodiment as well. These large-scale language models are artificial intelligence models that have been generated through extensive pre-training on natural language contained in a large number of documents and texts that exist in the human world. The number of parameters of these artificial intelligence models exceeds hundreds of millions. Furthermore, in addition to this, there are also models that incorporate reinforcement learning based on human feedback. An example of a base model is a model called Transformer. As an example of training these models, for example, Reference 1 is publicly available.

[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf

[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. More advanced models can even perform natural language question answering (also called dialogue or conversation), natural language suggestion generation, and programming code generation. Because these artificial intelligence models have a very large number of parameters, training requires enormous amounts of data and computing resources. Therefore, training this level of artificial intelligence for a specific application is extremely resource-inefficient. To address this, large-scale pre-training is performed to generate foundation models that can be applied to various uses. For example, the large-scale language model server 19001 shown in Figure 1A may be equipped with such a large-scale language model and configured to be usable by various terminals via an API (Application Programming Interface). Alternatively, the artificial intelligence response output device 10010 shown in Figure 1A may be equipped with a local large-scale language model and configured to use it itself. The training of any large-scale language model can be performed separately through large-scale pre-training to generate it, and the generated large-scale language model can then be duplicated and provided to the large-scale language model server 19001, artificial intelligence response output device 10010, etc. In this way, instead of performing pre-training for each application or terminal, duplicating the large-scale language model, which is the foundation model generated through large-scale pre-training, and using it on individual servers and terminals allows for the sharing of resource consumption used for training, resulting in better resource efficiency.

[0024] Furthermore, even if a large-scale language model is generated as a foundational model through extensive pre-training, it may be configured to perform additional training, such as transfer learning, on individual servers or devices, depending on the application and purpose.

[0025] Furthermore, large-scale language models can pre-train on natural language and perform input / output processing targeting natural language. In addition, multimodal large-scale language model artificial intelligence capable of processing not only natural language text information but also other types of information is also applicable to the embodiments of the present invention. Figure 1A shows a large-scale language model server 20001 having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model artificial intelligence include GPT-4 (see Reference 2) and Gato (see Reference 3), which are publicly available. These technologies may also be used in this embodiment. These multimodal large-scale language models are artificial intelligence models generated by performing large-scale pre-training on natural language and other types of information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. Furthermore, there are also models that incorporate reinforcement learning based on human feedback. Hereinafter, information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.

[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf

[0027] Next, using Figure 1B, we will describe an example configuration of an artificial intelligence response output device 10010 that receives user input to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as the large-scale language model to the user input.

[0028] The artificial intelligence response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, and the like. The artificial intelligence response output device 10010 may have a large screen, such as a so-called monitor or television.

[0029] The display unit 10011 may be a flat panel display, a screen that projects images from the back, or a display that projects an optical image into the air to show a floating image. If the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Alternatively, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which pixels emit light themselves. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be equipped with a touch operation input sensor and configured to accept touch operation input from the user 230's finger. In this case, the display unit 10011 may be configured as a touch panel. Through the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire user input that forms the basis of instructions (prompts) to the large-scale language model, which is artificial intelligence.

[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth® communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 1132 of the artificial intelligence response output device 10010 can communicate with the communication device 19011 connected to the internet 19000. The communication path between the communication unit 1132 and the communication device 19011 may include both wired and wireless sections, and may also pass through routers or repeaters. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN communication method. This allows the artificial intelligence response output device 10010 to communicate with various servers connected to the internet 19000.

[0031] The artificial intelligence response output device 10010 is equipped with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011 and the communication unit 1132, etc.

[0032] The power supply 1106 converts the AC current input from an external source via the external power input interface 1111 into DC current and supplies the necessary DC current to each part of the artificial intelligence response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. In addition, the secondary battery 1112 supplies power to each part that requires power via the external power input interface 1111 when power is not supplied from an external source.

[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit such as a remote controller, or an infrared light receiving unit, and inputs signals for operations other than touch operations by the user to the touch operation input sensor of the display unit 10011. Separately from the user who touches the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the artificial intelligence response output device 10010. Through the user's operation input via the operation input unit 1107, the artificial intelligence response output device 10010 can obtain user input that forms the basis of instructions (prompts) to the large-scale language model, which is artificial intelligence. Note that there may also be a modified configuration in which the touch operation input sensor of the display unit 10011 is included as part of the operation input unit 1107.

[0034] The video signal input unit 1131 receives video data by connecting an external video output device. Various digital video input interfaces are possible for the video signal input unit 1131. For example, it can be configured with an HDMI (High-Definition Multimedia Interface) standard video input interface, a DVI (Digital Visual Interface) standard video input interface, or a DisplayPort standard video input interface. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also use various USB interfaces.

[0035] The audio signal input unit 1133 receives audio data by connecting an external audio output device. The audio signal input unit 1133 may be configured as an HDMI standard audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface with integrated terminals and cables.

[0036] The audio output unit 1140 is capable of outputting audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 is also capable of outputting audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured as a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output audio signals as digital signals to external devices, such as the Audio Return Channel function specified in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output audio signals as analog signals to external devices such as headphones.

[0037] Microphone 1039 is a microphone that picks up sounds from the vicinity of the artificial intelligence response output device 10010, converts them into signals, and generates audio signals. The microphone may be configured to record human voices, such as the user's voice, and the control unit 1110, described later, may perform speech recognition processing on the generated audio signal to obtain textual information from the audio signal. Through the audio input from microphone 1139, the artificial intelligence response output device 10010 can obtain user input that forms the basis of instructions (prompts) for the large-scale language model, which is an artificial intelligence.

[0038] The imaging unit 1180 is a camera having an image sensor. The camera may be provided on the front of the display unit 10011 side of the artificial intelligence response output device 10010, or on the back of the display unit 10011 side. Both a front camera and a rear camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a rear camera.

[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be composed of a magnetic recording medium such as a hard disk drive (HDD) or a semiconductor memory such as a solid-state drive (SSD). For example, the storage unit 1170 may have various types of information, such as video data, image data, and audio data, pre-recorded in it at the time of product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from external devices or external servers via the communication unit 1132. The video data, image data, etc., recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, etc., recorded in the storage unit 1170 may also be output to external devices or external servers via the communication unit 1132.

[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may also be called a video processing circuit and may be composed of hardware such as an ASIC, FPGA, or video processor. The video control unit 1160 may also be called a video processing unit or image processing unit. The video control unit 1160 performs video switching control, such as determining which video signal to input to the display unit 10011 from among the video signals stored in the memory 1109 and the video signals (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform image processing control on the video signals input from the video signal input unit 1131 and the video signals stored in the memory 1109. Examples of image processing include scaling processing such as enlarging, reducing, and transforming images; brightness adjustment processing to change the brightness; contrast adjustment processing to change the contrast curve of an image; and retinex processing which decomposes an image into its light components and changes the weighting of each component.

[0041] The attitude sensor 1113 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the attitude of the artificial intelligence response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each connected part.

[0042] The non-volatile memory 1108 stores various data used by the artificial intelligence response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, data for various operations displayed on the display unit 10011 of the artificial intelligence response output device 10010, display icons, data and layout information for objects operated by the user. Memory 1109 stores video data and device control data displayed on the display unit 10011. The control unit 1110 may read various software from the storage unit 1170, expand it into memory 1109, and store it.

[0043] The local LLM processing unit 10028 has memory capable of holding a large-scale language model (LLM) and can perform inference on the LLM based on the control of the control unit 1110. The hardware can be a so-called GPU (Graphics Processing Unit). The local LLM processing unit 10028 may perform training as well as inference. Note that the local LLM processing unit 10028 is not necessarily required if the execution of LLM inference on the LLM in the local environment of the artificial intelligence response output device 10010 is not necessary.

[0044] The control unit 1110 controls the operation of each connected part. The control unit 1110 may also work in cooperation with a program stored in memory 1109 to perform calculation processing based on information acquired from each part within the artificial intelligence response output device 10010. One of the control states by the control unit 1110 is, for example, the output of responses from the large-scale language model of the local LLM processing unit 10028, or responses from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001, acquired via the communication unit 1132, through the display unit 10011 or the audio output unit 1140, which is a speaker or the like.

[0045] Furthermore, when input is received from the user via the touch panel, microphone 1139, or operation input unit 1107 as described above, the control unit 1110 can perform the control to generate an instruction sentence based on that input and send it to the local large-scale language model of the local LLM processing unit 10028 of the artificial intelligence response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and to obtain a response from these large-scale language models.

[0046] Furthermore, the storage unit 1170 may store a standard response database (which may also be referred to as a standard response DB) for outputting standard phrases as responses to instructions from the artificial intelligence response output device 10010. The control unit 1110 can then perform control to generate the response to be output using the data stored in the standard response database. Figure 1C shows an example of a standard response database. In the example in Figure 1C, the standard response to be output by the artificial intelligence response output device 10010 is stored for each condition assigned a condition number. For example, if the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, as in condition number 1, the response should be output using "Good morning" or "Today is the XXth of XX." as the standard response. The XX part, such as "XXth of XX," can be generated using information stored in the memory 1109 or other memory of the artificial intelligence response output device 10010.

[0047] Furthermore, in the example of a standard response phrase in the database shown in Figure 1C, if multiple standard response phrases separated by / are stored, the control unit 1110 can be controlled to randomly select one of the standard response phrases using a random number or the like and output a response. This can resolve and improve the situation where responses under the same conditions become monotonous. The explanation for the examples of condition numbers 2, 3, and 4 is the same as for the example of condition number 1. The control unit 1110 should be controlled to output using the standard response phrases shown in Figure 1C for each example of the condition content shown in Figure 1C.

[0048] Next, we will explain an example of condition number 5 shown in Figure 1C. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of user input obtained via the touch panel, microphone 1139, or operation input unit 1107 as natural language, or when there is an obvious grammatical error in the user input, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I may not know about that." By responding in this way, the system can prompt the user to input again and wait for corrected user input.

[0049] Next, we will explain an example of condition number 6 shown in Figure 1C. Condition number 6 is an example where the control unit 1110 detects an error (abnormal state) in any of the parts that make up the artificial intelligence response output device 10010 shown in Figure 1B, and user input is received via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 controls the system to output a response using the standard response phrase "It seems to be malfunctioning." By responding in this way, the system can explain to the user that the artificial intelligence response output device 10010 is malfunctioning and prompt the user to take action to address the error.

[0050] The artificial intelligence response output device 10010 may output a response using the response template database (response template DB) described with reference to Figure 1C, instead of a response from a large-scale language model such as the local large-scale language model provided by the artificial intelligence response output device 10010, the large-scale language model provided by the large-scale language model server 19001, or the multimodal large-scale language model provided by the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the response template database (response template DB).

[0051] The response template database (response template DB) shown in Figure 1C described above is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 can use it. However, the response template database (response template DB) shown in Figure 1C may also be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 should generate responses using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 should send the response generated using the response template database (response template DB) to the artificial intelligence response output device 10010 instead of the response generated by the large-scale language model stored in their respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a response template database (response template DB), it becomes possible to generate responses using a response template database (response template DB).

[0052] In the above description, the artificial intelligence response output device 10010 was described as having a display panel for a display screen using fixed pixels. This concept may also include a projection-type image display device (projector) in which a projection optical system is provided after the display panel for a display screen using fixed pixels, and the optical image of the image on the display panel for the display screen is projected onto a screen or wall.

[0053] In the examples shown in Figures 1A and 1B, an example was described in which the artificial intelligence response output device 10010 includes a display unit 10011. However, the artificial intelligence response output device 10010 according to the embodiment of the present invention does not necessarily have to include a display unit 10011. For example, even without a display unit 10011, the device can be configured to receive user input to the artificial intelligence via an audio signal input unit 1133 or a microphone 1139, and to output a response from the artificial intelligence, such as a large-scale language model, to the user input via an audio output unit 1140.

[0054] According to the artificial intelligence response output device and artificial intelligence response output system of Embodiment 1 of the present invention described above, it is possible to receive input from a user to artificial intelligence such as a large-scale language model and output a response to the user input generated by the inference of the artificial intelligence, such as a large-scale language model on a server device on the network or a local large-scale language model on the artificial intelligence response output device itself.

[0055] <Example 2> Next, as Embodiment 2 of the present invention, we will describe an example in which the artificial intelligence response output device 10010 described in Embodiment 1 is connected to the internet and operates by connecting to a server equipped with a large-scale language model artificial intelligence via the internet. In this embodiment, we will explain the differences from Embodiment 1, and repeating explanations of configurations similar to those in these embodiments will be omitted.

[0056] Using Figure 2A, an example of the connection state between the artificial intelligence response output device 10010 and the large-scale language model server 19001 of Embodiment 2 of the present invention will be explained. The artificial intelligence response output device 10010 according to Embodiment 2 may also be called a character conversation device. Furthermore, the system including the artificial intelligence response output device 10010 and the large-scale language model server 19001 according to Embodiment 2 may also be called a character conversation system. The display unit 10011 displayed by the artificial intelligence response output device 10010 shows an image of character 19051. The image of character 19051 is an image generated by rendering a 3D model of the character in a virtual space.

[0057] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can become an artificial intelligence (AI) assistant for the user. In this case, the character conversation device or character conversation system in this embodiment may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.

[0058] In the example shown in Figure 2A, the voice output unit 1140 of the artificial intelligence response output device 10010 is composed of a speaker. The artificial intelligence response output device 10010 is also equipped with a microphone 1139 that can pick up the user's voice. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the internet 19000 via a communication unit 1132. In the example shown in Figure 2A, the communication between the communication unit 1132 and the communication device 19011 is shown as wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the internet 19000 may have both wired and wireless sections. The artificial intelligence response output device 10010 can communicate with a large-scale language model server 19001 via the communication device 19011 and the internet 19000. Furthermore, the artificial intelligence response output device 10010 can communicate with a second server 19002, different from the large-scale language model server 19001, via the communication device 19011 and the internet 19000. The configuration including the artificial intelligence response output device 10010 and the large-scale language model server 19001 may be considered as a single system.

[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described using Figure 2B. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that in Figure 2B, the communication paths such as the internet 19000 shown in Figure 2A have been omitted. In Figure 2B, the user 230 of the artificial intelligence response output device 10010 is also shown.

[0060] Here, we will explain the sequence of operations of the artificial intelligence response output device 10010. The artificial intelligence response output device 10010 loads the character operation program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character operation program, thereby enabling the various processes described below.

[0061] First, the artificial intelligence response output device 10010 is equipped with a microphone 1139. When user 230 speaks to character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. The character operation program executed by the control unit 1110 then extracts the text of the words spoken by user 230 from the audio signal. This text is in natural language. The extraction of the text of the words spoken by user 230 may be performed continuously for all words, or it may be started when the user speaks within a predetermined period following a trigger keyword. For example, the trigger keyword could be when the user says "Hello" followed by the character's name. For example, if character 19051's name is "Koto," then "Hello, Koto!" can be used as the trigger keyword.

[0062] The character operation program of the artificial intelligence response output device 10010 creates an instruction (prompt) based on the text of the words spoken by the user 230, and sends the instruction to the large-scale language model server 19001 using an API. Here, the instruction may be metadata containing information written using notation such as markup format of a markup language using tags, notation using predetermined symbols such as Markdown format, or object notation of a predetermined script such as JSON. The instruction contains natural language text information as the main message. The types of instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001 include setting instruction statements that store instructions such as initial settings, and user instruction statements that reflect instructions from the user. Type identification information that identifies whether the instruction statement is a setting instruction statement or a user instruction statement may be stored in a part of the instruction statement other than the main message. When the character motion program of the artificial intelligence response output device 10010 creates an instruction sentence (prompt) based on the text of the words spoken by the user 230, it creates a user instruction sentence and sends it to the large-scale language model server 19001.

[0063] Next, the large-scale language model of the artificial intelligence in the large-scale language model server 19001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response containing natural language text information based on the result. The large-scale language model server 19001 sends the response to the artificial intelligence response output device 10010 using an API. The response contains natural language text information as the main message. Here, the response may also be metadata containing information written in the same format as the instruction mentioned above (notation using tags such as the markup format of a markup language, notation using predetermined symbols such as the Markdown format, or object notation of a predetermined script such as JSON). If the same format as the instruction mentioned above is used in the response, type identification information may be stored in a part other than the main message to indicate that it is a different type of information from the initial setting instruction and the user instruction mentioned above. For example, information indicating that it is a response from the large-scale language model may be stored.

[0064] Next, the artificial intelligence response output device 10010 receives a response from the large-scale language model server 19001 and extracts the natural language text information stored as the main message in that response. Based on the natural language text information extracted from the aforementioned response, the character operation program of the artificial intelligence response output device 10010 uses speech synthesis technology to generate natural language speech that serves as a response to the user, and outputs it from the speaker, the voice output unit 1140, so that it sounds as if it were the voice of character 19051. This process may also be described as the character "uttering".

[0065] As described above, the processing by the artificial intelligence response output device 10010 and the large-scale language model server 19001 allows for specific examples of the voice responses of character 19051 to words from user 230, as shown in conversation examples 1-5 in Figure 2C. In this way, user 230 can converse with character 19051 as if it were a real person.

[0066] As described above, with the artificial intelligence response output device 10010 or the system including the artificial intelligence response output device 10010 shown in Figure 2B, it is not necessary to install the large-scale language model itself, which requires a vast amount of data and computing resources for training, into the artificial intelligence response output device 10010 itself. Furthermore, the advanced natural language processing capabilities of the large-scale language model can be utilized via an API, enabling the character to provide more appropriate responses and engage in more suitable conversations when the user speaks to it.

[0067] Next, using Figure 2D, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2D is an example of the natural language text of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.

[0068] Furthermore, Figure 2D shows the exchange of instructions and responses in chronological order, from the display setting instruction to the first round of user instructions and their responses, up to the fourth round of user instructions and their responses.

[0069] As shown in Figure 2D, the large-scale language model server 19001 can be instructed by the configuration instructions to provide the large-scale language model of the artificial intelligence with initial settings such as the name of the large-scale language model itself, the role it should play, and the characteristics of the conversation. The user's name can also be made to understand it as an initial setting. As a result, the large-scale language model generates responses from the first round onward while adhering to the assigned role. When a user hears the voice of character 19051 based on these responses from the first round onward, it will feel as if character 19051 embodies the settings and personality of the person described in the configuration instructions. Furthermore, the large-scale language model server 19001 in this embodiment is equipped with memory that stores the content of the conversation until the series of conversations is completed, and is configured to generate responses after storing a series of user instructions and their responses. This makes it possible to realize a conversation like the one shown in Figure 2D.

[0070] Next, using Figure 2E, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2E is an example of the natural language text of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.

[0071] Figure 2E shows an example of a new conversation that takes place when user 230 speaks to character 19051 again after the series of conversations shown in Figure 2D has ended. In Figure 2E, the exchange of instructions and responses is shown chronologically, from the first round of user instructions and their responses to the third round of user instructions and their responses.

[0072] Here, "termination" of "continuation of a series of conversations" refers to the process by which the large-scale language model server 19001 erases the conversation memory it held while the series of conversations was continuing, when predetermined conditions are met. An example of predetermined conditions is when the artificial intelligence response output device 10010 instructs the large-scale language model server 19001 to "terminate" the "continuation of a series of conversations" using an instruction statement. Another example of predetermined conditions is when a predetermined amount of time has passed since the artificial intelligence response output device 10010 stopped sending instruction statements to the large-scale language model server 19001 regarding the series of conversations (timeout). Furthermore, in the connection between the artificial intelligence response output device 10010 and the large-scale language model server 19001, authentication may be lost due to factors such as communication interruption or the power of the artificial intelligence response output device 10010 being turned off while the instruction statements and responses are being exchanged after the authentication process has been performed.

[0073] Furthermore, once the "continuation of a series of conversations" ends, the large-scale language model server 19001 erases the memories of the conversations it held while the series of conversations was ongoing. Therefore, even though the conversation shown in Figure 2E takes place after the series of conversations shown in Figure 2D, the server response to the user instruction is one in which the server has no memory whatsoever of the character's name, the role to be played, the characteristics of the conversation, the user's name, etc., that were included in the setting instruction shown in Figure 2D. Similarly, the conversation shown in Figure 2E is a response in which there is no memory whatsoever of the series of conversations shown in Figure 2D. In other words, the "end" of the "continuation of a series of conversations" shown in Figure 2D means that the conversation in Figure 2E starts from a state in which the large-scale language model of the artificial intelligence of the large-scale language model server 19001 has been initialized.

[0074] This can make user 230 feel as if character 19051 has lost their memories of them, or as if they are a completely different person. From user 230's perspective, the character's response feels very unnatural, resulting in a lonely and disappointing experience. Such behavior presents a challenge in that it is impossible to ensure the identity of settings and memories such as the name, role, conversational characteristics, and personality of character 19051 displayed on the artificial intelligence response output device 10010.

[0075] Next, using Figure 2F, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2F is an example of the natural language text of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.

[0076] Figure 2F shows an example of a case where, after the continuation of the series of conversations shown in Figure 2D has ended, user 230 speaks to character 19051 again to start a new conversation. Unlike the process in Figure 2E, in the process in Figure 2F, when a new conversation is started, the artificial intelligence response output device 10010 sends a setting instruction as the first instruction to the large-scale language model server 19001. This setting instruction contains the same natural language text as the initial setting instruction in Figure 2D. This may also be referred to as the reconfiguration text. This setting instruction then contains natural language text that explains the history of past conversations. This may also be referred to as the conversation history text. The history of past conversations can be recorded by the artificial intelligence response output device 10010 as natural language text information in the storage unit 1170, linked to the date and time information of the conversation, while the continuation of the series of conversations described in Figure 2D is taking place. If there are conversations on different dates, each conversation can be recorded linked to the date and time information, and the conversation history can be accumulated. When generating the initial instruction for a later conversation, as shown in Figure 2F, the natural language text information of the conversation and the date and time information of the conversation recorded in the storage unit 1170 can be read and used to generate the instruction.

[0077] When using natural language text information from past conversation history to generate the setting instruction statement, the format can be determined somewhat freely, as this data is sent to a large-scale language model. However, as shown in Figure 2F, it is advisable to prepare natural language prefixes and suffixes such as "I said the following on [date]," and "You said the following on [date]," and then combine them with the recorded conversation's natural language text information to generate the text of the setting instruction statement. Additionally, the date and time information of the conversation read from the storage unit 1170 may be combined with the "[date]" portion and used as part of the text of the setting instruction statement.

[0078] Even if user 230 speaks to character 19051 again to initiate a new conversation after a series of conversations has ended, performing the generation and transmission process of the setting instruction text shown in Figure 2F as described above will ensure that subsequent user instruction texts reflect the character's role, name, conversational characteristics, personality, and / or conversational characteristics settings and conversation history from the previous conversation. This is preferable because it allows the user to perceive a greater degree of consistency in the character's role, name, conversational characteristics, or personality settings and memories from the previous conversation.

[0079] Next, using Figure 2G, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2G is an example of the natural language text of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.

[0080] Figure 2G shows an example of a series of conversations in the same conversation as shown in Figure 2F, specifically the first round of user instructions and their responses, followed by the third round of user instructions and their responses. In Figure 2G, the exchange of instructions and responses is shown chronologically. The content of the setting instructions is the same as shown in Figure 2F, so repeated descriptions are omitted.

[0081] As shown in the natural language text of the server response in the table in Figure 2F, by using the setting instructions shown in Figure 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 will reflect the settings and conversation history of the character, such as the character's role, name, conversational characteristics, or personality, at the time of the previous conversation. This is preferable because, from the user's perspective, it is perceived that the identity of the character's settings and memories, such as the character's role, name, conversational characteristics, or personality, at the time of the previous conversation is better maintained. This can also be called pseudo-identity of the character from the user's perspective, as the character can be perceived as identical.

[0082] Furthermore, from the user's perspective, they can share memories with the character, resulting in a more enjoyable character conversation experience.

[0083] Next, using Figure 2H, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2H shows an example of operation in which the character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among multiple character candidates. The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 can switch the displayed character based, for example, on operation inputs input to the operation input unit 1107 or operations detected by the touch operation input sensor of the display unit 10011.

[0084] In the example in Figure 2H, in addition to character 19051 (named "Koto") used in the explanations of Figures 2A to 2G, characters 19052 (named "Tom") and 19053 (named "Necco") are shown. Characters 19051 (named "Koto") and 19052 (named "Tom") are human characters, while character 19053 (named "Necco") is a cat character. Switching the display of characters on the display unit 10011 can be done by rendering different virtual 3D characters for each character and then switching the displayed image on the display unit 10011.

[0085] Furthermore, when the character operation program executed by the control unit 1110 switches the display of the characters shown on the display unit 10011, it is preferable that the synthesized voice used for each character's "utterance" is also changed. This can be done by pre-storing synthesized voice data with corresponding voice to each character in the storage unit 1170, and then performing the synthesized voice change process when switching the display of the characters.

[0086] In the example shown in Figure 2H, the system is configured so that user 230 can converse with any of the characters. The artificial intelligence response output device 10010 in Figure 2H assigns different roles, names, conversational characteristics, or personalities to each of these characters. Furthermore, the memories of each character based on their conversation history are managed separately for each character.

[0087] Therefore, the artificial intelligence response output device 10010 constructs the database shown in Figure 2I in the storage unit 1170, and uses this database to manage character settings and the character's conversation history.

[0088] Next, using Figure 2I, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2I is an explanatory diagram of the database 19200 for managing character settings and character conversation history for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.

[0089] The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 constructs the database 19200 in the storage unit 1170, for example. The character ID is an identification number that identifies each of the multiple characters that can be displayed by the artificial intelligence response output device 10010, and may be a natural number or use the alphabet, etc. The name is data of the name of each of the multiple characters that can be displayed by the artificial intelligence response output device 10010.

[0090] The initial setup instruction is natural language text information that describes the settings of each of the multiple characters that can be displayed by the artificial intelligence response output device 10010, such as their role, name, conversational characteristics, or personality. Since this initial setup instruction is the main data of the setting instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, it is desirable that the content be such that the large-scale language model of the artificial intelligence in the large-scale language model server 19001 can read it directly.

[0091] The conversation history, numbered 1, 2, ..., is a record of the conversations between each character and the user, and is recorded separately for each character. Since this conversation history will be included in the natural language text information, which is the main data of the setting instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, it is desirable that the content be such that the large-scale language model of the artificial intelligence in the large-scale language model server 19001 can read it directly.

[0092] The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010, when the character displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched, uses the database 19200 in Figure 2I to select and switch the initial setting instruction text and conversation history used for the natural language text information, which is the main data of the setting instruction text sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, so that they correspond to the character displayed on the display unit 10011 of the artificial intelligence response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the history of that conversation in the conversation history area of ​​the database 19200 in Figure 2I that corresponds to the character displayed on the display unit 10011.

[0093] By using the database 19200, the character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 uses the same large-scale language model of the same artificial intelligence from the same large-scale language model server 19001 to establish a conversation between the user 230 and the character. However, from the user's perspective, the uniqueness of each character's personality and other settings is preserved, and the memory of different conversations continues for each character. From the user's perspective, this is preferable because it is perceived that the identity of the character's role, name, conversational characteristics, or personality settings and memories from previous conversations is more readily maintained for each character. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.

[0094] Therefore, even when the artificial intelligence response output device 10010 is configured to switch between displaying characters from among multiple character candidates on the display unit 10011, the operation using the database 19200 described above will result in a less jarring experience for the user in conversations with each character, and will allow them to share memories with each of the multiple characters, providing a more enjoyable character conversation experience.

[0095] Furthermore, if the user is prevented from editing the initial setting instructions for multiple characters, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained in a state close to the intentions of the provider of the artificial intelligence response output device 10010 or the creator of the character's content. Alternatively, the user may be allowed to edit the initial setting instructions for characters in response to input from the operation input unit 1107 or the like. In this case, the user can customize the character's role, name, conversational characteristics, or personality, and converse with a character they have set up themselves. In this case, the character's 3D model, its rendered image, and the type of synthesized voice may also be replaced accordingly.

[0096] Next, using Figure 2J, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service using a character conversation device with the artificial intelligence response output device 10010, or a character conversation system with the artificial intelligence response output device 10010 and the large-scale language model server 19001, at a lower cost will be described.

[0097] As explained in Figure 2B, training a large-scale language model to achieve this level of artificial intelligence for a specific application is extremely resource-inefficient. Therefore, it is more resource-efficient to generate a foundation model that can be applied to various uses by performing large-scale training, and then making it available on various devices via an API (Application Programming Interface). In this case, the provider of the large-scale language model often recovers the cost used to train the model from the user of the device as a usage fee for the device's API. In natural language models, the API usage fee is often charged based on the number of tokens (word units that divide sentences) that are processed.

[0098] Therefore, in the artificial intelligence response output device 10010 of Embodiment 2 of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010, and a character conversation service using a character conversation system with the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.

[0099] For example, by using the processing and configurations shown in Examples 1 to 3 in the table in Figure 2J, it is technically possible to reduce the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API.

[0100] Example 1 is an example of a method to reduce the number of tokens in the conversation history text stored and transmitted in the API configuration instructions, specifically by using document summarization processing to shorten the conversation history text and reduce the number of tokens. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. While document summarization can be performed at the start of the next conversation, it is more time-efficient to perform it at the end of a "series of conversations".

[0101] Alternatively, the text summarization process may be requested from the large-scale language model server 19001 itself. However, in this case, the token saving effect is low. Therefore, for example, if the second server 19002 provides natural language text summarization processing via an API at a lower cost than the large-scale language model server 19001, the text summarization processing can be requested from the second server 19002 via the API, and the text summary of the conversation history can be stored in a configuration instruction message to the large-scale language model server 19001 and transmitted.

[0102] Furthermore, if only text summarization processing is required, it can also be done on the terminal side. The control unit 1110 may execute a document summarization program, which is loaded into the memory 1109 of the artificial intelligence response output device 10010, to perform text summarization. In this case, the token saving effect is high. Also, even if the conversation history becomes long, by specifying an upper limit on the number of characters after summarization in the text summarization processing, the upper limit on the length of the conversation history text is determined, so an upper limit on tokens can be set, and token saving is possible.

[0103] Furthermore, since the amount of text information for initial character settings, such as character roles, names, conversational characteristics, or personalities, does not increase as much as the amount of conversation history, it is efficient and preferable to maintain the text information in the initial character setting instructions while reducing the number of text tokens in the conversation history.

[0104] The process described in Example 1 can be carried out by the character motion program executed by the control unit 1110, which controls each part.

[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored and transmitted in the API configuration instructions. For example, the number of tokens can be reduced by deleting older conversation histories from the conversation history with the character recorded in the storage unit 1170. Specifying an upper limit on the number of characters in the conversation history determines the upper limit on the length of the conversation history text, thus setting an upper limit on the number of tokens and enabling token saving. Alternatively, a predetermined period for the conversation history can be specified, and conversation histories exceeding that period can be deleted. In this case as well, token saving is possible. In Example 2 as well, since the text information of the character's initial settings, such as the character's role, name, conversation characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information in the character's initial settings instructions and reduce the number of tokens in the conversation history text information.

[0106] The process described in Example 2 can be carried out by the character motion program executed by the control unit 1110, which controls each part.

[0107] Example 3 is a method to reduce the number of tokens by reducing the frequency of sending configuration instructions using the API. Specifically, after the device power is turned on or after the display character is switched, and even after the video settings and synthesized voice settings for the displayed character are completed, the configuration instructions are not sent in advance. Instead, the configuration instructions are sent to the large-scale language model server 19001 only when the control unit 1110 determines that the natural language text information contained in the user's speech picked up by the microphone 1139 is text information that should be used with a large-scale language model of artificial intelligence. This reduces the frequency of sending configuration instructions to the large-scale language model server 19001 and thus reduces the number of tokens.

[0108] Specifically, for example, after the device is powered on (ON) or after an operation input to switch the displayed character, the display unit 10011 is displayed on the display unit 10011 as shown in Figure 2H, due to the display processing of the display unit 10011 under the control of a character operation program executed by the control unit 1110. At this time, for example, if a synthesized voice for the character 19051's appearance is stored in the storage unit 1170 or the like, the synthesized voice for the character's appearance, such as "Good morning. I'm Koto," "Good afternoon. I'm Koto," or "Good evening. I'm Koto," may be output from the speaker, which is the audio output unit 1140. At this time, the image of character 19051 is already set as the image of the character displayed on the display unit 10011, and the synthesized voice output from the speaker, which is the audio output unit 1140, is set to the synthesized voice corresponding to character 19051.

[0109] Here, as already explained, the inference processing of the large-scale language model of artificial intelligence in the large-scale language model server 19001 also takes time if the instruction sentence is long. In particular, if the setting instruction sentence includes text information about past conversation history, the number of tokens in the instruction sentence increases, and the inference processing time becomes especially long. The setting instruction sentence itself and its response are not output to the user 230. From the user's response to the setting instruction sentence, the synthesized speech as the character's "utterance" is output from the speaker, which is the speech output unit 1140. At first glance, it seems preferable to send the setting instruction sentence in advance from the artificial intelligence response output device 10010 to the large-scale language model server 19001 and complete the inference processing of the large-scale language model for the setting instruction sentence in advance, because this would speed up the output of the synthesized speech of the character 19051's "utterance" after the user 230 speaks to the character 19051.

[0110] However, if the large-scale language model server 19001 is pre-processed by sending a setting instruction to the setting instruction before user 230 speaks, and the inference processing of the large-scale language model for the setting instruction is completed in advance, for example, if user 230 turns off the power of the artificial intelligence response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or if user 230 switches the displayed character from character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, then the number of tokens processed by the large-scale language model server 19001 after the setting instruction was pre-processed will be the number of processing tokens for which usage fees have been wasted. This hinders the provision of character conversation devices using the artificial intelligence response output device 10010, and character conversation services using the character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001, to users at a lower cost.

[0111] Therefore, it is desirable that the artificial intelligence response output device 10010, after the device is powered on (ON) or after an operation input to switch the displayed character, controls the character operation program executed by the control unit 1110 to set the image of character 19051 as the image of the character displayed on the display unit 10011, and sets the synthesized voice output from the speaker, which is the sound output unit 1140, to the synthesized voice corresponding to character 19051, but continues not to send setting instruction text to the large-scale language model server 19001 until it recognizes that the user 230 is going to speak to character 19051.

[0112] Here, the point at which it is recognized that user 230 is speaking to character 19051 may be, for example, the point at which the trigger keyword described in Figure 2B is detected, or the point at which the text of the words spoken by user 230 is extracted. In this way, the number of processing tokens that unnecessarily waste usage fees can be reduced, and character conversation services using the character conversation device with artificial intelligence response output device 10010, or the character conversation system with artificial intelligence response output device 10010 and large-scale language model server 19001, can be provided to users at a lower cost.

[0113] Furthermore, even after the point in time when the system recognizes that user 230 is speaking to character 19051, it is desirable to continue not sending setting instructions to the large-scale language model server 19001, for example, if the text information extracted from user 230's voice picked up by microphone 1139 corresponds to preset keywords that do not require inference processing by a large-scale language model. Specifically, examples of preset keywords include those used by user 230 to request character 19051 to react, such as by performing an animation or emitting synthesized speech, such as "Try jumping" or "Try dancing." In this case, the character motion program executed by the control unit 1110 can read the motion data, animation video, and / or synthesized speech data corresponding to the reaction stored in the storage unit, and use this data to generate the video to be displayed on the display unit 10011 and to output synthesized speech from the speaker, which is the audio output unit 1140.

[0114] Such processing does not necessarily require the inference processing of the large-scale language model on the large-scale language model server 19001. If, after such processing, the user 230 turns off the power of the artificial intelligence response output device 10010 by operating via the touch input sensor of the operation input unit 1107 or the display unit 10011, or if, for example, the user 230 switches the displayed character from character 19051 to another character by operating via the touch input sensor of the operation input unit 1107 or the display unit 10011, if the setting instruction statement is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be a waste of usage fees.

[0115] Therefore, even after the point in time when the system recognizes that user 230 is speaking to character 19051, it is desirable to continue not sending the configuration instruction to the large-scale language model server 19001 until, for example, it is determined whether the text information extracted from user 230's voice picked up by microphone 1139 corresponds to text information of preset keywords that do not require inference processing of the large-scale language model. Only when this determination determines that inference processing of the large-scale language model is necessary should the configuration instruction be sent to the large-scale language model server 19001 and the inference processing of the large-scale language model be advanced.

[0116] Furthermore, the process described in Example 3 can be carried out by the character motion program executed by the control unit 1110, which controls each part.

[0117] As described above, the methods for reducing (saving) the number of processing tokens for large-scale language models, as shown in the examples in Figure 2J, allow us to provide users with character conversation services using the artificial intelligence response output device 10010 and the character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.

[0118] Next, an example of the display of the character conversation device (artificial intelligence response output device 10010) of Embodiment 2 of the present invention will be described using Figure 2K. The example in Figure 2K shows an example of displaying the response from the large-scale language model to the user instruction sentences described in Figures 2A to 2J on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, it is an example of displaying the text 10063, which is the response from the large-scale language model, together with the image of character 19051 on the display unit 10011. The text 10063, which is the response from the large-scale language model, may be displayed superimposed in front of the image of character 19051, as shown in Figure 2K. Alternatively, the text 10063, which is the response from the large-scale language model, may be displayed together with the image of character 19051 without being superimposed on it.

[0119] The display in Figure 2K is just one example, but for instance, if user 230 adjusts the volume of the voice output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to the minimum or sets the voice output to OFF by operating via the touch input sensors of the operation input unit 1107 or the display unit 10011, user 230 will not be able to confirm the response from the large-scale language model by voice.

[0120] Therefore, in this case, the control unit 1110 may be controlled to start a display mode in which the text 10063, which is the response from the large-scale language model, is displayed together with the image of the character 19051, as shown in Figure 2K. In this way, even when it is desired to reduce voice output, the user 230 can more conveniently use the character conversation device (artificial intelligence response output device 10010). Furthermore, the user 230 may be configured to manually switch ON / OFF the display mode in which the text 10063, which is the response from the large-scale language model, is displayed together with the image of the character 19051, by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011.

[0121] Next, using Figure 2L, we will explain an example of a response template database (response template DB) in a character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in Figures 2H and 2I. In the example in Figure 2L, the condition numbers and condition contents are the same as in Figure 1C. For these conditions, in the example in Figure 2L, individual response templates are set for each of the multiple characters. For example, for each of the three characters described in Figures 2H and 2I—character 1: Koto, character 2: Tom, and character 3: Necco—a response template for each condition is stored. The output control of the response templates is the same as in Figure 1C, so a repeated explanation will be omitted.

[0122] In the example shown in Figure 2L, the control unit 1110 selects a corresponding predefined response from the predefined response database (predefined response DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses it to control the output as a response issued by the character. For example, in the example of the predefined response database (predefined response DB) in Figure 2L, even under the same conditions, the predefined response is changed to an expression or content that corresponds to the personality of each character. As a result, the character conversation device (artificial intelligence response output device 10010) can provide the user with conversations that correspond to the personality of the displayed character. The user can feel that each character is a being with a more consistent personality. This makes it possible to realize a character conversation device (artificial intelligence response output device 10010) that gives multiple characters a greater sense of reality.

[0123] The response template database (response template DB) shown in Figure 2L, as described above, is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 can use it. However, the response template database (response template DB) shown in Figure 2L may also be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 should generate responses using the said response template database (response template DB). The control unit of the large-scale language model server 19001 should send the response generated using the said response template database (response template DB) to the artificial intelligence response output device 10010 instead of the response generated by the large-scale language models stored in each server. In this way, even if the artificial intelligence response output device 10010 is not equipped with a response template database (response template DB), it is possible to generate responses using the response template database (response template DB).

[0124] As described above, the character conversation device and character conversation system according to Embodiment 2 can reduce the sense of incongruity that users feel when conversing with the character displayed on the artificial intelligence response output device 10010. Furthermore, the character conversation device and character conversation system according to Embodiment 2 can provide character conversation services to users at a lower cost.

[0125] In the above description of Example 2, an example was described in which the large-scale language model possessed by the large-scale language model server 19001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in Figure 1B, and the large-scale language model possessed by the local LLM processing unit 10028 may be used instead of the large-scale language model possessed by the large-scale language model server 19001. In this case, in the above description of Example 2, the large-scale language model possessed by the large-scale language model server 19001 should be read as the large-scale language model possessed by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).

[0126] In this case as well, the sense of incongruity the user feels from conversations with the character displayed on the artificial intelligence response output device 10010 can be reduced. Furthermore, if the large-scale language model of the local LLM processing unit 10028 is used instead of the large-scale language model of the large-scale language model server 19001, the need to consider usage fees based on the number of processing tokens will decrease. However, even with the large-scale language model of the local LLM processing unit 10028, reducing the number of processing tokens can reduce the consumption of resources such as power required for inference. In this case, a character conversation service with lower power consumption can be provided to the user.

[0127] In the above description of Example 2, an example was described in which the conversation history with the character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). Alternatively, the conversation history with the character may be recorded and stored in a second server 19002 connected to the Internet 19000 or other cloud server. In this case, when a new conversation is started between the user and the character, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or other cloud server, obtains (downloads) the past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create instruction sentences for the large-scale language model. The specific method of using past conversation history to create instruction sentences for the large-scale language model is as described in the figures of Example 2, so a repeated explanation will be omitted.

[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) only needs to transmit (upload) the conversation history of the character up to that point to the second server 19002 or other cloud server at predetermined times, such as each time a conversation takes place between the user and the character, or when the conversation between the user and the character ends. In other words, the character conversation device (artificial intelligence response output device 10010) uploads the conversation history with the character to the second server 19002 or other cloud server at predetermined timings, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) downloads the latest conversation history from the second server 19002 or other cloud server and uses it to generate instruction sentences for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) used by the user the previous day and the character conversation device (artificial intelligence response output device 10010) that the user will use now are different devices capable of displaying the same character. When the user and the same character between these different devices converse multiple times at different times, it is possible to realize a conversation that appears as if the character's memory has been pseudo-carried over from the previous conversation, which is more preferable for the user.

[0129] The process described above, in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with a character to a second server 19002 or other cloud server to simulate the transfer of the character's memory, is also effective when dealing with a database 19200 containing the conversation histories of multiple characters, as described in Figures 2H and 2I. In other words, by configuring the database 19200 described in Figure 2I to be uploaded and downloaded to a second server 19002 or other cloud server, it is possible to achieve a conversation that simulates the transfer of each character's memory from the previous conversation, not only for one character but for multiple characters, when the user converses with each character multiple times at different times between different devices, between different characters, which is more preferable for the user.

[0130] <Example 3> Next, Embodiment 3 of the present invention is an improvement on the character conversation device (artificial intelligence response output device 10010) and character conversation system described in the figures of Embodiment 2. In this embodiment, the differences from Embodiment 2 will be explained, and repeating explanations of configurations similar to those in these embodiments will be omitted.

[0131] Similar to Example 2, the character in Example 3 can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can become an artificial intelligence (AI) assistant for the user. In this case, the character conversation device or character conversation system in this embodiment may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.

[0132] An example of a character conversation device and character conversation system according to Embodiment 3 of the present invention will be described using Figure 3A. In the character conversation system of Embodiment 3, a large-scale language model server 20001 is provided instead of the large-scale language model server 19001 in Figure 2A, and it is connected to the Internet 19000.

[0133] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence, but it is a multimodal large-scale language model artificial intelligence that can process not only natural language text information, which could be processed by the large-scale language model server 19001, but also other types of information besides natural language text information.

[0134] Furthermore, the artificial intelligence response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (artificial intelligence response output device 10010) of Embodiment 2, as an example.

[0135] In Embodiment 3, the artificial intelligence response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the internet 19000 using an API.

[0136] The character conversation system in Example 3 includes a mobile information processing terminal 20010 used by user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.

[0137] Here, an example of a mobile information processing terminal 20010 will be described using Figure 3B. The mobile information processing terminal 20010 includes a display panel 20011 which is a touch operation input panel, a control unit 20012, an external power input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.

[0138] The display panel 20011 is equipped with a touch input sensor and can accept touch input from the user 230's finger. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel and can display images. The display panel 20011 may also be called a display unit.

[0139] The communication unit 20020 can be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and memory, and the control unit controls the display panel 20011 and the communication unit 20020. Furthermore, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000 using one of the communication methods of the communication unit 20020. As a result, the mobile information processing terminal 20010 can communicate with various servers connected to the Internet 19000.

[0140] Power supply 20014 converts AC current input from an external source via the external power input interface 20013 into DC current and supplies the necessary DC current to each part of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied by power supply 20014. In addition, the secondary battery 20015 supplies power to each part that requires power via the external power input interface 20013 when external power is not supplied.

[0141] The video signal input section 20023 receives video data by connecting an external video output device. Various digital video input interfaces are possible for the video signal input section 20023. For example, it can be configured with an HDMI (High-Definition Multimedia Interface) standard video input interface, a DVI (Digital Visual Interface) standard video input interface, or a DisplayPort standard video input interface. Alternatively, analog video input interfaces such as analog RGB or composite video may be provided. The video signal input section 20023 may also use various USB interfaces.

[0142] The audio signal input unit 20024 receives audio data by connecting an external audio output device. The audio signal input unit 20024 may be configured as an HDMI audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface, etc. The audio signal input unit 20024 may also be various USB interfaces, etc. In the case of an HDMI interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface with integrated terminals and cables.

[0143] The audio output unit 20021 is capable of outputting audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 is also capable of outputting audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured as a speaker. In addition, the audio output unit 20021 may output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output as a digital signal to an external device, such as the Audio Return Channel function specified in the HDMI standard.

[0144] Microphone 20022 is a microphone that picks up sounds from the surrounding area of ​​the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may be configured to record human voices, such as the user's voice, and the control unit 20012, described later, may perform speech recognition processing on the generated audio signal to obtain text information from the audio signal.

[0145] The imaging unit 20025 is a camera having an image sensor. The camera may be provided on the front of the display panel 20011 side of the mobile information processing terminal 20010, or on the back of the display panel 20011 side. Both a front camera and a rear camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a rear camera.

[0146] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be composed of a magnetic recording medium such as a hard disk drive (HDD) or a semiconductor memory such as a solid-state drive (SSD). For example, the storage unit 20016 may have various types of information, such as video data, image data, and audio data, pre-recorded in it at the time of product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices or external servers via the communication unit 20020. The video data, image data, etc., recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc., recorded in the storage unit 20016 may also be output to external devices or external servers via the communication unit 20020.

[0147] The video control unit 20017 performs various controls related to the video signals input to the display panel 20011. The video control unit 20017 may also be called a video processing circuit and may be composed of hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be called a video processing unit or image processing unit. For example, the video control unit 20017 controls video switching, such as determining which video signal to input to the display panel 20011 from among the video signals to be stored in memory 20026 and the video signals (video data) input to the video signal input unit 20023. The video control unit 20017 may also perform control to perform image processing on the video signals input from the video signal input unit 20023 and the video signals to be stored in memory 20026. Examples of image processing include scaling processing, such as enlarging, reducing, and transforming images; brightness adjustment processing, which changes the brightness; contrast adjustment processing, which changes the contrast curve of an image; and retinex processing, which decomposes an image into its light components and changes the weighting of each component.

[0148] The attitude sensor 20018 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each connected part.

[0149] The non-volatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the non-volatile memory 20027 includes, for example, data for various operations displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for objects used by the user. Memory 20026 stores video data and device control data displayed on the display panel 20011. The control unit 20012 may read various software from the storage unit 20016, expand it into memory 20026, and store it there.

[0150] The control unit 20012 controls the operation of each connected component. The control unit 20012 may also work in cooperation with a program stored in memory 20026 to perform calculations based on information acquired from each component within the mobile information processing terminal 20010.

[0151] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described using Figure 3C. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. In Embodiment 3 as well, the character conversation device (artificial intelligence response output device 10010) loads the character operation program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character operation program, thereby enabling the various processes described below to be realized.

[0152] In Example 2, the actions performed by User 230 to the character conversation device (artificial intelligence response output device 10010) were mainly through User 230's voice. In Example 2, the character conversation device (artificial intelligence response output device 10010) performed a series of operations starting with the process of picking up User 230's voice with a microphone. In contrast, the character conversation device (artificial intelligence response output device 10010) in Example 3 is also capable of performing the series of operations described in Example 2, starting with the process of picking up User 230's voice with a microphone. In addition, in the character conversation device (artificial intelligence response output device 10010) in Example 3, User 230 can perform actions to the character conversation device (artificial intelligence response output device 10010) through user operation via the operation input unit 1107 in Figure 1B. Here, an example of the operation input unit 1107 in Figure 1B is a mouse, keyboard, touch panel, etc.

[0153] Furthermore, in the character conversation device (artificial intelligence response output device 10010) of Embodiment 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by touch operation detected by the user's touch operation input sensor on the display unit 10011 in Figure 1B.

[0154] Furthermore, user 230 can also input user 230's operation input to the character conversation device (artificial intelligence response output device 10010) by operating the mobile information processing terminal 20010 and communicating from the mobile information processing terminal 20010 to the character conversation device (artificial intelligence response output device 10010).

[0155] Alternatively, the display panel 20011 of the mobile information processing terminal 20010 may display an information-storage image, such as a two-dimensional code containing information that the user wants to convey to the character conversation device (artificial intelligence response output device 10010), and the imaging unit 1180 of the character conversation device (artificial intelligence response output device 10010) may capture this display. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information-storage image, such as a two-dimensional code, captured by the imaging unit 1180, and obtain the information. Alternatively, the display panel 20011 of the mobile information processing terminal 20010 may display an image that the user wants to convey to the character conversation device (artificial intelligence response output device 10010), and the imaging unit 1180 of the character conversation device (artificial intelligence response output device 10010) may capture this display. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180 and obtain the result of said image recognition processing.

[0156] Thus, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are greater than those of the character conversation device (artificial intelligence response output device 10010) described in Example 2. As a result, the character conversation device (artificial intelligence response output device 10010) of Example 3 can acquire the results of actions performed by the user 230 other than the user's voice, and generate an instruction sentence (prompt) to send to the large-scale language model server 20001 based on that. This makes it possible to more favorably include types of information other than natural language text information extracted from the user's voice in the instruction sentence sent to the large-scale language model server 20001. Examples of types of information other than natural language text information extracted from the user's voice include images, videos, and audio.

[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment sends instruction texts to the large-scale language model server 20001 using an API. In this embodiment as well, instruction texts may be metadata containing information written using notation such as markup format of a markup language using tags, notation such as Markdown format using predetermined symbols, or object notation of a predetermined script such as JSON. In this embodiment as well, there are two types of instruction texts: setting instruction texts that store instructions such as initial settings, and user instruction texts that reflect instructions from the user. Type identification information that identifies whether an instruction text is a setting instruction text or a user instruction text may be stored in a part of the instruction text other than the main message. In this case, the instruction text includes natural language text information as the main message. Furthermore, in this embodiment, in addition to natural language text information, the main message of the instruction text may include non-natural language information sources such as images, videos, or audio as a type of information other than natural language text information. A specific method for including non-natural language information sources in instruction texts will be described later.

[0158] The large-scale language model server 20001 in this embodiment has a multimodal large-scale language model that can process non-natural language information sources in conjunction with natural language text information. The large-scale language model server 20001 receives an instruction sentence from a character conversation device (artificial intelligence response output device 10010). Based on the instruction sentence, the multimodal large-scale language model performs inference and generates a response that includes natural language text information as a result of the inference. Here, since the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.

[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or audio stored as the main message in the response. The character operation program of the character conversation device (artificial intelligence response output device 10010) may use speech synthesis technology to generate natural language audio as a response to the user based on the natural language text information extracted from the aforementioned response, and output it from the audio output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.

[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display natural language characters that serve as a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010), based on the natural language text information extracted from the aforementioned response. In this case, the characters may be displayed together with character 19051, superimposed on the image of character 19051, or displayed in place of the image of character 19051. The video control unit 1160 may perform these specific processes.

[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) in order to present it to the user, based on the image information of the non-natural language information source extracted from the aforementioned response. In this case, the image may be displayed together with character 19051, superimposed on the image of character 19051, or displayed in place of the image of character 19051. These specific processes can be executed by the image control unit 1160.

[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video information of the non-natural language information source extracted from the aforementioned response on the display screen of the character conversation device (artificial intelligence response output device 10010) in order to present it to the user. In this case, the video may be displayed together with character 19051, superimposed on the video of character 19051, or displayed in place of the video of character 19051. These specific processes can be executed by the video control unit 1160.

[0163] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output speech generated based on the speech information of the non-natural language information source extracted from the aforementioned response from the speech output unit 1140, which is a speaker.

[0164] As described above, with the character conversation device (artificial intelligence response output device 10010) shown in Figure 3C, or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001, it is not necessary to install the large-scale language model itself, which requires a massive amount of data and computing resources for training, within the character conversation device (artificial intelligence response output device 10010). Furthermore, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via an API. In addition to responses based on natural language text, responses based on non-natural language information sources can be provided in response to user actions towards the character, enabling more appropriate conversations.

[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described using Figure 3D. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3D shows an example of the natural language text and non-natural language information source such as an image of the main message of the instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001, and an example of the natural language text and non-natural language information source such as an image of the main message of the server response that is the response to it. In this embodiment, images, videos, audio, etc. can be used as non-natural language information sources, but Figure 3D shows an example of an image as a non-natural language information source.

[0166] Furthermore, Figure 3D shows the exchange of instructions and responses in chronological order, from the first round of setting instructions and user instructions and their responses to the second round of user instructions and their responses. Here, the instructions and responses shown in Figure 3D include non-natural language information sources 20061 and 20062, which were not shown in Figure 2D of Example 2. In the example of Figure 3D, both non-natural language information sources 20061 and 20062 are images.

[0167] In Figure 3D, for the sake of simplicity, an image of the non-natural language information source 20061 is shown embedded within the instruction text. However, there are multiple methods for transmitting or specifying the data of the non-natural language information source 20061 in the instruction text sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) can use any one of these methods, or switch between them. An example of each method will be explained below.

[0168] The first method for transmitting or specifying non-natural language information source data in an instruction is used, for example, when the non-natural language information source to be specified is located on a server or other location connected to a network such as the Internet. A specific example of the first method is to specify a non-natural language information source file located on a network such as the Internet using information such as tags and symbols within the instruction, along with the network location information (so-called URL, etc.) and file name.

[0169] For example, a tag used to specify an image in a markup language. <img src=""****”"> By using this tag and writing the location and filename information of the image file in the **** part, you can specify an image that exists on a network such as the internet. Alternatively, you can use a tag that specifies a video in a markup language. <video src=""****”">You can also specify a video that exists on a network such as the internet by using the **** part and writing the location information and file name information of the video file. Alternatively, you can use a tag that specifies audio in a markup language. <audio src=""****”">You can specify audio files located on a network such as the internet by using the format and writing the location and filename information of the audio file in the **** section. Alternatively, if using JSON notation, you can specify images located on a network such as the internet by preparing a key such as img_src and writing the location and filename information of the image file as the value. For video and audio files, you just need to prepare the respective keys and values. The example given is just one example, and you may use other proprietary formats. In any case, the information specifying the location and filename information of the non-natural language information source file should be stored in the instruction statement.

[0170] As in the first method, when information specifying the location and filename of a non-natural language information source file is stored in the instruction statement, the instruction statement itself does not need to store the data of the non-natural language information source file. Therefore, the amount of data in the instruction statement can be reduced. In the first method, the large-scale language model server 20001 that receives an instruction statement specifying non-natural language information source data can use the location and filename information of the non-natural language information source file stored in the instruction statement to obtain the non-natural language information source file located on a server or other location connected to a network such as the Internet.

[0171] Here, we will explain how location information and file name information are input when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence using the first method. In Figure 3C, we have explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) have increased compared to Embodiment 2, in addition to the voice of the user 230. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, etc., through user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in Figure 1B.

[0172] Furthermore, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may work with the memory 1109 to execute a web browser program and display the GUI of the web browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). User operations on the GUI of the web browser program may be received via the operation input unit 1107 (for example, mouse, keyboard, touch panel) or by user touch operations detectable by the touch operation input sensor of the display unit 10011, and non-natural language information source data such as images, videos, and audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction statement. In this case, the web browser program should acquire the location information and file name information of the non-natural language information source data and pass it to the character operation program.

[0173] Alternatively, user 230 may operate the mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) and input location information such as a URL for specifying non-natural language information source data into the character conversation device (artificial intelligence response output device 10010). Alternatively, as explained in Figure 3C, location information such as a URL for specifying non-natural language information source data, file name information, etc., may be input by displaying an information-storing image such as a two-dimensional code on the display panel 20011 of the mobile information processing terminal 20010, performing image recognition processing on the image captured by the imaging unit 1180 of the character conversation device (artificial intelligence response output device 10010), and obtaining the result of the image recognition processing.

[0174] Furthermore, the use of the first method for transmitting or specifying non-natural language information source data in an instruction is not limited to cases where the non-natural language information source file already exists on a server or other location connected to a network such as the Internet. For example, if it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) in an instruction, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to a second server 19002 via the Internet 19000 and include the Internet location information (so-called URL, etc.) and file name of the uploaded non-natural language information source data on the second server 19002 in the instruction. In this case, the second server 19002 functions as a so-called intermediate server.

[0175] Similarly, if it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in the instruction text, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the internet location information (so-called URL, etc.) and file name of the non-natural language information source data on the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include the acquired internet location information (so-called URL, etc.) and file name of the non-natural language information source data uploaded to the second server 19002 in the instruction text.

[0176] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may work in cooperation with the memory 1109 and the storage unit 1170 to construct a media server within the character conversation device (artificial intelligence response output device 10010) that can be accessed from other servers via the internet 19000. In this case, when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction statement using the first method, it only needs to store in the instruction statement location information on the internet (such as a URL) indicating the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself, and the file name of the corresponding non-natural language information source data.

[0177] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction is, for example, simply to store (attach) the non-natural language information source data itself in the instruction (prompt) and send it. Generally, non-natural language information source data such as images, videos, and audio are larger in data size than natural language text information. Therefore, in this case, the data size of the instruction (prompt) itself will be larger than in the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) can store the non-natural language information source data that it wants to store (attach) in the instruction (prompt) in memory 1109, and when sending the instruction (prompt), it can store (attach) the data in the instruction (prompt) via the communication unit 1132 and output it to the large-scale language model server 20001. The non-natural language information source data that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the internet 19000, acquired by the communication unit 1132 from the mobile information processing terminal 20010, or read from the storage unit 1170 and stored in memory 1109.

[0178] As described above, the character conversation device (artificial intelligence response output device 10010) is capable of transmitting or specifying non-natural language information source data using instruction sentences.

[0179] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources in conjunction with natural language text information. As shown in the example in Figure 3D, through the first round of user instructions, it can acquire images of a swimming pool and poolside, which are non-natural language information sources 20061, and natural language text information. As a result of this inference, it can output natural language text information as shown in the figure, in response to the first round of user instructions.

[0180] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example in Figure 3D, in the response to the second round of user instructions, the large-scale language model server 20001 can include the non-natural language information source 20062 generated by the inference of the multimodal large-scale language model in its response and send it to the character conversation device (artificial intelligence response output device 10010). In Figure 3D, the non-natural language information source 20062 is an example of an image with a circle added to the image of a swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 3D, but may also be a video or audio.

[0181] When the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, the method can be the first method or a method similar to the second method used by the character conversation device (artificial intelligence response output device 10010) to transmit or specify non-natural language information source data in the instruction statement.

[0182] Specifically, in a method similar to the first method described above, the large-scale language model server 20001 may store information specifying the location and file name of the non-natural language information source file in the instruction statement in its response. The non-natural language information source 20062 itself, such as images, videos, and audio, may be kept by the large-scale language model server 20001, or it may be transferred to and kept by the second server 19002, which functions as an intermediate server. In either case, the large-scale language model server 20001 may store information specifying the location and file name of the non-natural language information source file in the instruction statement in its response. The character conversation device (artificial intelligence response output device 10010) that has received the response may use the location and file name information of the non-natural language information source file described in the instruction statement to access the large-scale language model server 20001 or the second server 19002 to obtain the non-natural language information source 20062.

[0183] Furthermore, specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and send it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction and use it for various outputs to the user 230.

[0184] As described above using Figure 3D, the operation of the character conversation device (artificial intelligence response output device 10010) and character conversation system of Embodiment 3 involves the transmission and reception of instruction sentences and responses between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230, enabling conversation using non-natural language information such as images, videos, and audio. This makes it possible to achieve more sophisticated and natural conversations, as shown in each message in Figure 3D.

[0185] Next, using Figure 3E, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3E is an example of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.

[0186] Figure 3E shows an example of a new conversation that takes place after the series of conversations shown in Figure 3D has ended, when user 230 speaks to character 19051 again. In the example in Figure 3E, no processing using the conversation history is performed, as explained in Figures 2F, 2G, and 2I of Example 2. Therefore, Figure 3E, like Figure 2E of Example 2, shows a response in which the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, the user's name, and the conversation history, which were included in the setting instruction, are not remembered at all.

[0187] Next, using Figure 3F, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3F is an example of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.

[0188] Figure 3F shows an example of a case where, after the series of conversations shown in Figure 3D has ended, user 230 speaks to character 19051 again to initiate a new conversation. Here, in Figure 3F, the method of storing a message explaining the history of past conversations in the setting instruction statement, as explained in Figure 2F of Embodiment 2, is also applied to the character conversation device (artificial intelligence response output device 10010) of Embodiment 3. Specifically, the message that constitutes the content of the setting instruction statement in Figure 3D is stored as a reset message in Figure 3F, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.

[0189] The large-scale language model server 20001 in Example 3 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. Therefore, in past instructions and responses, non-natural language information source data may have been transmitted or specified. Accordingly, in the example in Figure 3F, the conversation history message reflects not only the natural language text information in past instructions and responses, but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method of transmitting or specifying non-natural language information source data in the instructions in Figure 3F is the same as the transmission or specification of non-natural language information source data as explained in Figure 3D, so a repeated explanation will be omitted.

[0190] In the example in Figure 3D, the method of transmitting or specifying non-natural language source data can involve either storing (attaching) the non-natural language source data itself to the instruction statement, or not storing (attaching) the non-natural language source data to the instruction statement. The same applies to the instruction statement in Figure 3F.

[0191] Next, using Figure 3G, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3G is an example of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which forms the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.

[0192] Figure 3G shows an example of a series of conversations in the same conversation as shown in Figure 3F, specifically the first round of user instructions and their responses, followed by the third round of user instructions and their responses. In Figure 3G, the exchange of instructions and responses is shown chronologically. The content of the setting instructions is the same as shown in Figure 3F, so repeated descriptions are omitted.

[0193] As explained above, even when using the large-scale language model server 20001 which has a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information as in Example 3, even if user 230 speaks to character 19051 again to start a new conversation after a series of conversations has ended, if the setting instruction sentence generation process and transmission process shown in Figure 3F are performed, the subsequent user instruction sentence response will reflect the settings and conversation history of the character at the time of the previous conversation, such as the character's role, name, conversational characteristics, personality, and / or conversational characteristics, as shown in Figure 3G. This is preferable because it allows the user to perceive a greater degree of consistency in the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation.

[0194] Next, using Figure 3H, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3H is an explanatory diagram of the database 20200 for managing the character settings and character conversation history for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, Figure 3H uses the example described in Figure 2H of Embodiment 2 for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.

[0195] Furthermore, the database 20200, shown in Figure 3H, which manages character settings and character conversation history, has the same format as the database 19200 shown in Figure 2I of Example 2. In Figure 3H, only the differences from the database 19200 shown in Figure 2I will be explained. In addition, the contents of the character "Koto" in the database will be explained, and the contents of other characters will be omitted.

[0196] As described above, the large-scale language model server 20001 of Example 3 is a multimodal large-scale language model that can process non-natural language information sources in addition to natural language text information. Therefore, both the instruction sentences from the character conversation device (artificial intelligence response output device 10010) and the responses from the large-scale language model server 20001 include not only natural language text information but also the transmission or specification of non-natural language information source data. Accordingly, in the database 20200 shown in Figure 3H, the conversation history data records not only the natural language text information contained in these instruction sentences and responses but also the information on the transmission or specification of non-natural language information source data. The specific method of transmitting or specifying non-natural language information source data in the recording of the conversation history is the same as the specification of transmission or specification of non-natural language information source data described in Figure 3D, so a repeated explanation is omitted.

[0197] In the example in Figure 3D, there are two methods for transmitting or specifying non-natural language source data: one where the non-natural language source data itself is stored (attached) to the instruction, and another where the non-natural language source data is not stored (attached) to the instruction. The same applies to the conversation history in Figure 3H. However, in the conversation history in Figure 3H, if the method for specifying non-natural language source data involves specifying the location information and file name information of a non-natural language source file on a server located on a network such as the Internet (the second server 19002 that functions as an intermediate server or other cloud server), there is a possibility that the non-natural language source file on that server may be deleted if the conversation history period becomes long. In that case, it may become impossible to retrieve the non-natural language source file at a later date using the location information and file name information, potentially resulting in the loss of conversation record information.

[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts instruction and response messages into a conversation history and records them, it can obtain the non-natural language information source file specified in the instruction and response from a server on the network using its location information and file name information, and store it in the storage unit 1170. Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) can rewrite the location information and file name of the non-natural language information source file to location information on the internet (so-called URL, etc.) indicating the media server of the media server built within the character conversation device (artificial intelligence response output device 10010), and then record it in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself deletes the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be lost from the conversation record information, making it more suitable for preserving the conversation record.

[0199] Using the database shown in Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between displaying multiple character candidates on the display unit 10011, the user experiences less discomfort in conversations with each character, can share memories with each of the multiple characters, and can obtain the effect shown in Figure 2I of Embodiment 2, resulting in a more enjoyable character conversation experience. Furthermore, this effect can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information.

[0200] In the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, a multimodal large-scale language model artificial intelligence is used in the large-scale language model server 20001, which is capable of processing not only natural language text information but also non-natural language information other than natural language text information.

[0201] Here, communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 is conducted using an API. In a multimodal large-scale language model, it is possible that API usage fees may be charged based on the amount of data from non-natural language information sources, in addition to the number of natural language text information units called tokens that are used to divide sentences.

[0202] Therefore, in order to provide the character conversation service using the character conversation system according to this embodiment to users at a lower cost, the following modifications may be used.

[0203] In the first variation, the conversation history record in the database shown in Figure 3H also records information about the transmission or specification of non-natural language information source data. However, the character and the user exchange conversations about the non-natural language information source data using natural language text information, and the content of these conversations is recorded in natural language text information. Therefore, even if the recording of information about the transmission or specification of the non-natural language information source data is omitted in the conversation history record in the database shown in Figure 3H, the conversation about the non-natural language information source data itself will still be recorded to some extent as natural language text information. Thus, if a certain degree of information reduction is acceptable, the recording of information about the transmission or specification of the non-natural language information source data may be omitted in the conversation history record in the database shown in Figure 3H. In this case, the information about the transmission or specification of the non-natural language information source data is also omitted from the conversation history message of the setting instruction in Figure 3F. This makes it possible to reduce the amount of data of non-natural language information sources communicated using the API.

[0204] Next, as a second variation, in the recording of the conversation history in the database of Figure 3H, instead of recording information on the transmission or specification of non-natural language information source data, natural language text information describing the content of the non-natural language information source data is recorded. The natural language text information describing the content of the non-natural language information source data may be obtained, for example, by initiating a conversation between the large-scale language model of the large-scale language model server 20001 and the character conversation device (artificial intelligence response output device 10010), separate from the conversation as a character, and having the large-scale language model server 20001 describe the content of the non-natural language information source data with a predetermined character limit. Alternatively, the content of the non-natural language information source data may be obtained by having a conversation with another large-scale language model on another server, which can be used at a lower cost than the large-scale language model of the large-scale language model server 20001, and having it describe with a predetermined character limit. Furthermore, if alternative text data is prepared from the time of acquisition of the non-natural language information source data, that alternative text data may be used as the natural language text information describing the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is the tags of a markup language. <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">This is text information that is written in the **** section, etc.

[0205] Furthermore, if using JSON format notation, an object can be stored that associates the location information and filename information of the non-natural language source data, which are key-value pairs indicating the location information of the non-natural language source data, with a key corresponding to the alternative text and a value that is the alternative text data itself.

[0206] In this case as well, the recording of information regarding the transmission or specification of the non-natural language information source data can be omitted in the conversation history recording of the database in Figure 3H, and the information regarding the transmission or specification of the non-natural language information source data is also omitted from the conversation history message of the setting instruction in Figure 3F. This makes it possible to reduce the amount of data of the non-natural language information source communicated using the API.

[0207] Next, as a third variation, at the point of the first round of user instructions in Figure 3D, information on the transmission or specification of non-natural language information source data is not stored in the user instructions, but is replaced with natural language text information that describes the content of the non-natural language information source data. For example, in the first round of user instructions in Figure 3D, the information on the transmission or specification of non-natural language information source data 20061 can be replaced with natural language text information such as, "This image is of a swimming pool with a seat and parasol by the poolside. There is water in the swimming pool. There are drinks on the table next to the seat." In this case, the description may be obtained by having a conversation with another large-scale language model on another server, which can be used at a lower cost than the large-scale language model of the large-scale language model server 20001, to describe the content of the non-natural language information source data with a predetermined character limit. Alternatively, the description may be obtained from a server of various other services that can obtain an overview or description of the content of non-natural language information source data such as images, videos, and audio. Furthermore, if alternative text data is available at the time of acquisition of non-natural language source data, this alternative text data may be used as natural language text information that explains the content of the non-natural language source data.

[0208] Next, an example of a display example of the character conversation device (artificial intelligence response output device 10010) of Embodiment 3 of the present invention will be described using Figure 3I. The example in Figure 3I shows an example of displaying the response from the large-scale language model to the user instruction sentences described in Figures 3A to 3H on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying the text 10063 of the natural language information source data, the image 10064 of the non-natural language information source data, and / or the video 10065 of the non-natural language information source data, which are the response from the large-scale language model, together with the video of the character 19051 on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are the response from the large-scale language model, may be displayed superimposed in front of the video of the character 19051, as shown in Figure 3I.

[0209] Furthermore, the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on it. The display in Figure 3I is just one example, but for example, if user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to the minimum or sets the audio output to OFF by operating via the touch operation input sensor of the operation input unit 1107 or the display unit 10011, user 230 will not be able to confirm the response from the large-scale language model by voice. In this case, the control unit 1110 may control the system to start a display mode in which the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in Figure 3I.

[0210] In this way, even when the user 230 wishes to minimize voice output, the user 230 can more conveniently use the character conversation device (artificial intelligence response output device 10010). Furthermore, the user 230 may manually switch ON / OFF the display mode, which displays the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, along with the image of the character 19051, via the operation input unit 1107 or the touch operation input sensor of the display unit 10011. As shown in the display example in Figure 3I, it becomes possible to more conveniently output responses from the large-scale language model in a multimodal character conversation device (artificial intelligence response output device 10010).

[0211] As described above, the character conversation device and character conversation system according to Example 3 offer users a more advanced conversational experience that includes non-natural language information in addition to natural language information, by using a multimodal large-scale language model, in addition to the effects of the character conversation device and character conversation system according to Example 2. Furthermore, the character conversation device and character conversation system according to Example 3 can provide character conversation services to users at a lower cost.

[0212] In the above description of Example 3, an example was described in which the large-scale language model possessed by the large-scale language model server 20001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be equipped with the local LLM processing unit 10028 shown in Figure 1B, and the multimodal large-scale language model possessed by the local LLM processing unit 10028 may be used. In this case, the multimodal large-scale language model possessed by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model possessed by the large-scale language model server 20001.

[0213] In this case, in the above description of Example 3, the multimodal large-scale language model possessed by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model possessed by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010). In this case as well, by using a multimodal large-scale language model, it is possible to provide users with a more advanced conversation experience that includes non-natural language information in addition to natural language information. When using the multimodal large-scale language model possessed by the local LLM processing unit 10028 instead of the multimodal large-scale language model possessed by the large-scale language model server 20001, the need to consider usage fees based on the number of processing tokens and the amount of data of non-natural language information sources is reduced. However, even with the multimodal large-scale language model possessed by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power for inference by reducing the number of processing tokens and the amount of data of non-natural language information sources. In this case, it is possible to provide users with a character conversation service that consumes less power.

[0214] Furthermore, the configuration described in Example 2, which involves uploading and downloading the conversation history with a character and the database data containing the conversation history to and from the second server 19002 or another cloud server, can also be used in the example using a multimodal large-scale language model described in Example 3. In this case as well, when a user interacts with one or more characters at different times across different devices, it is possible to achieve a conversation that appears as if the memories of each character have been artificially carried over from the previous conversation, which is more preferable for the user.

[0215] <Example 4> Next, Embodiment 4 of the present invention is an improvement on the artificial intelligence response output device 10010, character conversation device, or system described in the figures of Embodiment 2 or Embodiment 3. In this embodiment, the differences from Embodiment 2 or Embodiment 3 will be explained, and repeating explanations of configurations similar to those in those embodiments will be omitted.

[0216] Similar to the embodiments described above, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. The system including the artificial intelligence response output device 10010 and the large-scale language model server may also be referred to as an artificial intelligence response output system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system.

[0217] Using Figure 4A, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) of Embodiment 4 of the present invention will be explained. The database of Embodiment 4 shown in Figure 4A is an extension of the database described in Figure 2I or Figure 3I. Specifically, the database shown in Figure 4A assumes a case where multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores initial setting instructions and conversation history corresponding to each user and character in the database.

[0218] In the example in Figure 4A, for User 1, whose user ID is 1, the initial setup instructions and conversation history for each of the characters—Koto (character ID 1), Tom (character ID 2), and Necco (character ID 3)—are stored. In addition, for User 2, whose user ID is 2, and User 3, whose user ID is 3, the initial setup instructions and conversation history for each of the characters—Koto (character ID 1), Tom (character ID 2), and Necco (character ID 3)—are also stored.

[0219] These initial setup instructions and conversation history data are stored as separate data in different areas for each user-character combination. In Figure 4A, for illustrative purposes, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses the initial setup instructions and conversation history stored in different areas for each user-character combination, based on the user currently using (logged into) the character conversation device (artificial intelligence response output device 10010) or its system, thereby enabling it to more effectively maintain the consistency of the character's personality and the continuity of its memory for each different user.

[0220] Specifically, consider a scenario where User 1 has already conversed with character Tom using a character conversation device (artificial intelligence response output device 10010), and User 2 is unaware of that conversation, and then User 2 subsequently converses with character Tom. In this case, if the artificial intelligence response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify the user, the response output from the artificial intelligence response output device 10010 may be based on a conversation history that User 2 does not remember, potentially leading to inconsistencies in the conversation between User 2 and the character of the artificial intelligence response output device 10010.

[0221] In contrast, even in a similar situation, using the database shown in Figure 4A, the control unit 1110 of the character conversation device (artificial intelligence response output device 10010) identifies the user by ID, stores the initial setting instructions and conversation history in a different area for each user, and uses the initial setting instructions and conversation history stored in the different areas for each user to generate the artificial intelligence response. As a result, the initial setting instructions and conversation history used to generate the artificial intelligence response for each user are based on the user's operations or conversation history and are managed separately from the operations or conversation history of other users. This makes it possible to better maintain consistency in the conversation history between each user and each character of the artificial intelligence response output device 10010.

[0222] Furthermore, the database of initial setting instructions and / or conversation history, as explained in Figure 4A, may be stored in the storage unit 1170 of the artificial intelligence response output device 10010 and used by the control unit 1110. However, it is not limited to this, and the database of initial setting instructions and / or conversation history may also be stored on a server on the network. For example, if the artificial intelligence response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 in generating artificial intelligence responses, the database of initial setting instructions and / or conversation history, as explained in Figure 4A, may be stored on these servers themselves. In this way, the process of re-inserting the initial setting instructions and conversation history into the instructions and sending them from the artificial intelligence response output device 10010 to these servers can be omitted, and the number of transmission tokens for using the large-scale language model can be reduced.

[0223] When storing a database of initial setting instructions and / or conversation history, as explained in Figure 4A, the AI ​​response output device 10010 should send the user ID, character ID, and user instructions for subsequent conversations to these servers. The large-scale language models on these servers use the user ID and character ID obtained from the AI ​​response output device 10010 to retrieve the corresponding initial setting instructions and conversation history from the database of initial setting instructions and / or conversation history shown in Figure 4A. The large-scale language models on these servers should then perform inference using the initial setting instructions and conversation history, along with the user instructions for subsequent conversations sent from the AI ​​response output device 10010, generate an AI response, and send it to the AI ​​response output device 10010. In this way, the effect of more favorably maintaining the consistency of character personality and memory continuity for each different user can be obtained while saving the number of tokens sent for the use of the large-scale language models.

[0224] Next, using Figure 4B, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) of Embodiment 4 of the present invention will be described. The database in Embodiment 4 shown in Figure 4B is an extension of the database described in Figure 1C or Figure 2L. Specifically, the database shown in Figure 4B assumes a case where multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data of standard response phrases corresponding to each user and character in the database.

[0225] In the example in Figure 4B, for user 1 (user ID 1), the system stores the standard response data for each of the following characters: character Koto (character ID 1), character Tom (character ID 2), and character Necco (character ID 3). In addition, for user 2 (user ID 2) and user 3 (user ID 3), the system also stores the standard response data for each of the following characters: character Koto (character ID 1), character Tom (character ID 2), and character Necco (character ID 3).

[0226] These standard response phrases are stored as separate data in different areas for each user-character combination. In Figure 4B, for illustrative purposes, the data stored in each area is denoted as standard response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, standard response phrase data 101 is stored as a database, such as a table, corresponding to the standard response phrases corresponding to condition numbers 1-7 for character 1: Koto, as shown in Figure 2L. Data 201 in Figure 4B is stored as a database, such as a table, corresponding to the standard response phrases corresponding to condition numbers 1-7 for character 2: Tom, as shown in Figure 2L.

[0227] Data 301 in Figure 4B is stored as a database, such as a table, corresponding to the standard response phrases corresponding to condition numbers 1 to 7 for character 3: Necco shown in Figure 2L. Data 102, 202, and 302 in Figure 4B store standard response phrases modified for user 2 in a similar format. Data 103, 203, and 303 in Figure 4B store standard response phrases modified for user 3 in a similar format. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses standard response phrase data stored in different areas for each combination of user and character, based on the user currently using (logged into) the character conversation device (artificial intelligence response output device 10010) or its system.

[0228] In this way, even with the same character, it becomes possible to provide responses using different predefined response templates for each user. In other words, even with the same character, it may be preferable to change the content of the predefined response template depending on the relationship between the character and the user. For example, depending on the relationship between the character's set age and the user's age registered in the artificial intelligence response output device 10010 or the system, the user may be older, the same age, or younger than the character. In this case, changing the content of the character's predefined response template for older users, users of the same age, and younger users will result in a more appropriate or natural conversation between the user and the character. In other words, by performing the operation using the database shown in Figure 4B, it is possible to create a more appropriate or natural conversation by varying the content of the predefined response template according to the relationship between the character and the user.

[0229] As described above, the response template database (response template DB) shown in Figure 4B is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 can use it. However, the response template database (response template DB) shown in Figure 4B may also be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 should generate responses using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 should send the response generated using the response template database (response template DB) to the artificial intelligence response output device 10010 instead of the response generated by the large-scale language model stored in their respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a response template database (response template DB), it becomes possible to generate responses using a response template database (response template DB).

[0230] As described above, the character conversation device and character conversation system according to Embodiment 4 make it possible to create more suitable or natural conversations depending on the relationship between the character and the user, as well as the conversation history.

[0231] <Example 5> Next, Embodiment 5 of the present invention is an improvement on the artificial intelligence response output device 10010 or artificial intelligence response output system described in Figures 1, 2, and 3 of Embodiment 1. Specifically, this is an example of switching the response generation process of the artificial intelligence response output device 10010 from response generation processing using a large-scale language model on the network to response generation processing using a local large-scale language model (such as the local LLM processing unit 10028) provided by the artificial intelligence response output device 10010, or response generation processing using a response template database. In this embodiment, the differences from these embodiments will be explained, and repeated explanations of configurations similar to those embodiments will be omitted.

[0232] Similar to the embodiments described above, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. The system including the artificial intelligence response output device 10010 and the large-scale language model server may also be referred to as an artificial intelligence response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system.

[0233] Using Figure 5A, an example of the switching process for response generation in the artificial intelligence response output device 10010 of Embodiment 5 of the present invention will be explained. The table in Figure 5A shows examples of the switching process for response generation in the artificial intelligence response output device 10010, from Example 1 to Example 9. In the table in Figure 5A, the "Switching Overview" column shows an overview of the switching process for each example. The "State Before Switching of LLM (API-Connected LLM) on the Network" column shows the state before the response generation process by the large-scale language models on the network (large-scale language models connected using APIs), such as the large-scale language model provided by the large-scale language model server 19001 in Figure 1 and the multimodal large-scale language model provided by the large-scale language model server 20001, is switched to another response generation process. The "Switching Occurrence Conditions" column shows the conditions under which the switching process for response generation occurs. The column "Switching Destination from Network LLM (API-connected LLM)" indicates the switching destination for the response generation process of the artificial intelligence response output device 10010, switching from large-scale language models on the network (large-scale language models connected using APIs), such as the large-scale language model provided by the large-scale language model server 19001 and the multimodal large-scale language model provided by the large-scale language model server 20001. The control unit 1110 of the artificial intelligence response output device 10010 should control the system to switch to the large-scale language model, database, or corresponding indicated in "Switching Destination from Network LLM (API-connected LLM)" when the conditions indicated in "Switching Occurrence Conditions" occur in the state shown in "State Before Switching from Network LLM (API-connected LLM)" in Figure 5A.

[0234] The following describes each example shown in the table in Figure 5A. Example 1 is an example of switching depending on the network connectivity status of the artificial intelligence response output device 10010, as shown in the "Switching Overview". In Example 1, the "State before switching of the LLM (API-connected LLM) on the network" indicates that the network connectivity status of the artificial intelligence response output device 10010 is in a connectable state. Here, in Example 1, the "Condition for switching" is indicated as "When network connectivity becomes impossible". That is, this means when network connectivity between the artificial intelligence response output device 10010 and the large-scale language model on the network (large-scale language model connected using an API) becomes impossible. Specifically, this connectivity failure may be due to a communication failure on the connection path from the artificial intelligence response output device 10010 to the Internet 19000. Alternatively, this connectivity failure may be due to a communication failure on the Internet 19000. Alternatively, this connectivity failure may be due to the large-scale language model on the network (large-scale language model connected using an API) itself being unable to connect to the Internet 19000. Furthermore, in Example 1, "local LLM" is indicated as the "switching destination from the network-based LLM (API-connected LLM)." Specifically, this means switching to the response generation process performed by the local LLM processing unit 10028 of the artificial intelligence response output device 10010. In other words, in Example 1, even if for some reason connection to the large-scale language model on the network (a large-scale language model connected using an API) becomes impossible and the response generation process by the large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable, the response generation process is switched to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This makes it possible to continue the response generation process using the large-scale language model, although there may be performance differences between the large-scale language models.

[0235] Next, let's explain Example 2 in Figure 5A. In Example 2, the "switching destination from the LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "response template DB (database)". The response generation process using this "response template DB (database)" is the same as the process explained in Figure 1C, Figure 2L, or Figure 4B, so a repeated explanation will be omitted. In other words, in Example 2, if for some reason it becomes impossible to connect to the large-scale language model on the network (large-scale language model connected using an API), and the response generation process using the large-scale language model on the network (large-scale language model connected using an API) is unavailable, it is possible to switch to the response generation process using the response template database, thereby generating a response with a simpler process and outputting that response to the user.

[0236] Next, let's explain Example 3 in Figure 5A. Example 3 is a variation of Example 1 in which the "switching destination from the network LLM (API-connected LLM)" is changed from "local LLM" to "non-response handling". This "non-response handling" means that even if a user input requests a response from the large-scale language model via the touch panel, microphone 1139, or operation input unit 1107, no response is generated for this input, or even if a user input requests a response from the large-scale language model, no response is output for it. In other words, Example 3 makes it possible to simplify the handling of cases where, for some reason, connection to the network large-scale language model (large-scale language model connected using an API) becomes impossible, and the response generation process by the network large-scale language model (large-scale language model connected using an API) is unavailable.

[0237] Next, we will explain Example 4 in Figure 5A. As shown in the "Switching Overview," Example 4 is an example of switching due to a response delay of the LLM on the network. In Example 4, the "State before switching of the LLM on the network (API-connected LLM)" is shown as a state in which a response from the LLM on the network is obtained within a predetermined time. Here, in Example 4, the "Condition for switching" is shown as when a response from the LLM on the network is not obtained within a predetermined time and exceeds the predetermined time. Also, in Example 4, the "Destination for switching from the LLM on the network (API-connected LLM)" is shown as the "Local LLM." The destination "Local LLM" is the same as in Example 1, so we will omit the repeated explanation. In other words, in Example 4, even if for some reason the response from the LLM on the network (Large-Scale Language Model connected using an API) exceeds a predetermined time and the response generation process by the LLM on the network (Large-Scale Language Model connected using an API) cannot be used smoothly, the response generation process by the local LLM processing unit 10028 of the artificial intelligence response output device 10010 is switched to. This means that, despite performance differences among large-scale language models, it is possible to continue using large-scale language models for response generation.

[0238] Next, let's explain Example 5 in Figure 5A. Example 2 is the same as Example 4, but with the "switching destination from the LLM on the network (API-connected LLM)" changed from "local LLM" to "response template DB (database)". The response generation process using this "response template DB (database)" is the same as the process explained in Figure 1C, Figure 2L, or Figure 4B, so a repeated explanation will be omitted. In other words, in Example 5, if for some reason the response from the LLM on the network (large-scale language model connected using an API) exceeds a predetermined time, and the response generation process using the LLM on the network (large-scale language model connected using an API) cannot be used smoothly, it is possible to switch to a response generation process using the response template database, thereby generating a response with a simpler process and outputting that response to the user.

[0239] Next, we will explain Examples 6 to 9 in Figure 5A. As shown in the "Switching Overview," Examples 6 to 9 are examples of switching due to reaching the upper limit of API usage or usage fees. Here, as explained in Example 2, providers of large-scale language models often recover the costs used to train the large-scale language model from the user of the device as API usage fees. In this case, natural language models often charge API usage fees based on the number of tokens, which are units of words that divide sentences. Here, various billing and limiting methods can be considered for API usage fees. One example is to define the upper limit of the amount of large-scale language model usage service that a user can receive under normal circumstances using the number of tokens processed.

[0240] In this case, users can access services utilizing large-scale language models at a predetermined API usage fee until they reach their usage limit (or corresponding usage fee). Once they reach their usage limit (or corresponding usage fee), certain restrictions may be imposed, such as being unable to access services utilizing large-scale language models at their normal performance or frequency.

[0241] Examples 6 to 9 in Figure 5A illustrate the switching control of the response generation process by the control unit 1110 of the artificial intelligence response output device 10010 when such limitations occur in the large-scale language model utilization service. Specifically, in Example 6, the "state before switching of LLM on the network (API-connected LLM)" is a state in which the API usage amount and API usage fees have not reached a predetermined upper limit. This means that the usage amount of LLM on the network (API-connected LLM) has not reached a predetermined upper limit. At this time, the user can use LLM on the network (API-connected LLM) in the normal state.

[0242] In Example 6, the "condition for switching" is shown as when the API usage or API usage fee reaches a predetermined limit. This means when the usage of the LLM on the network (API-connected LLM) reaches a predetermined limit. Also, in Example 6, the "destination for switching from the LLM on the network (API-connected LLM)" is shown as a second LLM on a different network from the LLM that was being used under normal circumstances (which may be called the first LLM). An example of a second LLM on the network is an LLM that is cheaper than the first LLM that was being used under normal circumstances. Since it is a lower-priced service, the performance of the second LLM is likely to be lower than that of the first LLM. Even in this case, there is still a significant advantage if a large-scale language model can be used cheaply even after the usage / usage fee limit of the first LLM has been reached.

[0243] Next, let's explain Example 7 in Figure 5A. In Example 7, the "switching destination from the network-based LLM (API-connected LLM)" in Example 6 is changed from a second network-based LLM (which may be called the first LLM) that was normally used to a "local LLM". In Example 7, even if the API usage or API usage fees reach a predetermined limit, that is, even if the usage of the network-based LLM (API-connected LLM) reaches a predetermined limit, it is possible to continue using the large-scale language model for response generation by switching to a response generation process using a local LLM, which is not subject to restrictions based on the network-based LLM usage, API usage, or API usage fees.

[0244] Next, we will explain Example 8 in Figure 5A. In Example 8, the "switching destination from the LLM on the network (API-connected LLM)" in Example 7 is changed from "local LLM" to "response template DB (database)". The response generation process using this "response template DB (database)" is the same as the process explained in Figure 1C, Figure 2L, or Figure 4B, so repeated explanations will be omitted. In Example 8, even if the API usage or API usage fees reach a predetermined limit, that is, even if the usage of the LLM on the network (API-connected LLM) reaches a predetermined limit, the system switches to a response generation process using a response template database that is not subject to restrictions based on the usage of the LLM on the network, API usage, or API usage fees. This makes it possible to generate a response with a simpler process and output that response to the user.

[0245] Next, let's explain Example 9 in Figure 5A. Example 9 is a variation of Example 7 in which the "switching destination from the LLM on the network (API-connected LLM)" is changed from "local LLM" to "non-response handling." This "non-response handling" means not generating a response to the user or not outputting a response to the user. In Example 9, it becomes possible to simplify the handling of situations where the response generation process by the large-scale language model on the network (large-scale language model connected using APIs) is unavailable because the API usage or API usage fees have reached a predetermined limit, that is, because the usage of the LLM on the network (API-connected LLM) has reached a predetermined limit.

[0246] As described above, the switching control of the response generation process of the artificial intelligence response output device 10010, as shown in Examples 1 to 9 of Figure 5A, allows for more appropriate switching or response depending on the situation, even when the response generation process by the LLM (Large-Scale Language Model connected using APIs) on the network is not available as usual.

[0247] Note that the switching control for Examples 1 to 9 in FIG. 5A may be performed by combining a plurality of examples. For example, the switching control for Examples 1 to 3 may be combined with any one of the controls for Examples 4 to 9 respectively. Similarly, the control for Example 4 or Example 5 may be combined with any one of the controls for Examples 1 to 3 or Examples 6 to 9 respectively. Similarly, the control for Examples 6 to 9 may be combined with any one of the controls for Examples 1 to 5 respectively.

[0248] Next, an example of a display example of an AI assistant or a character will be described with respect to the case where the artificial intelligence response output device 10010 of Example 5 is configured as an AI assistant device or a character conversation device, using FIGS. 5B to 5D.

[0249] First, FIG. 5B is a display example of an AI assistant or a character in the artificial intelligence response output device 10010 when performing the switching control of Example 3 in FIG. 5A. In the example of FIG. 5B, the display state of the AI assistant or the character is changed according to whether the network connection state of the artificial intelligence response output device 10010 is network connectable or network unavailable. Since the states where the artificial intelligence response output device 10010 is network connectable and network unavailable are as described in FIG. 5A, repeated explanations will not be given.

[0250] In the example of FIG. 5B, when the artificial intelligence response output device 10010 is network - connectable, it displays the AI assistant or character in the normal state where it is awake. However, when it is not network - connectable, it displays the AI assistant or character in the "sleeping" state. In the switching control of Example 3 in FIG. 5A, when the artificial intelligence response output device 10010 is not network - connectable, it does not generate a response or output a response even if there is an input of an instruction sentence from the user. At this time, the user may feel a sense of discomfort when the AI assistant or character displayed by the artificial intelligence response output device 10010 is in the normal awake state. However, if the AI assistant or character displayed by the artificial intelligence response output device 10010 is displayed in the sleeping state, the user can understand that "the reason the AI assistant or character does not respond is that it is sleeping", and it becomes possible to further reduce the sense of discomfort felt by the user.

[0251] In addition, in the case of FIG. 5B(2), before the user input requests a response from the large - language model via the touch panel, microphone 1139 or operation input unit 1107 of the artificial intelligence response output device 10010, it is desirable for the user to understand that "the reason the AI assistant or character does not respond is that it is sleeping". Therefore, in the case of (2) in FIG. 5B where it is not network - connectable, the start timing of the state of displaying the AI assistant or character in the "sleeping" state is preferably immediately after the control unit 1110 of the artificial intelligence response output device 10010 determines that it is not network - connectable, which is before the user input that requests a response from the large - language model.

[0252] Next, as another display example, the display example in Figure 5C will be explained. The display example in Figure 5C is an example in which the display state of the AI ​​assistant or character is changed according to the state of "Switching destination from LLM on the network (API-connected LLM)" in the table in the switching control in Figure 5A. Specifically, Figure 5C shows: (1) an example of the display of the AI ​​assistant or character when the artificial intelligence response output device 10010 is able to connect to a large-scale language model on the network (a large-scale language model connected using an API) and is in a state where response generation processing by the large-scale language model on the network is available (referred to as the normal state in this figure); (2) an example of the display of the AI ​​assistant or character when the artificial intelligence response output device 10010 has switched to response generation processing using an LLM or response template database that is less powerful than the large-scale language model on the network (a large-scale language model connected using an API); and (3) an example of the display of the AI ​​assistant or character when the artificial intelligence response output device 10010 has switched to the non-response response described in Figure 5A.

[0253] In the example in Figure 5C, for example, if (1) the artificial intelligence response output device 10010 is in the "normal state," the artificial intelligence response output device 10010 will display the AI ​​assistant or character in a state where there are no particular problems. Note that "normal state" in Figure 5C can be considered as any state other than states (2) and (3). Also, for example, if (2) the artificial intelligence response output device 10010 switches to response generation processing using an LLM or response template database which is less powerful than the large-scale language model on the network (a large-scale language model connected using an API), the artificial intelligence response output device 10010 will display the AI ​​assistant or character in a "sleepy" state. Note that "displaying the AI ​​assistant or character in a 'sleepy' state" can also be expressed as "displaying a state in which the AI ​​assistant or character is feeling sleepy."

[0254] The response generation process in (2) is less efficient than the response generation process using a large-scale language model on the network (a large-scale language model connected using an API) in the normal state in (1). Therefore, by displaying the AI ​​assistant or character in a "sleepy" state, it is possible to implicitly inform the user that the AI ​​assistant or character has low response performance. This makes it possible to further reduce the discomfort the user feels with low-performance responses. The switching conditions for the artificial intelligence response output device 10010 to switch to a response generation process using an LLM or a response template database, which is less efficient than the large-scale language model on the network (a large-scale language model connected using an API), are as explained in Figure 5A, so a repeated explanation will be omitted.

[0255] Furthermore, in the case of Figure 5C(2), it is desirable to implicitly inform the user that the AI ​​assistant or character has low response performance before the user input requests a response from the large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 10010. Therefore, it is desirable that the start timing of the state in Figure 5C(2) where the AI ​​assistant or character is displayed in a "sleepy" state is immediately after the point in time when the artificial intelligence response output device 10010 switches to response generation processing using an LLM or response template database, which has lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user input requests a response from the large-scale language model.

[0256] Furthermore, for example, when the artificial intelligence response output device 10010 switches to the non-response mode described in Figure 5A, the artificial intelligence response output device 10010 displays the AI ​​assistant or character in a "sleeping" state. As explained in Figure 5B, by displaying the AI ​​assistant or character displayed by the artificial intelligence response output device 10010 in a "sleeping" state, the user can understand that "the AI ​​assistant or character is not responding because it is sleeping," thereby further reducing the sense of unease the user may feel. The conditions for the artificial intelligence response output device 10010 to switch to the non-response mode described in Figure 5A are as explained in Example 3 or Example 9 in Figure 5A, so a repeated explanation will be omitted. In the case of Figure 5C(3), it is desirable that the user understands that "the AI ​​assistant or character is not responding because it is sleeping" before the user makes user input requesting a response from the large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 10010. Therefore, the timing of the start of the state in Figure 5C (3) where the AI ​​assistant or character is displayed in a "sleeping" state is preferably immediately after the point in time when the artificial intelligence response output device 10010 switches to the non-response mode described in Figure 5A, prior to the user input requesting a response from the large-scale language model.

[0257] In the display example shown in Figure 5C, the artificial intelligence response output device 10010 implicitly reflects the change in the state of the AI ​​assistant or character as a change in state, without directly providing the user with a technical explanation of the state of the artificial intelligence response output device 10010 regarding the response generation process. This reduces the sense of unease the user may feel compared to directly providing the user with a technical explanation of the state of the artificial intelligence response output device 10010 regarding the response generation process. Furthermore, it reduces the sense of unease the user may feel compared to keeping the display state of the AI ​​assistant or character the same as the normal state, even though the state of the artificial intelligence response output device 10010 regarding the response generation process has changed.

[0258] However, some users may want a more accurate explanation of the technical status in each state. Therefore, an example of a display to accommodate such users will be explained using Figure 5D. The rows for the device status and the display status explanation in the table shown in Figure 5D are exactly the same as those in Figure 5C, so repeated explanations will be omitted. Also, the display example of the AI ​​assistant or character shown in the row for the AI ​​assistant or character display example is almost identical to that in Figure 5C, but differs in that a question mark (?) is displayed in the display example. This question mark (?) is a mark that the user operates when requesting an explanation from the artificial intelligence response output device 10010, and may be called a help mark.

[0259] In the example in Figure 5D, when a user selects the question mark (?) through user operation via the touch panel of the operation input unit 1107 or display unit 10011 in Figure 1B, the display of the AI ​​assistant or character of the artificial intelligence response output device 10010 changes to the display example shown in the row of the display example after user operation. Specifically, regardless of whether the device state is (1), (2), or (3), a technical explanation of the state for each state is displayed. For example, in the example in Figure 5D, if the device state is (1) normal state, a display such as "Normal state" should be shown, explaining that it is a normal state with no particular technical limitations. Also, if the device state is (2) using a low-performance LLM or response template database, a display such as "Low-performance mode" should be shown, technically explaining that it is a low-performance state. This display can also be considered an explanation of the factors that cause the AI ​​assistant or character to display in a "sleepy" state.

[0260] In this case, a more technically detailed explanation may be provided. Specifically, a message such as "Low-performance LLM usage mode" or "Standard response mode" may be displayed. Also, if the device status is (3) unresponsive, a message such as "Network connection unavailable" may be displayed to technically explain the reason for switching to unresponsive mode. If the reason for switching to unresponsive mode is that the response from the LLM (Large-Scale Language Model connected using an API) on the network has exceeded a predetermined time, a message such as "Response from LLM is delayed" may be displayed. Also, if the reason for switching to unresponsive mode is that the usage of LLM on the network, API usage, or API usage fees have reached their limits, a message such as "LLM usage has reached its limit," "API usage has reached its limit," or "API usage fees have reached a predetermined amount" may be displayed. These messages can be considered as explanations of the reasons why the AI ​​assistant or character is displayed in a "sleeping" state.

[0261] As shown in the example display in Figure 5D described above, even if there are technical limitations in the response generation process of the artificial intelligence response output device 10010, by implicitly indicating the status of the device through changes in the display state of the AI ​​assistant or character, rather than directly explaining it to the user, the sense of discomfort the user may feel can be further reduced. This display is more suitable for users who do not need a technical explanation. Furthermore, by displaying an operation mark to explain the technical status, the status of the response generation process in the artificial intelligence response output device 10010 (normal state or state with technical limitations) is displayed to the user who operates the mark. This makes it possible to provide a more suitable display for users who want to know the technical status accurately.

[0262] In the examples of Figures 5B, 5C, and 5D, the "sleeping" state is shown as an example of the display state of the AI ​​assistant or character when it is "unresponsive," but this is just one example, and the embodiments of this example are not limited to this. Instead of the "sleeping" state, other display states that implicitly indicate an unresponsive situation, such as "on break," may be used. Also, in the examples of Figures 5C and 5D, the "sleepy" state is shown as an example of the display state of the AI ​​assistant or character when a low-performance LLM or response template database is being used, but this is just one example, and the embodiments of this example are not limited to this. Other display states that implicitly indicate the low responsiveness of the AI ​​assistant or character, such as "hungry," may be used.

[0263] As described above, the artificial intelligence response output device and artificial intelligence response output system according to Embodiment 5 make it possible to more effectively switch the response generation process used by the artificial intelligence response output device depending on the connection status between the large-scale language model on the network and the artificial intelligence response output device, the response delay status from the large-scale language model on the network, or the amount of utilization of the large-scale language model on the network. Furthermore, when the artificial intelligence response output device according to Embodiment 5 is configured as an AI assistant device or a character conversation device, it becomes possible to display information that is less unnatural to the user.

[0264] <Example 6> Next, Embodiment 6 of the present invention is an improvement on the artificial intelligence response output device 10010 or artificial intelligence response output system described in Figures 1 to 5 of Embodiments. Specifically, this embodiment is an example in which the response generation process of the artificial intelligence response output device 10010 is more preferably combined with a response generation process using a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided by the artificial intelligence response output device 10010, and a response generation process using a response template database to generate a response output. In this embodiment, the differences from these embodiments will be explained, and repeated explanations of configurations similar to these embodiments will be omitted.

[0265] Similar to the embodiments described above, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. The system including the artificial intelligence response output device 10010 and the large-scale language model server may also be referred to as an artificial intelligence response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system.

[0266] An example of the response generation process in the artificial intelligence response output device 10010 of Embodiment 6 of the present invention will be explained using Figure 6. Figure 6 shows an example of the response generation process in the artificial intelligence response output device 10010 of Embodiment 6 of the present invention. Specifically, it shows a time axis progressing from top to bottom, a processing flow, and an example of a response output. The response output shown in the example of a response output can be performed via display by the display unit 10011 of the artificial intelligence response output device 10010 or via sound output by the sound output unit 1140.

[0267] In the example in Figure 6, first, at time t0, the artificial intelligence response output device 10010 receives user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107, and the control unit 1110 of the artificial intelligence response output device 10010 acquires this user input (step 600). Next, at time t1, the control unit 1110 begins preparation for response output using the response template database stored in the storage unit 1170, and begins response output using the response template database (step 601). In the example in Figure 6, at time t2, response output using the response template database has started, and as shown in the figure, the template response is being output but is not yet complete. The "Good morning" in the figure shows the output up to the middle of the sentence that follows, "Good morning...".

[0268] At time t3, before the response output using the response template database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and sends the generated instruction sentence to the large-scale language model on the network or to the local large-scale language model (such as the local LLM processing unit 10028) of the artificial intelligence response output device 10010, initiating a request for a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the response template database is completed, the control unit 1110 begins acquiring a response from the large-scale language model (step 603).

[0269] At time t5, an example of a response output completed using the response template database is shown. For example, Figure 6 shows an example where, at time t5, the display of the response "Good morning. Today is [Month] [Day]." has been completed using the template text stored in the response template database and date information stored in memory. Here, at time t4, before the completion of the response output using the response template database, the control unit 1110 has already started acquiring responses from the large-scale language model. Therefore, at time t6, following time t5 when the display of the response output using the response template database is completed, the control unit 1110 starts outputting responses from the large-scale language model following the response output using the response template database (step 604). Subsequently, at time t7, the response from the large-scale language model is output following the response output using the response template database. Once the response output from the large-scale language model is completed, the response output according to the processing flow shown in Figure 6 is completed (step 605).

[0270] Next, the effects of the processing flow shown in FIG. 6 of the present invention will be described. Processing of large language models requires a great deal of computing resources. Generally, even if inference, which requires fewer computing resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several seconds to more than ten seconds from when the control unit starts a response request to the large language model until a response can be obtained from the large language model. This period corresponds to the period from time t3 to time t4 shown in FIG. 6. Also, from time t0 when there was user input until time t4, since the control unit 1110 has not obtained a response output from the large language model, it cannot output a response from the large language model to the user.

[0271] Therefore, in a processing flow where there is no start of preparation for response output using the response template sentence database in step 601 shown in FIG. 6 and no start of response output using the response template sentence database, the user may continue to wait without a response from the artificial intelligence response output device 10010 for a period exceeding several seconds to more than ten seconds from time t0 when the user input was made until time t4. For example, when the artificial intelligence response output device 10010 is configured as an AI assistant device or a character conversation device, etc., this waiting time may give the user a sense of discomfort.

[0272] In contrast, in the processing flow according to Embodiment 6 of the present invention shown in FIG. 6, the control unit 1110 starts processing of response output using a response template sentence database that requires fewer computing resources than processing of the large language model before starting to obtain a response from the large language model. As a result, the user does not have to continue to wait without a response from the artificial intelligence response output device 10010 from time t0 until time t4. For the user, whether it is a response output using the response template sentence database or a response output from the large language model, it is still a response from the artificial intelligence response output device 10010.

[0273] Therefore, in the processing flow shown in Figure 6, by adding step 601 before step 603, the response of the artificial intelligence response output device 10010 to the user can be made to appear faster. This makes it possible to further reduce the user's discomfort caused by the length of waiting time. Furthermore, by outputting the response from the large-scale language model in step 604, following the response using the response template database, the user can perceive these outputs as a series of more natural outputs.

[0274] As described above, the artificial intelligence response output device and artificial intelligence response output system according to Embodiment 6 can reduce the user's waiting time for a response from the artificial intelligence response output device, thereby further reducing the discomfort the user feels.

[0275] <Example 7> Next, Embodiment 7 of the present invention is an improvement on the artificial intelligence response output device 10010 or artificial intelligence response output system described in Figures 1 to 6 of Embodiments. Specifically, in the response generation process of the artificial intelligence response output device 10010, this embodiment is an example in which a response output is generated by more preferably combining a response generation process using a large-scale language model on the network or a response generation process using a local large-scale language model (such as the local LLM processing unit 10028) that the artificial intelligence response output device 10010 has internally, and a response generation process using a response template database. In this embodiment, the differences from these embodiments will be explained, and repeated explanations of configurations similar to those embodiments will be omitted.

[0276] Similar to the embodiments described above, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. The system including the artificial intelligence response output device 10010 and the large-scale language model server may also be referred to as an artificial intelligence response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system. The same applies to the artificial intelligence response output device in each of the subsequent embodiments.

[0277] First, using Figure 7, an example of the configuration of the artificial intelligence response output device and artificial intelligence response output system of Embodiment 7 of the present invention will be described. The artificial intelligence response output system in Figure 7 is an improvement on the character conversation system shown in Figure 3C. In Figure 7, the same components as those shown in Figure 3C are denoted by the same reference numerals, and redundant explanations are omitted.

[0278] In the example shown in Figure 7, as described in Explanation 7010, the artificial intelligence response output device 10010 includes, for example, a client application as software that is deployed in the memory 1109 shown in Figure 1B and executed by the control unit 1110.

[0279] The client application can control the input of instruction statements (including setting instructions and user instructions) as described in Figures 1 to 6 of Examples. The client application can accept user input (hereinafter also referred to as user input). The client application can also, for example, send information to a large-scale language model using instruction statements based on user input. The client application can also, for example, send control information to the large-scale language model that includes information about emotion types used to classify user emotions obtained from user input.

[0280] As illustrated in Explanation 7010, the client application receives user input and, based on that input, detects the user's emotions (hereinafter referred to as "user emotions"), quantifies the user emotions, and manages a counter for the user emotions. As will be explained in more detail later, in this example, the client application obtains the evaluation value (also called the user emotion evaluation value) of at least one of the emotion words and emotion elements in the user input, and the user emotion counter value. In other words, the client application obtains parameters related to user emotions as information about the emotion type used to classify the user emotions. The client application then sends control information, including this information about emotion types, to a large-scale language model, either along with or included in the instruction statement.

[0281] In the example in Figure 7, as explained in 7020, an LLM (Large-Scale Language Model) application exists on the large-scale language model side (server side) that controls the input and output of information to the large-scale language model. As in the example in Figure 7, when the LLM application (abbreviation for large-scale language model application; the same applies hereinafter) is on the large-scale language model server 20001 side, the LLM application is loaded into the memory of the large-scale language model server 20001 and executed by the control unit of the large-scale language model server 20001. When the LLM application receives an instruction sentence containing control information sent by the client application, it generates a response (response sentence) according to the received instruction sentence. As shown in an example in 7020, the LLM application determines the response mode of the large-scale language model (also called the AI ​​response mode) according to the emotion type (user emotion) based on the control information, and generates a response (also called the LLM response) to the instruction sentence based on the determined AI response mode. The LLM application also sends the generated LLM response to the artificial intelligence response output device 10010.

[0282] When the client application receives an LLM response generated by the LLM application, it controls the components shown in Figure 1B, including the display unit 10011 and the audio output unit 1140, to output the received LLM response to the user. In this example, the client application outputs the received LLM response to the user according to output conditions corresponding to the AI ​​response mode, that is, output conditions corresponding to the emotion type (user emotion).

[0283] The client application can generate the instructions (prompts) for the large-scale language model in response to user input. Examples of user input interfaces that accept user input include a mouse, keyboard, and touch panel, as illustrated in the operation input unit 1107 in Figure 1B, similar to the explanation in Figure 3C. A microphone 1139 that captures the user's voice can also be considered a user input interface. Furthermore, a communication unit 1132 that communicates with the mobile information processing terminal 20010 used by the user, such as a smartphone or tablet, can also be considered a user input interface. Output interfaces that the client application uses to output LLM responses to the user include, for example, a display unit 10011 that outputs in media such as text, images, or video, and an audio output unit 1140 that outputs the response from the large-scale language model as audio.

[0284] Incidentally, the client application can control the insertion of predefined phrases into the LLM responses output by the artificial intelligence response output device 10010. This includes the output control of predefined phrases as described in Figures 1C, 2L, 4B, 5A, or 6. Examples of such predefined phrases include the insertion of predefined greetings and predefined acknowledgments. These predefined phrases can be stored, for example, in the storage unit 1170 in Figure 1B of the artificial intelligence response output device 10010 as a response predefined phrase database, which the client application can then use.

[0285] Furthermore, the LLM application can accept preset instructions to adjust the output (response) of the large-scale language model to specific specifications. While user-defined instructions change with each user input, the preset instructions for adjusting the output of the large-scale language model to specific specifications are given on a regular basis and do not change with individual user inputs. Therefore, these preset instructions may also be called regular instructions or initial setting instructions. They may also be called instructions for customizing the output of the large-scale language model. For example, a preset instruction could be used to insert a predetermined interjection into the natural language response generated by the large-scale language model. The LLM application's preset instructions may be set by sending instructions from the client application of the artificial intelligence response output device 10010 to the LLM application. Alternatively, the LLM application's preset instructions may be set by communicating with the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user.

[0286] Furthermore, the example in Figure 7 shows an example where the large-scale language model has an LLM application that controls the input and output of information to the large-scale language model. In contrast, as another modification, the artificial intelligence response output device 10010 may be configured to have an LLM application that controls the input and output of information to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. In this case, the LLM application may be software that is loaded into memory 1109 and executed by the control unit 1110. Also in this case, the LLM application that controls the input and output of information to the local LLM processing unit 10028 and the client application described above may be different applications, but they may also be the same application.

[0287] Here, the artificial intelligence response output device 10010 can send not only instruction text but also control information to the large-scale language model server 20001, as described above. This control information may be sent together with the instruction text. Furthermore, the transmission of the control information may occur before the transmission of the instruction text. The client application may control the transmission of this control information and instruction text, but it is not necessarily required to be done by the client application.

[0288] Next, an example of the details of the control information will be described. This control information stores information about the emotion type acquired by the client application. In this example, the client application acquires the evaluation values ​​of emotion words and emotion elements that appear in the user input, as well as a counter value for the user's emotion. The control information stores information about the emotion type used to classify the user's emotion, such as parameters related to the user's emotion acquired by the client application in this way.

[0289] Furthermore, the control information may also store identification and authentication information for logging into the LLM application. For example, a user may create an account in the LLM application in advance and generate authentication information including identification information such as a user ID and a password. The creation of this account may be performed by communicating with the LLM application via the communication unit 1132 of the artificial intelligence response output device 10010. Alternatively, the creation of this account may be performed by communicating with the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user. The control information containing the authentication information is sent from the communication unit 1132 of the artificial intelligence response output device 10010 to the LLM application, and login is performed through the authentication process. If authentication is successful, the artificial intelligence response output device 10010 and the LLM application establish communication as the user corresponding to the authentication information. As a result, the user can use the information available to their account from the information stored in the memory area of ​​the LLM application.

[0290] Furthermore, if the LLM application is an LLM application on the large-scale language model server 20001, the memory area for the LLM application can be provided in the memory or storage unit of the large-scale language model server 20001. Also, if the LLM application is an LLM application that controls the input and output of information to and from the local LLM processing unit 10028 of the artificial intelligence response output device 10010, the memory area for the LLM application can be provided in the storage unit 1170 or memory 1109 in Figure 1B.

[0291] The information available to the user's account includes the preset instructions mentioned above. Furthermore, as shown in Figures 2H, 2I, 2L, 4A, or 4B, if the artificial intelligence response output device 10010 can switch and display multiple characters, the LLM application may store information for preset instructions corresponding to each of the multiple characters in its memory area. In this case, if the control information transmitted from the artificial intelligence response output device 10010 to the LLM application includes a character ID to identify the character, the LLM application can determine which character's preset instruction the user wants to apply. In other words, this is equivalent to including information on switching preset instructions for each character in the control information, and enabling the LLM application to switch preset instructions according to that switching information.

[0292] Furthermore, if user 230 sets character-specific preset instructions for the LLM application, the user can store the setting information for these preset instructions in the control information sent from the artificial intelligence response output device 10010 to the LLM application and send it. Alternatively, the user can send the setting information for these preset instructions to the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user. Upon receiving the setting information for these preset instructions, the LLM application stores the preset instruction information based on this setting information in its memory. As described above, the preset instruction information may be stored for each character. Alternatively, it may be stored for each user account. Once the preset instruction settings are complete, the user can transmit a character ID that identifies the desired character from the artificial intelligence response output device 10010 to the LLM application, allowing the LLM application to select the preset instruction to apply to that character and apply it to the large-scale language model. In this case, it is not necessary to transmit information equivalent to the preset instruction (setting information) to the LLM application each time an instruction is sent, thus reducing the amount of communication data.

[0293] Another example is that, each time an instruction is sent, the control information transmitted from the artificial intelligence response output device 10010 to the LLM application may contain information equivalent to a preset instruction and transmit it. In this case, although the transmission frequency of information equivalent to a preset instruction will increase, it will be possible to omit preparations such as prior user account registration and prior setting of preset instructions. Alternatively, each time an instruction is sent, the information equivalent to the preset instruction may be stored in the setting instruction area of ​​the instruction rather than the user instruction area and transmitted. In this case as well, it will be possible to omit preparations such as prior user account registration and prior setting of preset instructions.

[0294] Next, using Figures 8A, 8B, and 8C, several examples of the deployment and execution of client applications and LLM applications in an artificial intelligence response output system will be described. The examples in Figures 8A, 8B, and 8C are schematic diagrams that focus on the deployment, execution, and communication of client applications and LLM applications in an artificial intelligence response output system configuration that includes an artificial intelligence response output device 10010 and a large-scale language model server 20001. Other components besides the client applications and LLM applications are omitted in notation and description for the sake of simplicity.

[0295] First, in the example shown in Figure 8A, the client application 8010 is executed in the artificial intelligence response output device 10010. Specifically, the client application 8010 is loaded into the memory 1109 shown in Figure 1B and executed by the control unit 1110. In addition, the server LLM application 8020 is executed in the large-scale language model server 20001. Specifically, the server LLM application 8020 is loaded into the memory provided by the large-scale language model server 20001 and executed by the control unit provided by the large-scale language model server 20001. The server LLM application 8020 may also be an application that controls the entire large-scale language model possessed by the large-scale language model server 20001. The server LLM application 8020 may also be an application that controls the input and output of information to and from the large-scale language model possessed by the large-scale language model server 20001. The client application 8010 and the server LLM application 8020 communicate via, for example, the communication unit 1132 shown in Figure 1B, the communication device 19011 shown in Figure 1A, a network such as the Internet 19000, and the communication unit provided by the server. In the artificial intelligence response output system example in Figure 8A, the client application 8010 and the server LLM application 8020 cooperate to output a more suitable artificial intelligence response to the user. Here, the large-scale language model possessed by the large-scale language model server 20001 may also be called the server large-scale language model. The server LLM application 8020 may also be called the server large-scale language model application.

[0296] Next, in another example, the example in Figure 8B, the client application 8010 is executed in the artificial intelligence response output device 10010. Similar to Figure 8A, the client application 8010 is deployed to the memory 1109 shown in Figure 1B and executed by the control unit 1110. In addition, the local LLM application 8015 is executed in the artificial intelligence response output device 10010. Specifically, the local LLM application 8015 is deployed to the memory 1109 shown in Figure 1B in the artificial intelligence response output device 10010 and executed by the control unit 1110. The local LLM application 8015 may be an application that controls the entire large-scale language model possessed by the local LLM processing unit 10028 shown in Figure 1B. The local LLM application 8015 may also be an application that controls the input and output of information to the large-scale language model possessed by the local LLM processing unit 10028. The client application 8010 and the local LLM application 8015 communicate with each other, for example, via a communication channel such as a bus within the artificial intelligence response output device 10010. In the artificial intelligence response output system shown in the example of Figure 8B, the client application 8010 and the local LLM application 8015 work together to output a more suitable artificial intelligence response to the user. Here, the large-scale language model possessed by the local LLM processing unit 10028 may also be called the local large-scale language model. The local LLM application 8015 may also be called the local large-scale language model application.

[0297] Next, in another example, the example in Figure 8C, the client application 8010 and the local LLM application 8015 are executed in the artificial intelligence response output device 10010. The details of the client application 8010 and the local LLM application 8015 are as described in Figures 8A and 8B, so a repeated explanation will be omitted. Also, the server LLM application 8020 is executed in the large-scale language model server 20001. The details of the server LLM application 8020 are as described in Figure 8A, so a repeated explanation will be omitted. The client application 8010 and the server LLM application 8020 communicate via, for example, the communication unit 1132 shown in Figure 1B, the communication device 19011 shown in Figure 1A, a network such as the Internet 19000, and the communication unit provided by the server. The client application 8010 and the local LLM application 8015 communicate via, for example, a communication path such as a bus within the artificial intelligence response output device 10010. In the artificial intelligence response output system shown in Figure 8C, the client application 8010, the local LLM application 8015, and the server LLM application 8020 work together to output a more suitable artificial intelligence response to the user.

[0298] By the way, in the artificial intelligence response output system according to Example 7, as briefly explained earlier, when conversing with user 230, the large-scale language model takes into account the user's emotions and generates a response to the user instruction sentence (user input sentence). In the artificial intelligence response output system according to Example 7, as an example, the client application 8010 and the LLM application work together to take into account the user's emotions and generate a response to the user instruction sentence (user input sentence). Furthermore, output adjustments are applied to the generated response output according to the user's emotions.

[0299] More specifically, in the artificial intelligence response output system according to Embodiment 7, the artificial intelligence response output device 10010 acquires parameters related to user emotion as information about emotion types for classifying user emotion, and transmits control information including the acquired parameter information to a large-scale language model. The large-scale language model determines a response mode according to the emotion type (user emotion) based on the control information received from the artificial intelligence response output device 10010, and generates a response to the instruction sentence according to the determined response mode. In this example, the control unit 1110 of the artificial intelligence response output device 10010 performs emotion type discrimination, that is, user emotion discrimination, and transmits control information including the result of emotion type discrimination to the large-scale language model. Based on the control information, the large-scale language model determines a response mode according to the emotion type, and generates a response to the instruction sentence according to the determined response mode. Then, the artificial intelligence response output device 10010 outputs the response generated by the large-scale language model to the user according to output conditions corresponding to the response mode. The generation and output of responses that take user emotion into account in the artificial intelligence response output system will be described in detail below.

[0300] Figures 9A and 9B illustrate an example of a conversation between an artificial intelligence response output system and a user. Figure 9A shows an example where the large-scale language model creates a response without taking into account the user's emotions. Figure 9B shows an example where the large-scale language model takes into account the user's emotions. In Figures 9A and 9B, the right side of the figures shows an example of a user instruction sentence (user input sentence) that the user inputs. The details of how to input the user instruction sentence are as explained in Examples 1 to 6, so a repeated explanation will be omitted. The left side of the figures shows an example of the output from the artificial intelligence response output device 10010 of the artificial intelligence response output system, that is, the output of the response (LLM response) generated by the large-scale language model. Both the input of the user instruction sentence and the output of the LLM response are shown in chronological order from top to bottom.

[0301] First, as shown in Figure 9A, when a large-scale language model creates a response without taking into account the user's emotions, it is easy for the conversation between the artificial intelligence response output system and user 230 to become disjointed. In this example, the first user instruction sentence, user input sentence 1, is the greeting "I'm home," and in response, the artificial intelligence response output device 10010 outputs the standard greeting "Welcome back" as LLM response 1. Note that LLM response 1 may be generated by the large-scale language model, but since a standard phrase can be used, it may also be inserted by the artificial intelligence response output device 10010.

[0302] Next, the user input 2 is "Haa..." (a sigh). At this point, the presence of a sigh indicates that the user does not seem to have had a good day. In other words, if the user's emotions were interpreted from the sigh, it would be possible to understand that user 230 does not seem to have had a good day. However, in this example, the artificial intelligence response output system does not interpret the user's emotions (or atmosphere) from the sigh. Therefore, in response to the sigh (user input 2), the LLM response 2 outputs "Did anything good happen today?", which is not appropriate for the user's emotions. This LLM response 2 is a standard phrase that is output following LLM response 1, "Welcome back," and is output in the context of conversation with the user to initiate conversation or fill gaps.

[0303] Subsequently, if the user input sentence 3 is "Nothing good has happened," in response to the output of LLM response 2, the artificial intelligence response output system can output an LLM response 3 that is appropriate for user input sentence 3, such as "What happened? I'm here to listen if you'd like to talk," which is a response sentence generated by a large-scale language model.

[0304] Furthermore, if the user input sentence 4 is "I think I've lost my favorite pendant," in response to LLM response 3, the artificial intelligence response output system can output LLM response 5, which is appropriate for user input sentence 4, as a response sentence generated by a large-scale language model, such as "Don't give up. I'm sure you'll find it."

[0305] In the example in Figure 9A, if the system understands the user's emotions from the user sighing when user input sentence 2 is entered, the output of LLM response 2, "Did anything good happen today?", becomes unnecessary. Furthermore, if LLM response 2 is not output, user input sentence 3, "Nothing good happened," in response to LLM response 2, will not be entered either. The exchange between user 230 and the artificial intelligence response output system regarding LLM response 2 and user input sentence 3 can be said to be unnecessary if the artificial intelligence response output system understands the user's emotions. Such unnecessary exchanges in a conversation between user 230 and the artificial intelligence response output system may cause user 230 to feel uncomfortable.

[0306] In contrast, as shown in Figure 9B, when a large-scale language model extracts user emotions from user input and creates an LLM response (response sentence) to the user input, it is less likely that the conversation between user 230 and the artificial intelligence response output system will not make sense.

[0307] In the example shown in Figure 9B, similar to the example in Figure 9A, the user first inputs a greeting, "I'm home," as user input sentence 1. In response, the artificial intelligence response output device 10010 outputs a standard greeting, "Welcome back," as LLM response 1. Next, the user inputs "Haa..." (a sigh) as user input sentence 2. In the example in Figure 9B, the artificial intelligence response output system recognizes user input sentence 2, which is a sigh, as an emotional expression (emotional element), and from this emotional expression, it infers the user's feelings and understands that the user does not seem to have had a good day. Therefore, in response to user input sentence 2, which is a sigh, the artificial intelligence response output system can output a response sentence appropriate to user input sentence 2, such as "Why did you sigh?", as LLM response 2.

[0308] Then, when user input 3 such as "I think I've lost my favorite pendant" is input in response to LLM response 2, the artificial intelligence response output system can output a response sentence suitable for user input 3, such as "Don't give up. I'm sure you'll find it," as LLM response 3.

[0309] As shown in the example in Figure 9B, by appropriately understanding the user's emotions, unnecessary interactions can be suppressed in conversations between user 230 and the artificial intelligence response output system. Therefore, regardless of user 230's emotions, the discomfort user 230 feels when conversing with the artificial intelligence response output system can be reduced. Furthermore, it can feel as if the user is conversing with a human being. Consequently, conversations between user 230 and the artificial intelligence response output system tend to proceed more smoothly.

[0310] In the artificial intelligence response output system according to Example 7, as described above, when interpreting user emotions from user input, the response policy (response mode) of the large-scale language model is appropriately determined based on the interpreted user emotions. The determination of the response mode of the large-scale language model can also be described as setting the emotions of the large-scale language model. As a result, the large-scale language model generates an LLM response with appropriate content for the user input, and the generated LLM response is output to the user 230 from the artificial intelligence response output device 10010 under appropriate output conditions. In other words, the artificial intelligence response output system makes it easier to output responses that match the user's emotions.

[0311] The response output processing in the artificial intelligence response output system according to Example 7, particularly the process of determining the response policy (response mode), will be explained in more detail below. First, the series of processes in the artificial intelligence response output system from user input to the output of the LLM response will be described. Figure 10 is a diagram showing an example of the processing flow in the artificial intelligence response output system according to Example 7.

[0312] Here, the various processes in the artificial intelligence response output device 10010 are described primarily as processes performed by the control unit 1110, but they can also be described as processes performed by the client application executed by the control unit 1110. Similarly, the various processes in the large-scale language model server 20001 are described primarily as processes performed by the large-scale language model, but they can also be described as processes performed by the LLM application executed by the control unit of the large-scale language model server 20001.

[0313] As shown in Figure 10, first, as a pre-processing step for response output processing, user authentication is performed by the control unit 1110 of the artificial intelligence response output device 10010 when the artificial intelligence response output system is started (step S01). This user authentication verifies whether the user is a registered user. The method of user authentication is not particularly limited, and any existing method may be used. If the user is not a registered user, user registration processing is performed, and initial setup processing is performed. User registration processing corresponds to the authentication information generation processing described above.

[0314] During the initial setup process, various information (configuration information) necessary for generating and outputting LLM responses is registered. This configuration information corresponds to the preset instructions mentioned above, and includes, for example, information about the user's preferred response style. A response style can also be described as how the large-scale language model responds to user input sentences. In this example, during the initial setup process, the user selects one of several pre-configured response styles. Examples of response styles include "cheerful (want to be treated in a cheerful manner)," "calm (want to be treated in a calm, composed manner)," "agreeable (want to be agreed with)," "encouraging (want to be encouraged, cheered up)," and "scolding (want to be scolded, want to be treated in a strict manner)." Note that the types of response styles are not limited to those listed above and can be set arbitrarily. In addition, the initial setup process basically only needs to be performed once, but it can be performed again, for example, if the user's preferred response style changes.

[0315] The configuration information registered during the initial setup process is stored in the storage unit 1170 or memory 1109 of the artificial intelligence response output device 10010, associated with user registration information such as the user ID. If the initial setup process, including user registration, is completed, the next time a user uses the artificial intelligence response output system, the user authentication in step S01 will read out the configuration information, including the user's preferred response style, by entering the user ID, etc. Note that the artificial intelligence response output system may be made available for temporary use without user registration. However, in this case, the user will need to register the above configuration information each time.

[0316] Once user authentication is complete and user 230 provides input to the artificial intelligence response system (user input), the control unit 1110 of the artificial intelligence response output device 10010 receives the user input in step S02. User input is provided, for example, via the operation input unit 1107, the voice signal input unit 1133, or the microphone 1139, and the input content is received as the user input text (conversation text, etc.) as described in Figure 9B.

[0317] Next, in step S03, the control unit 1110 performs user emotion detection based on user input. In this example, as part of user emotion detection, the control unit 1110 obtains information regarding emotion types that classify user emotions based on user input. For example, the control unit 1110 extracts words representing user emotions (emotion words) from the content of user input (user input sentence), and obtains the evaluation value of the extracted emotion words as information regarding emotion types (user emotions). Furthermore, the control unit 1110 determines the emotion type from the evaluation value of the emotion words, etc. The control unit 1110 determines which of the pre-set emotion types the user emotion at the time of user input (which can also be called the current emotion) corresponds to.

[0318] The determination of emotion types by the control unit 1110 can also be described as the detection of user emotions. In this embodiment, the determination of emotion types can also be achieved by acquiring information by capturing images of the user with a camera and analyzing facial expressions and gestures, or by acquiring information on differences in the user's voice and speaking style (volume, speed, tone, etc.) and differences in facial expressions and gestures compared to the user's normal state. Furthermore, by combining this additional information with evaluation values ​​of emotion words and using it for determination, it becomes possible to determine emotion types with higher accuracy and in more detail. More specifically, the characteristics of each user's normal voice and speaking style, as well as their facial expressions, are stored in the non-volatile memory 1108 or storage unit 1170 of the artificial intelligence response device 10010, or in the memory or storage unit of the large-scale language model server 20001, and the differences from the normal state are determined by comparing them with the acquired additional information. While emotional vocabulary alone may allow us to identify a positive emotion, for example, we may not be able to determine which specific emotion it is, and we may have to rely on inference. However, by combining additional information, it becomes possible to accurately determine whether a user's emotion is "fun" or "joy," or even to identify a more detailed emotion, such as one that falls somewhere between "fun" and "joy."

[0319] The above method for detecting user emotions by the control unit 1110 is merely an example and is not particularly limited; existing methods may also be used. Examples of methods for detecting user emotions include the technologies described in Japanese Patent Publication No. 2023-71231, Japanese Patent Publication No. 2018-132623, and Japanese Patent Publication No. 2003-162294.

[0320] The control unit 1110 then creates control information that includes information about the emotion type, such as the evaluation value of the emotion phrase. In this example, the emotion type determination result (user emotion detection result) by the control unit 1110 is included in the information about the emotion type.

[0321] Furthermore, when a user inputs data, emotional elements such as actions (behaviors) that represent the user's emotions may be extracted, and information regarding the emotion type may be obtained based on these extracted emotional elements. For example, as explained in the example in Figure 9B, a sigh made by the user may be extracted as an emotional element (emotional expression), and an evaluation value of the extracted emotional element may be obtained. This makes it easier to determine the emotion type more accurately. In other words, it makes it easier to detect the user's emotions more accurately. The acquisition of information regarding the emotion type will be described in more detail later.

[0322] Next, the process proceeds to step S04, where the control unit 1110 creates an LLM instruction (prompt) for the large-scale language model (LLM) provided by the large-scale language model server 20001, based on the user input. This LLM instruction instructs the creation of an LLM response (response statement) to the user input (also called a user instruction). The LLM instruction may also be generated in a form that includes control information. In other words, the LLM instruction generated in step S03 may include the control information created in step S02 along with the user input.

[0323] Next, the control unit 1110 sends the created LLM instruction to the large-scale language model (LLM) of the large-scale language model server 20001 (step S05). If initial setup processing was performed in step S01, the configuration information is sent to the large-scale language model along with the user instruction. Of course, the configuration information may be sent to the large-scale language model separately from the user instruction, or even before the user instruction.

[0324] The Large-Scale Language Model (LLM) receives an LLM instruction sent from the artificial intelligence response output device 10010 (step S06). If the received LLM instruction contains configuration information, the LLM model extracts the configuration information from the LLM instruction and stores it in the storage unit of the LLM Language Model server 20001. Next, in step S07, control information is extracted from the received LLM instruction. Specifically, information regarding the emotion type contained in the control information is extracted from the LLM instruction.

[0325] Next, in step S08, the response policy (AI response mode) to user input is determined based on the emotion type information extracted from the control information. In this example, the large-scale language model determines the AI ​​response mode based on the emotion type discrimination result by the artificial intelligence response output device 10010. More specifically, the large-scale language model determines the AI ​​response mode based on the emotion type discriminated by the artificial intelligence response output device 10010 and the response style selected by the user 230 during the initial setup process. As an example, the large-scale language model selects one from a plurality of pre-configured AI response modes that corresponds to the emotion type and the user 230's preferred response style. The user 230's preferred response style information is stored, for example, as configuration information in the storage unit. Furthermore, the AI ​​response mode does not necessarily have to be determined based on the response style; for example, it may be determined based solely on the emotion type.

[0326] Depending on the AI ​​response mode determined in this way, the conditions for generating LLM response statements by the large-scale language model are appropriately changed. In other words, the conditions for generating LLM response statements by the large-scale language model are appropriately set for each AI response mode. Examples of LLM response statement generation conditions include the expression of the LLM response statement (phrasing, wording, etc.). Note that the LLM response statement generation conditions corresponding to each AI response mode are not particularly limited and can be set arbitrarily. Furthermore, information on the LLM response statement generation conditions corresponding to each AI response mode is stored, for example, in the memory or storage unit of the large-scale language model server 20001.

[0327] Then, the large-scale language model creates an LLM response to the user instruction (user input) contained in the LLM instruction, according to the AI ​​response mode selected in step S08 (step S09). That is, the large-scale language model creates an LLM response to the user instruction with a response policy that corresponds to the user's emotion (emotion type). Subsequently, in step S010, the LLM response created by the large-scale language model is transmitted to the artificial intelligence response output device 10010.

[0328] When the control unit 1110 of the artificial intelligence response output device 10010 receives an LLM response transmitted from the large-scale language model (step S011), it outputs the received LLM response in step S012. In this example, the LLM response sentence created by the large-scale language model is primarily output to the user via the speech output unit 1140. At this time, the artificial intelligence response output device 10010 outputs the LLM response to the user under output conditions corresponding to the AI ​​response mode selected by the large-scale language model.

[0329] For example, information regarding output conditions corresponding to the AI ​​response mode is transmitted from the large-scale language model to the artificial intelligence response output device 10010 along with the LLM response sentence. Alternatively, information regarding the AI ​​response mode is transmitted from the large-scale language model to the artificial intelligence response output device 10010, and the output conditions corresponding to the AI ​​response mode may be stored, for example, as setting information in the storage unit 1170 of the artificial intelligence response output device 10010.

[0330] Specific examples of output conditions for voice output include, for example, the tone of voice when reading the LLM response (answer), the volume and tone of the voice, and the reading speed. In addition, LLM response output may be provided via the display unit 10011 as text or other media output (text output), either along with or instead of voice output. In this case, output conditions include, for example, the display speed of the LLM response, the font size, and the font color. However, the output conditions corresponding to each AI response mode are not limited to these and can be set arbitrarily. For example, while using the output conditions corresponding to each AI response mode as a basis, fine adjustments such as lowering the tone of voice may be made to those conditions depending on the situation before outputting.

[0331] As described above, in the artificial intelligence response output system according to Embodiment 7, during a conversation with the user, the system takes into account the user's emotions to determine a response mode, and a large-scale language model generates an LLM response under predetermined generation conditions according to the determined response mode. The artificial intelligence response output device 10010 then outputs the LLM response to the user 230 under predetermined output conditions corresponding to the response mode. This suppresses unnecessary exchanges during conversations between the user 230 and the artificial intelligence response output system. Therefore, regardless of the user's emotions at the time, the user 230 is less likely to feel uncomfortable in conversations with the artificial intelligence response output system. Conversations between the user 230 and the artificial intelligence response output system tend to proceed more smoothly.

[0332] <Information about emotional types> Next, we will explain in more detail the information regarding the emotion type acquired in step S03 shown in Figure 10. Figure 11 is a diagram illustrating the information regarding the emotion type acquired in the artificial intelligence response output device of Example 7.

[0333] As shown in Figure 11, in the artificial intelligence response output system of Embodiment 7, the artificial intelligence response output device 10010 includes an emotion evaluation program 8050 and an emotion vocabulary database (emotion vocabulary DB) 8060 in the storage unit 1170, etc. The emotion program 8050 and the emotion vocabulary DB 8060 may be provided by an external device connected to the artificial intelligence response output device 10010. The emotion evaluation program 8050 is a type of client application, is loaded into memory 1109, and is executed by the control unit 1110. When the artificial intelligence response output device 10010 receives input from a user 230 (user input), the emotion evaluation program 8050 detects the user's emotion, quantifies the user's emotion, and manages the user emotion counter based on this user input. More specifically, the emotion evaluation program 8050 acquires an evaluation value (also called the user emotion evaluation value) of at least one of the emotion vocabulary and emotion elements in the user input.

[0334] As shown in explanation 7030, user input includes voice input via the voice signal input unit 1133 or microphone 1139, or character input via the operation input unit 1107. In addition to voice input of phrases, words, or sentences composed of multiple phrases or words (hereinafter referred to as "phrases, etc."), voice input other than phrases, etc. Examples of voice input other than phrases, etc. include input of the user's sighs, cries, sneezes, coughs, etc. Although sneezes and coughs cannot be called emotional phrases, by inputting them, it becomes possible to provide an AI response that takes the user's physical condition into consideration, achieving the same effect as providing a response that understands the user's emotions, that is, the user can get the feeling that they are interacting with a person when interacting with the AI. User input to the artificial intelligence response output device 10010 can also be done via the mobile information terminal 20010.

[0335] When user input is received by the artificial intelligence response output device 10010, as described in explanation 7040, the emotion evaluation program 8050, as its first process, checks whether the user input (the user input sentence received by the artificial intelligence response output device 10010) contains emotion phrases as expressions of user 230's emotions. In other words, it extracts emotion phrases from the user input sentence. The emotion phrase DB 8060 provided by the artificial intelligence response output device 10010 has emotion phrases pre-registered that represent user 230's emotions (for example, emotions such as joy, anger, sadness, and happiness). The emotion evaluation program 8050 checks whether the emotion phrases registered in the emotion phrase DB (hereinafter also referred to as registered emotion phrases) are included in the user input sentence. That is, the emotion evaluation program 8050 extracts phrases (emotion phrases) from the user input sentence that match the registered emotion phrases. In this example, multiple emotion types are set for user 230's emotions, and each registered emotion phrase is registered in association with one of these emotion types.

[0336] As a second process, the emotion evaluation program 8050 determines whether or not emotional elements are present in the emotional expression at the time of user input. In other words, the emotion evaluation program determines whether or not elements representing emotions (emotional elements) are included in the actions (behaviors) of user 230 at the time of user input. To put it another way, the emotion evaluation program 8050 extracts elements corresponding to emotional elements from the actions of user 230 at the time of user input. Emotional elements here may include, for example, specific actions (behaviors) of user 230 such as sighing, as well as user 230's state that is different from the usual state. For example, if user input is done via voice input, and user 230's state at that time, such as pitch of voice, speaking speed, tone of voice, etc., is significantly different from user 230's usual state, then user 230's state is extracted as an emotional element. Also, if user input is done via text input, and user 230's state at that time, such as input speed or the frequency of input errors (typos), is significantly different from the usual state, then user 230's state is extracted as an emotional element.

[0337] The method for determining whether or not user input contains an emotional element is not particularly limited, but for example, similar to the case of emotional words, actions (emotional elements) that match actions (registered emotional elements) that are pre-registered in the database are extracted from the user's actions at the time of user input. Furthermore, if the emotional element is a user state that is different from the normal state, it is determined whether or not the user state exceeds a pre-set determination threshold. For example, if the emotional element is a difference in the pitch of the user's voice, it is determined whether or not the pitch of the user's voice exceeds the determination threshold. If the user state exceeds the determination threshold, it is determined that the user 230's state is an emotional element. Note that the registered emotional elements and determination thresholds can be set arbitrarily. Furthermore, registered emotional elements may be stored in the emotional word DB together with registered emotional words, or they may be stored in a separate database from the emotional word DB, such as in the storage unit 1170. Furthermore, registered emotional elements are stored associated with each emotional type, similar to registered emotional words.

[0338] The emotion evaluation program 8050 obtains evaluation values ​​for at least one of the emotion words and emotion elements extracted from user input as information about the emotion type. In this example, the registered emotion words and registered emotion elements are associated with emotion types as described above and are registered in a predetermined database along with pre-set evaluation values. Therefore, the emotion evaluation program 8050 can obtain the evaluation values ​​of emotion words and emotion elements from the database. The evaluation values ​​for each emotion word and emotion element (hereinafter sometimes referred to as "emotion words, etc.") may be set to the same value (e.g., "1"), or they may be set to different values. Furthermore, the evaluation value for each emotion word, etc. may be a fixed value, or it may be varied according to occurrence conditions such as frequency of occurrence. In other words, the evaluation value for each emotion word, etc. may be weighted according to the occurrence conditions of the emotion word, etc. The method for calculating the evaluation value in this case is not particularly limited. The evaluation value can be calculated using any algorithm.

[0339] The evaluation values ​​of emotional words and phrases obtained in this way are aggregated for each corresponding emotional type. Furthermore, the emotional evaluation program 8050 determines the user's emotion at the time of user input based on the aggregated evaluation values. More specifically, the emotional evaluation program 8050 determines which of several emotional types the user's emotion falls under. For example, as shown in Figures 12A and 12B, the evaluation values ​​obtained by the emotional evaluation program 8050 are aggregated for each emotional type and recorded as an information table along with the counter value of the emotional type (user emotion). Such an information table is stored, for example, in memory 1109 or storage unit 1170.

[0340] Figures 12A and 12B show examples of information tables containing evaluation values ​​and counter values. The information table shown in Figure 12A is an example of a table created and updated without identifying user 230, and shows the aggregated evaluation values ​​and counter values ​​for a single use by user 230, i.e., a single login. The information table shown in Figure 12B is an example of a table created for each user 230, and shows the aggregated evaluation values ​​and counter values ​​for multiple uses by the same user, i.e., multiple logins.

[0341] Here, the extraction of emotional terms and other related information and the acquisition of evaluation values ​​by the emotion evaluation program 8050 are performed each time user input is received. That is, when emotional terms and other related information are extracted in response to each user input, the evaluation value of the extracted emotional terms and other related information is obtained. The obtained evaluation value is added to the evaluation value of the emotion type to which the extracted emotional terms and other related information belong. In other words, the obtained evaluation values ​​are aggregated for each emotion type.

[0342] As shown in Figures 12A and 12B, in this example, five emotion types are pre-set as emotion types for classifying user emotions: "neutral," "joy," "anger," "sad," and "happy." Each emotion word or phrase is associated with one of these five emotion types. When a user input is received, for example, if an emotion word or phrase belonging to the emotion type "joy" is extracted, the evaluation value of the extracted emotion word or phrase (e.g., "1") is added to the evaluation value of the emotion type "joy." As mentioned above, the evaluation value of the emotion word or phrase does not have to be "1" and can be set arbitrarily.

[0343] The emotion evaluation program 8050 determines the emotion type, for example, after aggregating the evaluation values ​​of all emotion words and phrases for a single user input. Specifically, the emotion evaluation program 8050 selects one of the five emotion types above as the user emotion at the time of user input, based on the evaluation values ​​in the information table. The criteria for selecting the emotion type are not particularly limited and can be set arbitrarily. In this example, the emotion evaluation program 8050 selects the emotion type with the highest evaluation value as the emotion type at the time of user input (also known as the current emotion type). In the example shown in Figure 12A, the evaluation value for the emotion type "joy" is "5", which is the highest. Similarly, in the example shown in Figure 12B, the evaluation value for the emotion type "joy" is "7", which is also the highest. Therefore, in the examples shown in Figures 12A and 12B, the emotion type "joy" is selected as the emotion type (user emotion) at the time of user input. At that time, "1" is added to the counter value of the selected emotion type. In other words, the "counter value" included in the information table represents the number of times each emotion type has been selected.

[0344] In the example shown in Figure 12A, the information table is created without identifying user 230. In other words, the information table containing evaluation values ​​and counter values ​​is stored without being associated with user information. In this case, the information table needs to be created and updated each time user 230 uses the service, i.e., each time they log in. On the other hand, in the example shown in Figure 12B, the information table is created for each user. In other words, the information table containing evaluation values ​​and counter values ​​is stored associated with user information such as the user ID. Therefore, it is also possible to record and store the results of accumulating evaluation values ​​and counter values ​​from multiple uses by user 230, i.e., multiple logins, as an information table.

[0345] When the cumulative results of evaluation values ​​and counter values ​​for each user are recorded, it becomes possible to understand the emotional trends of users over a certain period, and by treating this as one of the input pieces of information, it becomes possible to determine user emotions with higher accuracy. For example, it becomes possible to understand that user 230 is generally in a good mood during the spring season every year, or tends to feel down on Monday mornings. Based on this information, or by combining this information with emotion determination information at the time of user input, if it is possible to infer from the trend that, for example, since it is Monday morning today, the user should be feeling down, it becomes possible to reduce some of the processing in the AI ​​response mode to subsequent user input, thereby reducing the processing load or speeding up processing.

[0346] <Presentation of emotion types> Here, together with the evaluation value and counter value information of emotional phrases and the like, information on the emotional type selected by the emotion evaluation program 8050 may be presented to the user 230. For example, as shown in FIGS. 13A and 13B, in an information table including the evaluation value and counter value of emotional phrases and the like, a mark (for example, a star mark) indicating the emotional type selected by the emotion evaluation program 8050 is attached and presented to the user 230. In the examples shown in FIGS. 13A and 13B, the emotion type "joy" is selected as the user's emotion at the time of user input by the emotion evaluation program 8050.

[0347] Alternatively, as shown in FIG. 14A, the evaluation value and counter value of emotional phrases and the like are represented as a bar graph, and a mark (for example, a star mark) indicating the emotional type selected by the emotion evaluation program 8050 is attached to this bar graph and presented to the user 230. Also, for example, as shown in FIG. 14B, the evaluation value and counter value of emotional phrases and the like are represented as a so-called radar chart, and a mark (for example, a star mark) indicating the emotional type selected by the emotion evaluation program 8050 is attached to this radar chart and presented to the user 230. In the example shown in FIG. 14B, the radar chart is divided into ranges corresponding to each emotional type by the boundary line shown by the dotted line in the figure, and a mark (star mark) indicating the emotional type selected by the emotion evaluation program 8050 is attached to the range corresponding to the emotional type "joy". From this information, it can be seen that the emotional type is within the range of "joy" but closer to "normal". Note that such information does not necessarily have to be presented to the user 230 and may be used in the internal processing of the emotion evaluation program 8050.

[0348] <Determination of AI response mode> The emotion evaluation program 8050 then creates control information including an information table as information on the emotional type (step S04), and transmits the created control information together with the user instruction text to the large language model (step S05).

[0349] The large-scale language model (server LLM application 8020) receives an LLM instruction message containing a user instruction message sent from the artificial intelligence response output device 10010. After extracting control information from the LLM instruction message, it determines a response policy according to the emotion type (user emotion) in step S08 shown in Figure 10. As an example, the LLM application determines a response mode (AI response mode) based on the emotion type selected by the emotion evaluation program 8050. In this example, the LLM application determines the AI ​​response mode based on the emotion type as well as the preferred response style selected by the user 230. The user 230's preferred response style can be registered as setting information during the initial setup process, as described above. In this example, the user 230 can specify one of the following response styles: "cheerful," "calm," "agreeable," "encouraging," and "scolding." Note that the response style does not necessarily have to be specified. In other words, the user can also set the response style to "not specified."

[0350] Here, as shown in Figure 15 as an example, multiple AI response modes are set, including "Normal," "Sociable," "Intelligent," "Agreeable (Joyful)," "Friendly," "Agreeable (Angry)," "Gentle," "Agreeable (Sad)," "Encouraging," "Agreeable (Happy)," and "Positive." Each AI response mode is pre-registered and associated with the above emotion types and response styles. Therefore, when the LLM application obtains information on emotion types and response styles, it can appropriately select an AI response mode corresponding to these user emotions and response styles. The table shown in Figure 15 is an example where the AI ​​response output device 10010 selects the emotion type "Normal" as the user emotion and "None" as the user's preferred response style, and "Normal" is selected as the AI ​​response mode. When an AI response mode is selected by the LLM application, the valid flag for the selected AI response mode is changed from "0" to "1."

[0351] Once the AI ​​response mode is determined by the LLM application, that is, once the response strategy of the large-scale language model is determined, the large-scale language model then creates an LLM response according to the determined AI response mode (step S09). This allows the large-scale language model to create an LLM response that is appropriate for each user sentiment. For example, by switching the AI ​​response mode as needed each time user input is received, the large-scale language model can more easily create an LLM response that is more appropriate for each user input.

[0352] Subsequently, the LLM response created by the large-scale language model is sent to the artificial intelligence response output device 10010. Upon receiving the LLM response, the client application provided by the artificial intelligence response output device 10010 outputs an LLM response to the user according to the output conditions corresponding to the AI ​​response mode (step S012). By having the client application output an LLM response according to the output conditions corresponding to the AI ​​response mode in this way, it becomes easier to output an LLM response that matches the user's emotions and even matches the user's preferences.

[0353] Information regarding output conditions corresponding to the AI ​​response mode is transmitted from the large-scale language model to the artificial intelligence response output device 10010 along with the LLM response statement, for example. As shown in an example in Figure 16, the output conditions for the LLM response are associated with the AI ​​response mode and are stored, for example, in the memory or storage unit of the large-scale language model server 20001. In the example in Figure 16, the voice tone and reading speed during speech output are registered as output conditions for the LLM response. However, as mentioned above, the output conditions for the LLM response are not limited to these and can be set arbitrarily. Furthermore, information regarding the output conditions for the LLM response may be pre-stored as setting information in the storage unit 1170 of the artificial intelligence response output device 10010, for example. In this case, in step S010 of Figure 10, the large-scale language model only needs to transmit the information regarding the selected AI response mode along with the LLM response to the artificial intelligence response output device 10010.

[0354] Furthermore, in the artificial intelligence response output system according to Example 7, as described above, the emotion evaluation program 8050 determines the current emotion type of the user 230 and transmits the information of the emotion type determination result to the large-scale language model server 20001 as information related to the emotion type. The LLM application then determines the AI ​​response mode based on the information related to the emotion type. However, the procedure for determining the AI ​​response mode is not limited to this. For example, the emotion evaluation application 8050 may also select an AI response mode according to the user's current emotion type and transmit the information of the AI ​​response mode selection result to the large-scale language model server 20001 as information related to the emotion type. In this case, the LLM application creates an LLM response according to the AI ​​response mode included in the information related to the emotion type.

[0355] Furthermore, for example, the emotion evaluation program 8050 may calculate evaluation values ​​for each emotion type and transmit this information to the large-scale language model server 20001 as control information including information about the emotion type. The LLM application on the large-scale language model server 20001 may then determine the user's current emotion type based on the information about the emotion type. Also, the large-scale language model on the large-scale language model server 20001 is a multimodal large-scale language model as described above. Therefore, when the large-scale language model determines the emotion type, the imaging unit 1180 of the artificial intelligence response output device 10010 may capture an image or video of the user 230's face and transmit the capture results along with evaluation values ​​to the large-scale language model. The large-scale language model can then determine the emotion type (user emotion) more appropriately by determining the emotion type based on the control information including the image or video of the user 230's face.

[0356] <Example 8> Embodiment 8 of the present invention is a modification of Embodiment 7. More specifically, Embodiment 8 is a modification of the method for selecting emotion types using the emotion evaluation program 8050, and the configuration of the artificial intelligence response output system is the same as in Embodiment 7. Embodiment 8 will mainly describe the differences from Embodiment 7, and repeating explanations of configurations similar to those in Embodiment 7 will be omitted.

[0357] The procedure for outputting an LLM response in the artificial intelligence response output system according to Example 8 is the same as in Example 7 (see Figure 10), but the user emotion detection process in step S03, that is, the method for selecting the emotion type and determining the AI ​​response mode, differs from that of Example 7. The method for selecting the emotion type and determining the AI ​​response mode according to Example 8 will be described below.

[0358] In the artificial intelligence response output device 10010 according to Example 8, when user input is received, the emotion evaluation program 8050 extracts emotion words, etc. from the user input and obtains the evaluation value of the extracted emotion words, etc. as information about the emotion type (user emotion). The obtained evaluation values ​​of emotion words, etc. are aggregated for each corresponding emotion type. Based on the aggregated evaluation values, the emotion evaluation program 8050 determines the emotion type as the user emotion at the time of user input. As explained in Example 7, the evaluation values ​​of emotion words, etc. are aggregated for each emotion type and recorded as an information table together with the emotion type counter value (see Figures 12A and 12B). Based on this information (information table) of the evaluation values ​​and counter values ​​of the emotion type, the emotion evaluation program 8050 selects the emotion type.

[0359] In Example 7, the emotion evaluation program 8050 selected the one with the highest evaluation value as the emotion type when the user inputted it. In contrast, in Example 8, the emotion evaluation program 8050 selects two emotion types as the emotion type when the user inputted it. In other words, the emotion evaluation program 8050 selects a primary emotion type (indicated as "Emotion Type [Primary]" in the diagram) and a secondary emotion type (indicated as "Emotion Type [Secondary]" in the diagram) as the emotion type when the user inputted it.

[0360] As shown in Figure 17 as an example, in this example, five emotion types are pre-set as primary emotion types: "neutral," "joy," "anger," "sad," and "happy." In addition, multiple secondary emotion types are set for each primary emotion type. The secondary emotion types set for each primary emotion type may be the same, but in this example, some are different depending on the primary emotion type. Specifically, as shown in Figure 17, for the primary emotion type "neutral," the secondary emotion types set are "none (not set)," "joy," "anger," "sad," and "happy." For the primary emotion type "joy," the secondary emotion types set are "neutral," "happy," and "none." Similarly, for the primary emotion type "anger," the secondary emotion types set are "neutral," "sad," and "none." For the primary emotion type "sad," the secondary emotion types set are "neutral," "anger," and "none." For the primary emotion type "happy," the secondary emotion types set are "neutral," "joy," and "none."

[0361] In the artificial intelligence response output system according to Example 8, evaluation values ​​of emotional words and other items extracted from user input are aggregated for each corresponding emotion type (see Figures 12A and 12B), and the emotion evaluation program 8050 selects a primary emotion type and a secondary emotion type as the user's emotion at the time of user input based on the aggregated evaluation values ​​and other information (information table). The selection criteria for the primary emotion type and secondary emotion type are not particularly limited and can be set arbitrarily. In Example 8, as an example, the primary emotion type and secondary emotion type are selected based on the height of the evaluation values ​​aggregated as an information table. Specifically, among multiple emotion types, the one with the highest evaluation value in the information table is selected as the primary emotion type. Furthermore, among the emotion types, the one with the next highest evaluation value in the information table is selected as the secondary emotion type.

[0362] For example, in the example in Figure 12A, the emotion type with the highest evaluation score, "joy," is selected as the primary emotion type, and the emotion type with the second highest evaluation score, "fun," is selected as the secondary emotion type. Similarly, in the example in Figure 12B, the emotion type with the highest evaluation score, "joy," is selected as the primary emotion type, and the emotion type with the second highest evaluation score, "normal," is selected as the secondary emotion type. Note that the emotion type counter value may be added when the primary emotion type is selected, or it may be added separately when the primary emotion type and secondary emotion type are selected.

[0363] In addition, in Example 8, AI response modes, which are response strategies for a large-scale language model, are registered in association with the primary and secondary emotion types selected in this manner. As shown in Figure 17 as an example, in Example 8, multiple types of AI response modes, including "normal," "sociable," "intellectual," "dedicated," "proactive," "agreeable (joyful)," "cheerful," "friendly," "agreeable (angry)," "cautious," "gentle," "agreeable (sad)," "kind," "encouraging," "agreeable (happy)," "optimistic," and "positive," are stored in association with the primary and secondary emotion types.

[0364] Therefore, when the LLM application obtains information on the primary and secondary emotion types, it can appropriately select an AI response mode corresponding to these primary and secondary emotion types. The example shown in Figure 17 is an example where the emotion type "normal" is selected as the primary emotion type and the emotion type "none" is selected as the secondary emotion type, and "normal" is selected as the AI ​​response mode. When the LLM application selects an AI response mode, the valid flag for the selected AI response mode is changed from "0" to "1". Incidentally, the secondary emotion type "none" is selected, for example, when the evaluation value of each emotion type is "0" except for the emotion type selected as the primary emotion type. After that, as in Example 7, the large-scale language model creates an LLM response according to the AI ​​response mode (step S09).

[0365] As described above, in Example 8, the emotion e...

Claims

1. A response output system, Large-scale language models and, An input section that accepts user input, A control unit that generates an instruction sentence for the large-scale language model based on the user input and obtains the response generated by the large-scale language model to the instruction sentence, The control unit comprises an output unit that outputs the response acquired by the control unit to the user, The control unit transmits information regarding the emotion type used to classify user emotions obtained from user input to the large-scale language model. The large-scale language model generates the response according to a response mode selected according to the emotion type, The output unit outputs the response to the user according to the output conditions corresponding to the response mode. Response output system.

2. In the response output system according to claim 1, The selection of the response mode is made based on the emotion type and the response style specified by the user. Response output system.

3. In the response output system according to claim 1, The information relating to the emotion types includes information indicating the primary emotion type and information indicating secondary emotion types that further classify the primary emotion type. The aforementioned emotion type is determined by the combination of the primary emotion type and the secondary emotion type. Response output system.

4. In the response output system according to claim 1, The selection of the response mode is performed each time the response is generated by the large-scale language model. Response output system.

5. In the response output system according to claim 1, The selection of the response mode is performed each time the number of times the response has been generated by the large-scale language model reaches a preset number. Response output system.

6. In the response output system according to claim 1, The control unit acquires an evaluation value for at least one of the emotional words and emotional elements in the user input. The aforementioned emotion type is determined based on the aforementioned evaluation value. Response output system.

7. In the response output system according to claim 6, It has a memory unit, The evaluation value is stored in the storage unit while the user input is performed multiple times. The selection of the response mode is performed on the condition that the accumulated evaluation value, which is the sum of the evaluation values, reaches a preset accumulation threshold. Response output system.

8. In the response output system according to claim 6, The control unit determines the emotion type based on the evaluation value and transmits the information of the emotion type determination result to the large-scale language model as information related to the emotion type. Response output system.

9. In the response output system according to claim 6, The control unit transmits information regarding the evaluation values ​​of the emotion phrases and emotion elements to the large-scale language model as information regarding the emotion type. The large-scale language model determines the emotion type based on the evaluation value. Response output system.

10. In the response output system according to claim 6, The control unit determines the emotion type based on the evaluation value, selects the response mode according to the emotion type, and transmits the information of the response mode selection result to the large-scale language model as information regarding the emotion type. Response output system.

11. In the response output system according to claim 6, The aforementioned large-scale language model is configured as a group of large-scale language models including a plurality of response models corresponding to each of the response modes, The control unit determines the emotion type based on the evaluation value, selects the response mode according to the emotion type, and transmits the instruction message to the response model corresponding to the selected response mode. Response output system.

12. In the response output system according to claim 1, The output unit outputs the response generated by the large-scale language model to the user as audio. The output conditions include conditions for the reading speed of the response, the tone of the reading voice, or the volume of the reading voice. Response output system.

13. In the response output system according to claim 1, The output conditions include a time condition from the time the user input is received by the input unit until the output is produced by the output unit. Response output system.

14. A response output system according to claim 1, The aforementioned large-scale language model is stored on a server connected to the internet. The control unit is stored in a device different from the server, The large-scale language model and the control unit transmit and receive information via the internet. Response output system.

15. A response output device, An input section that accepts user input, A control unit that generates an instruction sentence for a large-scale language model based on the user input and obtains the response generated by the large-scale language model for the instruction sentence, The control unit comprises an output unit that outputs the response acquired by the control unit to the user, The control unit transmits information regarding the emotion type for classifying the user emotion obtained from the user input to the large language model, and receives the response generated by the large language model according to the response mode selected according to the emotion type. The output unit outputs the response to the user according to the output conditions corresponding to the response mode. Response output device.

16. In the response output device according to claim 15, The control unit acquires an evaluation value for at least one of the emotion phrases and emotion elements in the user input, and transmits the evaluation value information for the emotion phrases and emotion elements to the large-scale language model as information regarding the emotion type. Response output device.

17. In the response output device according to claim 15, The control unit acquires an evaluation value for at least one of the emotion phrases and emotion elements in the user input, determines the emotion type based on the evaluation value, and transmits the information of the emotion type determination result to the large-scale language model as information related to the emotion type. Response output device.

18. In the response output device according to claim 15, The control unit acquires an evaluation value for at least one of the emotional words and emotional elements in the user input, determines the emotional type based on the evaluation value, selects the response mode according to the emotional type, and transmits the information of the response mode selection result to the large-scale language model as information regarding the emotional type. Response output device.

19. In the response output device according to claim 15, The output unit outputs the response as audio to the user. The output conditions include conditions for the reading speed of the response, the tone of the reading voice, or the volume of the reading voice. Response output device.

20. In the response output device according to claim 15, The output conditions include a time condition from the time the input unit receives the user input until the output unit outputs the result. Response output device.