Response output device
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MAXELL LTD
- Filing Date
- 2023-08-08
- Publication Date
- 2026-08-06
AI Technical Summary
【0007】 本発明によれば、より好適な応答出力技術を提供できる。これ以外の課題、構成および効果は、以下の実施形態の説明において明らかにされる。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a response output device. [Background technology]
[0002] A response output technique using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special table 2019-528512 publication Summary of the Invention [Problem to be solved by the invention]
[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.
[0005] An object of the present invention is to provide a more suitable response output technique. [Means for solving the problem]
[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problem, and an example of such a response output device may include a control unit that acquires a response to an instruction sentence for a large-scale language model from the large-scale language model, a display unit, and a voice output unit, and the control state by the control unit may include a state in which the response from the large-scale language model is output via the display unit or the voice output unit. Effect of the Invention
[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become apparent from the following description of the embodiments. [Brief description of the drawings]
[0008] [Figure 1A] 1 is a diagram showing an example of an artificial intelligence response output device and system according to an embodiment of the present invention; [Figure 1B] 1 is a diagram showing an example of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 1C] FIG. 2 is a diagram showing an example of the operation of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 2A] 1 is an explanatory diagram of an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2B] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2C] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2D] FIG. 2 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2E] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2F] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2G] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2H] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2I] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2J] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2K] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 2L] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3A] 1 is an explanatory diagram of an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3B] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3C] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3D] FIG. 2 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3E] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3F] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3G] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3H] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 3I] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 4A] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 4B] FIG. 2 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. [Figure 5A] FIG. 2 is an explanatory diagram of an example of the operation of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5B] FIG. 2 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5C] FIG. 2 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5D] FIG. 2 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 6] FIG. 2 is an explanatory diagram of an example of a response generation process of an artificial intelligence response output device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, the embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. In addition, in all the drawings for explaining the present invention, the same reference numerals are given to parts having the same functions, and repeated explanations thereof may be omitted.
[0010] In addition, when the artificial intelligence response output device according to each embodiment of the present invention has a display screen, it may be called a display device. When the artificial intelligence response output device has a voice output function, it may be called a voice output device. The artificial intelligence response output device may simply be called an information processing device. A system including an artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence response output system. In addition, when the artificial intelligence response output device provides a response service of a large-scale language model that is an artificial intelligence to a user and helps the user, the artificial intelligence response output device or the display output of the artificial intelligence response output device can be an artificial intelligence (AI) assistant for the user. Therefore, in this case, the artificial intelligence response output device may be called an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an AI assistant system or an AI assistant display system. In addition, in this case, the artificial intelligence response output device serves as an interface between the user and the artificial intelligence, so it may be called an artificial intelligence interface device. In this case, a system including an artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence interface system.
[0011] <Example 1> As a first embodiment of the present invention, an AI response output device and a system thereof that output a response from a large-scale language model AI will be described.
[0012] An example of an artificial intelligence response output device 10010 of the present invention will be described with reference to Fig. 1A. In addition, in the case where the artificial intelligence response output device 10010 cooperates with a large-scale language model server 19001 through communication or the like, an example of a system in which the artificial intelligence response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.
[0013] In the example of FIG. 1A, the artificial intelligence response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen on which an image is projected from the rear, or a floating image on which an optical image is formed in the air. When the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Moreover, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which pixels emit light by themselves. Moreover, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.
[0014] In the example of Fig. 1A, the voice output unit 1140 of the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139 and can pick up the user's voice. By voice input from the microphone 1139 or operation input from the user via an operation input unit described later, the AI response output device 10010 can acquire user input that is the basis of instruction sentences (prompts) for the large-scale language model, which is the AI.
[0015] The AI response output device 10010 may include a local large-scale language model in the AI response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output of the display unit 10011 and / or an audio output of the audio output unit 1140.
[0016] In addition, the artificial intelligence response output device 10010 may not be equipped with a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or an audio output on the audio output unit 1140.
[0017] Alternatively, the artificial intelligence response output device 10010 may also include a local large-scale language model, and may be configured to communicate with an external large-scale language model server 19001 having a large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be switched and either one may be output as the display output of the display unit 10011 and / or the audio output of the audio output unit 1140. Alternatively, a response generated based on both the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be output as the display output of the display unit 10011 and / or the audio output of the audio output unit 1140.
[0018] The configuration in the case where the artificial intelligence response output device 10010 communicates with and cooperates with an external large-scale language model server 19001 or a large-scale language model server 20001 is as follows. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication may also be used. The communication path from the communication unit 1132 to the communication device 19011 may include a wired portion and a wireless portion, or may go through a router or a repeater. In addition, the communication path from the communication unit 1132 to the Internet 19000 may include a wired portion and a wireless portion, or may go through a router or a repeater. The artificial intelligence response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as one system.
[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided in the artificial intelligence response output device 10010, the large-scale language model provided in the large-scale language model server 19001, and the multi-modal large-scale language model provided in the large-scale language model server 20001.
[0020] In the example of FIG. 1A, the display unit 10011 displays each element in two display areas: a prompt display area 10051 where a user inputs a prompt to a large-scale language model, which is an artificial intelligence, and an artificial intelligence response display area 10061 where a response from the large-scale language model is displayed. In the example of FIG. 1A, the prompt display area 10051 displays an icon 10052 indicating a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, a video 10055 as a component of the prompt, and the like. In the example of FIG. 1A, the artificial intelligence response display area 10061 displays an icon 10062 indicating an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, a video 10065 as a component of the response from the artificial intelligence, and the like. The display example of the display unit 10011 of the artificial intelligence response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the artificial intelligence response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.
[0021] Here, the large-scale language model will be described. The large-scale language model is also written as LLM (Large Language Model). Specifically, various models such as GPT-1, GPT-2, GPT-3, InstructGPT, ChatGPT, etc. have been published. These technologies may be used in this embodiment as well. These large-scale language models are artificial intelligence models that have been generated by large-scale pre-learning on the natural language contained in numerous documents and texts that exist in the human world. The number of parameters in the artificial intelligence model exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of the base model is a model called Transformer. For example, Reference 1 has been published as an example of learning of these models.
[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Among them, more advanced ones are capable of natural language question answering (also called dialogue or conversation), natural language proposal generation, and programming code generation. Since the number of parameters of these artificial intelligence models is very large, huge amounts of data and computational resources are required for learning. Therefore, learning this level of artificial intelligence only for a specific purpose is very resource inefficient. Therefore, models are generated by performing large-scale pre-learning as foundation models that can be applied to various purposes. For example, the large-scale language model server 19001 shown in FIG. 1A may be configured to include such a large-scale language model and to be available on various terminals via an API (Application Programming Interface). In addition, the artificial intelligence response output device 10010 shown in FIG. 1A may be configured to include a local large-scale language model and to be used by the artificial intelligence response output device 10010 itself. The learning of any large-scale language model itself can be generated by separately conducting large-scale pre-learning, and the generated large-scale language model can be replicated and provided for the large-scale language model server 19001 and the artificial intelligence response output device 10010. In this way, instead of conducting pre-learning for each application or terminal, if a large-scale language model that is a base model generated by conducting large-scale pre-learning is replicated and used on each server or terminal, the resources consumed for learning can be shared, resulting in good resource efficiency.
[0024] In addition, even if a large-scale language model is used as a base model generated by performing large-scale pre-learning, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.
[0025] Moreover, the large-scale language model can pre-learn natural language and perform input / output processing targeting natural language. Furthermore, a multimodal large-scale language model artificial intelligence capable of processing not only natural language text information but also other types of information can be applied to the embodiment of the present invention. In FIG. 1A, a large-scale language model server 20001 is shown as a server having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model artificial intelligence include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may be used in the present embodiment. Note that these multimodal large-scale language models are artificial intelligence models that are generated by performing large-scale pre-learning on a large number of documents and texts that exist in the human world, natural language contained in the text, and types of information other than text information of natural language (e.g., images, videos, audio, etc.). In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. Hereinafter, types of information other than natural language text information such as images, videos, and audio may be referred to as non-natural language information sources.
[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the user input.
[0028] The AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power source 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, etc. The AI response output device 10010 may have a large screen such as a monitor or a television.
[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. When the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which pixels emit light by themselves. When the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can obtain a user input that is the basis of an instruction sentence (prompt) to a large-scale language model that is an artificial intelligence.
[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, etc. Using these communication methods, the communication unit 1132 of the artificial intelligence response output device 10010 can communicate with the communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include a wired portion and a wireless portion, or may go through a router or a repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and may communicate using a LAN communication method. This allows the artificial intelligence response output device 10010 to communicate with various servers connected to the Internet 19000.
[0031] The artificial intelligence response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.
[0032] The power supply 1106 converts an AC current input from the outside via the external power supply input interface 1111 into a DC current, and supplies the necessary DC current to each part of the AI response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. In addition, the secondary battery 1112 supplies power to each part requiring power via the external power supply input interface 1111 when power is not supplied from the outside.
[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit such as a remote controller, or an infrared light receiving unit, and inputs a signal for an operation different from a touch operation by a user to a touch operation input sensor of the display unit 10011. In addition to a user who touches the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the AI response output device 10010. By the operation input of the user via the operation input unit 1107, the AI response output device 10010 can obtain a user input that is the basis of an instruction sentence (prompt) to a large-scale language model that is an AI. Note that there may be a modified example in which the touch operation input sensor of the display unit 10011 is also included as a part of the operation input unit 1107.
[0034] The video signal input unit 1131 is connected to an external video output device and inputs video data. The video signal input unit 1131 may be implemented by various digital video input interfaces. For example, the video signal input unit 1131 may be implemented by a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may be implemented by various USB interfaces or the like.
[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured with an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may be any of various USB interfaces, or the like. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.
[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds and error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as an Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
[0037] The microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110 described later performs voice recognition processing on the generated audio signal to obtain text information from the audio signal. By using the voice input from the microphone 1139, the AI response output device 10010 can obtain a user input that is the basis of a command (prompt) for a large-scale language model, which is an AI.
[0038] The imaging unit 1180 is a camera having an image sensor. The camera may be provided on the front side of the display unit 10011 of the artificial intelligence response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 is described as having both a front camera and a back camera.
[0039] The storage unit 1170 is a storage device that records various information such as various data such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device such as a hard disk drive (HDD) or a semiconductor element memory such as a solid state drive (SSD). For example, various information such as various data such as video data, image data, and audio data may be recorded in the storage unit 1170 in advance at the time of product shipment. In addition, the storage unit 1170 may record various information such as various data such as video data, image data, and audio data acquired from an external device, an external server, etc. via the communication unit 1132. The video data, image data, etc. recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, etc. recorded in the storage unit 1170 may be output to an external device, an external server, etc. via the communication unit 1132.
[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be called a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may be called a video processing unit or an image processing unit. The video control unit 1160 performs control of video switching, such as which video signal is input to the display unit 10011 out of the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform control of image processing on the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling processing for enlarging, reducing, transforming, etc. an image, brightness adjustment processing for changing the luminance, contrast adjustment processing for changing the contrast curve of an image, and Retinex processing for decomposing an image into light components and changing the weighting of each component.
[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the artificial intelligence response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.
[0042] The non-volatile memory 1108 stores various data used by the AI response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, data of objects to be operated by user operations, layout information, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read out various software from the storage unit 1170 and expand and store it in the memory 1109.
[0043] The local LLM processing unit 10028 includes a memory capable of holding a large-scale language model (LLM), and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may execute learning in addition to inference. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of the large-scale language model in the local environment of the artificial intelligence response output device 10010.
[0044] The control unit 1110 controls the operation of each unit connected to it. The control unit 1110 may also cooperate with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the artificial intelligence response output device 10010. The control state by the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140 such as a speaker.
[0045] In addition, when there is input from the user via the above-mentioned touch panel, microphone 1139 or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 provided in the artificial intelligence response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001, and responses are obtained from these large-scale language models. All of this control can be performed by the control unit 1110.
[0046] The storage unit 1170 may also store a fixed response phrase database (which may be written as a fixed response phrase DB) for outputting a fixed phrase as a response to an instruction sentence of the artificial intelligence response output device 10010. The control unit 1110 may perform control to generate a response to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of a fixed response phrase database. In the example of FIG. 1C, a fixed response phrase output by the artificial intelligence response output device 10010 is stored for each condition to which a condition number is assigned. For example, as in condition number 1, when the user inputs "Good morning" via the above-mentioned touch panel, microphone 1139, or operation input unit 1107, a response may be output using "Good morning" or "Today is XX day of XX month, isn't it?" as a fixed response phrase. The 〇 part of "XX day of XX month" may be generated using information stored in the memory 1109 of the artificial intelligence response output device 10010.
[0047] In addition, in the example of the standard response phrase in the database shown in FIG. 1C, if multiple standard response phrases separated by / are stored, the control unit 1110 may control to randomly select one of the standard response phrases using a random number or the like and output a response. In this way, it is possible to eliminate and improve the situation where responses under the same conditions become monotonous. The explanation of the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control to output the standard response phrase of each example shown in FIG. 1C for the condition contents of each example shown in FIG. 1C.
[0048] Next, an example of condition number 5 shown in Fig. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I might not know about that." By responding in this way, the user can be prompted to re-enter the input, and the corrected user input can be waited for.
[0049] Next, an example of condition number 6 shown in FIG. 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in FIG. 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using "It seems to be not working properly" as a response template. By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is not working properly, and to encourage the user to deal with the error, etc.
[0050] The AI response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C, instead of a response from a large-scale language model such as a local large-scale language model provided in the AI response output device 10010, a large-scale language model provided in the large-scale language model server 19001, and a multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, a response may be output that combines the responses from these large-scale language models and a response using the fixed response phrase database (fixed response phrase DB).
[0051] The fixed response phrase database (fixed response phrase DB) of FIG. 1C described above may be stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 may use it. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the artificial intelligence response output device 10010, instead of a response generated by a large-scale language model stored in each server. In this way, even if the artificial intelligence response output device 10010 is not provided with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0052] In the above description, the artificial intelligence response output device 10010 has a display panel with a display screen using fixed pixels. This concept may include a projection type image display device (projector) that provides a projection optical system behind the display panel with a display screen using fixed pixels and projects an optical image of the image on the display panel of the display screen onto a screen or wall.
[0053] 1A and 1B, an example has been described in which the artificial intelligence response output device 10010 includes the display unit 10011. However, the artificial intelligence response output device 10010 according to the embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the artificial intelligence may be configured to receive input from a user to the artificial intelligence via the voice signal input unit 1133 or the microphone 1139, and output a response from the artificial intelligence, such as a large-scale language model, to the input from the user via the voice output unit 1140.
[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the above-described embodiment 1 of the present invention, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference of the artificial intelligence, such as a large-scale language model owned by a server device on a network or a local large-scale language model owned by the artificial intelligence response output device itself.
[0055] <Example 2> Next, as a second embodiment of the present invention, an example will be described in which the AI response output device 10010 described in the first embodiment is connected to the Internet, and is connected to a server equipped with a large-scale language model AI via the Internet to perform operation. In this embodiment, differences from the first embodiment will be described, and repeated explanations of configurations similar to those of the first embodiment will be omitted.
[0056] An example of a connection state between the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment of the present invention will be described with reference to FIG. 2A. The AI response output device 10010 according to the second embodiment may be called a character conversation device. A system including the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment may be called a character conversation system. An image of a character 19051 is displayed on a display unit 10011 displayed by the AI response output device 10010. The image of the character 19051 is an image generated by rendering a 3D model of the character in a virtual space.
[0057] In addition, the character in this embodiment can provide the user with the service of a large-scale language model, which is an artificial intelligence, and can be of assistance to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and the character conversation system in this embodiment may be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, and an AI assistant response output system.
[0058] In the example of FIG. 2A, the voice output unit 1140 of the AI response output device 10010 is composed of a speaker. In addition, the AI response output device 10010 is equipped with a microphone 1139, and can collect the user's voice. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A, an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication may also be used. The communication path from the communication unit 1132 to the Internet 19000 may include wired and wireless parts. The AI response output device 10010 can communicate with a large-scale language model server 19001 via the communication device 19011 and the Internet 19000. In addition, the AI response output device 10010 can communicate with a second server 19002 different from the large-scale language model server 19001 via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 may be considered as one system.
[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2B. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that Fig. 2B omits the illustration of communication paths such as the Internet 19000 shown in Fig. 2A. Fig. 2B also illustrates a user 230 of the artificial intelligence response output device 10010.
[0060] Here, we will explain a series of operations of the AI response output device 10010. The AI response output device 10010 expands a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing various processes described below.
[0061] First, the AI response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is a natural language. The extraction of the text of the words spoken by the user 230 may be performed continuously for all words, but may also be started when the user utters words within a predetermined period after a trigger keyword. For example, the trigger keyword may be when the user utters "Hello" followed by the character name. For example, if the name of the character 19051 is "Koto", "Hello, Koto!" may be the trigger keyword.
[0062] The character operation program of the AI response output device 10010 creates an instruction (prompt) based on the text of the words spoken by the user 230, and transmits the instruction to the large-scale language model server 19001 using an API. Here, the instruction may be metadata in which information described in a notation using tags such as a markup format of a markup language, a notation using predetermined symbols such as a Markdown format, or an object notation of a predetermined script such as JSON is stored. The instruction stores text information in a natural language as a main message. The types of instruction transmitted from the AI response output device 10010 to the large-scale language model server 19001 include a setting instruction that stores instructions such as initial settings, and a user instruction that reflects an instruction from a user. Type identification information that identifies whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. When the character operation program of the artificial intelligence response output device 10010 creates a prompt based on the text of the words spoken by the user 230, it creates the user prompt and transmits it to the large-scale language model server 19001.
[0063] Next, the large-scale language model of the artificial intelligence of the large-scale language model server 19001 executes inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response including text information in natural language based on the result. The large-scale language model server 19001 uses an API to transmit the response to the artificial intelligence response output device 10010. The response stores text information in natural language as a main message. Here, the response may be metadata that stores information described in a notation of the same format as the instruction (notation using tags such as a markup format of a markup language, notation using predetermined symbols such as a markdown format, or object notation of a predetermined script such as JSON, etc.). When the response uses the same format as the instruction, type identification information may be stored in a part other than the main message to indicate that the information is a different type from the initial setting instruction and the user instruction. For example, information indicating that it is an answer sentence from the large-scale language model may be stored.
[0064] Next, the AI response output device 10010 receives a response from the large-scale language model server 19001, and extracts the natural language text information stored as the main message in the response. The character operation program of the AI response output device 10010 uses a voice synthesis technique to generate a natural language voice as a reply to the user based on the natural language text information extracted from the response, and outputs the voice from the voice output unit 1140, which is a speaker, so that it sounds as if it is the voice of the character 19051. This process may be expressed as the character's "utterance."
[0065] As described above, the AI response output device 10010 and the large-scale language model server 19001 perform processing to generate a response voice of the character 19051 in response to a comment from the user 230. Conversation examples 1 to 5 in FIG. 2C show specific examples of the response voice of the character 19051 in response to a comment from the user 230. In this way, the user 230 can have a conversation with the character 19051 as if it were a real person.
[0066] According to the AI response output device 10010 of Fig. 2B or a system including the AI response output device 10010 described above, it is not necessary to mount a large-scale language model, which requires huge amounts of data and computational resources for learning, on the AI response output device 10010 itself. In addition, the advanced natural language processing capabilities of the large-scale language model can be utilized via the API, and when a user speaks to a character, a more suitable response can be given to the user, enabling a more suitable conversation to be held.
[0067] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2D. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2D shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the source of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response which is the response.
[0068] FIG. 2D also shows an exchange of instructions and responses in chronological order, from a display setting instruction, a first round of user instructions and their responses, to a fourth round of user instructions and their responses.
[0069] As shown in FIG. 2D, the setting instruction can instruct the large-scale language model of the artificial intelligence of the large-scale language model server 19001 to initially set the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, and the like. The user's name can also be understood as the initial setting. This allows the large-scale language model to generate responses from the first round onward while respecting the role. Then, from the perspective of the user who hears the voice of the character 19051 based on the responses from the first round onward, the character 19051 feels as if it has the setting and personality of the person described in the setting instruction. In addition, the large-scale language model server 19001 according to this embodiment is provided with a memory that stores the contents of the conversation until the series of conversations are completed, and is configured to store a series of user instruction statements and their responses and then generate a response. This allows a conversation as shown in FIG. 2D to be realized.
[0070] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2E. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2E shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the source of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response which is the response.
[0071] Fig. 2E shows an example of a case where the user 230 talks to the character 19051 again to start a new conversation after the series of conversations shown in Fig. 2D has ended. Fig. 2E shows an exchange of instructions and responses in chronological order from the first round of user instructions and their responses to the third round of user instructions and their responses.
[0072] Here, the "end" of the "continuation of a series of conversations" refers to a process in which the large-scale language model server 19001 erases the memory of the conversation that the large-scale language model server 19001 has held while the series of conversations is continuing, when a predetermined condition is satisfied. An example of the predetermined condition is, for example, a case in which the AI response output device 10010 issues an instruction to the large-scale language model server 19001 to "end" the "continuation of a series of conversations" by an instruction sentence. Another example of the predetermined condition is, for example, a case in which the AI response output device 10010 has not sent an instruction sentence to the large-scale language model server 19001 for the series of conversations for a predetermined time or more (timeout). Another example is a case in which, after performing authentication processing in the connection between the AI response output device 10010 and the large-scale language model server 19001, the authentication processing is lost due to factors such as communication disconnection or the power off (OFF) of the AI response output device 10010 while exchanging the instruction sentence and response.
[0073] When the "continuation of a series of conversations" is "ended", the large-scale language model server 19001 erases the memory of the conversation that was held while the series of conversations was continuing from the large-scale language model server 19001. Therefore, although the conversation shown in FIG. 2E is after the series of conversations shown in FIG. 2D, the server response to the user instruction is a response with contents of a state in which the name of the character set in the large-scale language model, the role to be played, the characteristics of the conversation, the user's name, etc., contained in the setting instruction shown in FIG. 2D, are not remembered at all. Similarly, the conversation shown in FIG. 2E is a response with contents of a state in which the series of conversations shown in FIG. 2D are not remembered at all. In other words, the "end" of the "continuation of a series of conversations" shown in FIG. 2D means that the conversation in FIG. 2E starts from a state in which the large-scale language model of the artificial intelligence of the large-scale language model server 19001 is initialized.
[0074] This causes the user 230 to feel as if the character 19051 has lost its memory of the user 230 or is a completely different person. From the user 230's perspective, the character's response is very unnatural, and the user 230 feels lonely and disappointed. With this type of behavior, there is a problem in that it is not possible to ensure the sameness of the settings and memories of the character 19051 displayed on the AI response output device 10010, such as the name, role, conversation characteristics, and personality.
[0075] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2F. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2F shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the source of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response which is the response.
[0076] FIG. 2F shows an example of a case where the user 230 speaks to the character 19051 again to have a new conversation after the series of conversations shown in FIG. 2D has ended. In the process of FIG. 2F, unlike the process of FIG. 2E, when starting a new conversation, the AI response output device 10010 transmits a setting instruction as the first instruction to the large-scale language model server 19001. The setting instruction stores the same natural language text as the setting instruction of the initial setting in FIG. 2D. This may be expressed as a reset text. The setting instruction stores natural language text that explains the history of past conversations. This may be expressed as a conversation history text. The AI response output device 10010 may record the history of past conversations as natural language text information in the storage unit 1170 while the series of conversations described in FIG. 2D are being continued, linking the history of the conversation to information on the date and time of the conversation as information on the date and time of the conversation. If there are conversations on different dates, each conversation may be recorded linked to information on the date and time, and the conversation history may be accumulated. When generating a setting instruction for the first instruction of a later conversation such as that shown in FIG. 2F, the natural language text information of the conversation recorded in storage unit 1170 and the information on the date and time when the conversation took place can be read out and used to generate the setting instruction.
[0077] When natural language text information of the history of past conversations is used to generate the setting instruction sentence, the format can be determined freely to a certain extent since it is data to be sent to a large-scale language model, but as shown in Fig. 2F, a prefix or suffix in natural language such as "I talked about the following on XX day," or "You talked about the following on XX day," can be prepared and merged with the natural language text information of the recorded conversation to generate the text of the setting instruction sentence. Also, information on the date and time of the conversation read from the storage unit 1170 can be merged with the "XX day" portion and used as part of the text of the setting instruction sentence.
[0078] Even if the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations has ended, if the generation process and transmission process of the setting instruction text in Fig. 2F described above are performed, the response to the user instruction text thereafter will reflect the settings and conversation history of the character's role, name, conversation characteristics, personality, and / or conversation characteristics at the time of the previous conversation. This is more preferable because it is recognized by the user as ensuring the sameness of the settings and memories of the character's role, name, conversation characteristics, or personality at the time of the previous conversation.
[0079] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2G. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2G shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the source of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response which is the response.
[0080] Figure 2G shows an example of a series of conversations shown in Figure 2F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 2G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 2F, so repeated description will be omitted.
[0081] As shown in the natural language text of the server response in the table of FIG. 2F, by using the setting instruction sentence shown in FIG. 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 reflects the settings and conversation history of the character's role, name, conversation characteristics, personality, etc. at the time of the previous conversation. This is more preferable because it is recognized from the user's perspective that the identity of the settings and memories of the character's role, name, conversation characteristics, personality, etc. at the time of the previous conversation is more ensured. Note that this means that the characters can be viewed as the same from the user's perspective, so it may be called a pseudo-identity of the characters from the user's perspective.
[0082] Furthermore, from the user's perspective, memories can be shared with the character, resulting in a more enjoyable character conversation experience.
[0083] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2H. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, FIG. 2H shows an example of the operation in which a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input input to the operation input unit 1107 or an operation detected by a touch operation input sensor of the display unit 10011.
[0084] In the example of FIG. 2H, in addition to the character 19051 (named "Koto") used in the description of FIGS. 2A to 2G, a character 19052 (named "Tom") and a character 19053 (named "Necco") are shown. The character 19051 (named "Koto") and the character 19052 (named "Tom") are human characters, and the character 19053 (named "Necco") is a cat character. The display of the characters displayed on the display unit 10011 can be switched by switching and displaying on the display unit 10011 an image generated by rendering the character in a virtual 3D space that differs for each character. The process for realizing the display of the rendered image of the 3D model of each character can be, for example, any of the first to third process examples described in FIG. 15A. Also, depending on the character, a dynamic 2D image may be displayed.
[0085] Furthermore, it is preferable that the character operation program executed by control unit 1110 also changes the synthetic voice used for the "speech" of each character when switching the display of the characters displayed on display unit 10011. This can be achieved by storing synthetic voice data of a tone of voice associated with each character in advance in storage unit 1170, and performing synthetic voice change processing when switching the display of the character.
[0086] In the example of Fig. 2H, the user 230 is configured to be able to converse with any of the characters. In the AI response output device 10010 of Fig. 2H, different roles, names, conversation characteristics, or personalities are set for each of these characters. In addition, the memories of each character based on the conversation history are also managed as different for each character.
[0087] Therefore, the AI response output device 10010 constructs a database shown in FIG. 2I in the storage unit 1170, and uses this database to manage character settings and character conversation histories.
[0088] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2I. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.
[0089] A character operation program executed by the control unit 1110 of the AI response output device 10010 constructs the database 19200 in the storage unit 1170, for example. The character ID is an identification number for identifying each of the multiple characters that can be displayed on the AI response output device 10010, and may be a natural number or may use the alphabet, etc. The name is data of the name of each of the multiple characters that can be displayed on the AI response output device 10010.
[0090] The initial setting instruction text is text information in a natural language that explains the settings such as the role, name, conversation characteristics, or personality of each of multiple characters that can be displayed on the artificial intelligence response output device 10010. Since the initial setting instruction text is natural language text information that is the main data of the setting instruction text transmitted from the artificial intelligence response output device 10010 to the large-scale language model server 19001, it is desirable that the description content can be read as is by the large-scale language model of the artificial intelligence of the large-scale language model server 19001.
[0091] The conversation histories 1, 2, ... are records of conversations between the user and each character, and are recorded separately for each character. The conversation histories are to be included in the natural language text information, which is the main data of the setting instruction text sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable to make the contents readable as is by the large-scale language model of the AI of the large-scale language model server 19001.
[0092] When the character displayed on the display unit 10011 of the AI response output device 10010 is switched, the character operation program executed by the control unit 1110 of the AI response output device 10010 uses the database 19200 in Fig. 2I to select and switch the initial setting instruction sentence and conversation history used for the natural language text information, which is the main data of the setting instruction sentence transmitted from the AI response output device 10010 to the large-scale language model server 19001, so as to correspond to the character displayed on the display unit 10011 of the AI response output device 10010. In addition, every time a conversation takes place between the user 230 and the character, the character operation program records the conversation history in the database 19200 in Fig. 2I in the conversation history area corresponding to the character displayed on the display unit 10011.
[0093] By using the database 19200 in this way, the character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 establishes a conversation between the user 230 and the character using the character's speech using the response of the same artificial intelligence large-scale language model of the same large-scale language model server 19001. From the user's perspective, it is more preferable because it is recognized that the sameness of the settings and memories of the role, name, conversation characteristics, or personality of the character at the time of the previous conversation is more secured for each character. This may be expressed as a pseudo-identity of the characters as seen by the user.
[0094] Therefore, even when the artificial intelligence response output device 10010 is configured to switch the character to be displayed on the display unit 10011 from among a number of character candidates, the operation using the database 19200 described above reduces the sense of discomfort felt by the user when conversing with each character, and the user can share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.
[0095] If the initial setting instruction texts of multiple characters cannot be edited by the user, the settings of the role, name, conversation characteristics, personality, etc. of each character can be maintained close to the intention of the provider of the AI response output device 10010 or the creator of the content of the character. On the other hand, the initial setting instruction texts of the characters may be edited by the user in response to input from the operation input unit 1107 or the like. In this case, the settings of the role, name, conversation characteristics, personality, etc. of the character can be set to the user's preference, and the user can converse with the character that he or she has set up. In this case, the 3D model of the character, its rendering image, and the type of synthetic voice of the character may be replaced accordingly.
[0096] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to Fig. 2J. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service at a lower cost by the character conversation device using the artificial intelligence response output device 10010 and the character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 will be described.
[0097] As explained in FIG. 2B, it is very resource inefficient to train a large-scale language model with this level of artificial intelligence only for a specific purpose. Therefore, it is more resource efficient to perform large-scale training to generate a model as a foundation model that can be applied to various purposes, and use it on various terminals via an API (Application Programming Interface). In such cases, the provider of the large-scale language model often collects the cost used in training the large-scale language model from the terminal user as a usage fee for the terminal API. In natural language models, the API usage fee is often charged based on the number of processing of units of words that divide sentences, called tokens.
[0098] Therefore, in the artificial intelligence response output device 10010 of the second embodiment of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010 and a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.
[0099] For example, by using the processing and configurations of Examples 1 to 3 as shown in the table of Figure 2J, it is technically possible to reduce the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API.
[0100] Example 1 is an example of shortening the conversation history text and reducing the number of tokens using a document summarization process, among the methods for reducing the number of tokens in the conversation history text stored in the API setting instruction text and transmitted. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. Although the text summarization may be performed at the start of the next conversation, it is more time-saving to perform it at the end of a "series of conversations."
[0101] Furthermore, the text summarization process may be requested to the large-scale language model itself of the large-scale language model server 19001. However, in this case, the effect of saving the number of tokens is low. Therefore, for example, if the second server 19002 provides text summarization process in natural language via an API at a lower cost than the large-scale language model of the large-scale language model server 19001, the text summarization process may be requested to the second server 19002 via the API, and the text summary of the conversation history may be stored in a setting instruction statement for the large-scale language model server 19001 and transmitted.
[0102] Furthermore, if text summarization is the only process, it can be performed on the terminal side, and text summarization can be performed by the control unit 1110 executing a document summarization program loaded in the memory 1109 of the AI response output device 10010. In this case, the effect of saving the number of tokens is high. Even if the conversation history becomes long, if an upper limit on the number of characters after summary is specified in the text summarization process, the upper limit on the text length of the conversation history is determined, so that an upper limit on the tokens can be set and tokens can be saved.
[0103] In addition, since the text information of the character's initial settings, such as the character's role, name, conversational characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information of the character's initial setting instructions and reduce the number of tokens of the text information in the conversation history.
[0104] The process described in Example 1 may be performed by a character operation program executed by the control unit 1110 to control each unit.
[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored in the setting instruction text of the API and transmitted. For example, the number of tokens is reduced by deleting the older conversation history among the conversation histories with the characters recorded in the storage unit 1170. If an upper limit on the number of characters in the conversation history is specified, the upper limit on the length of the text in the conversation history is determined, so that the upper limit value of the tokens can be determined and tokens can be saved. Alternatively, a method may be used in which a predetermined period of the conversation history is specified and conversation history that exceeds the period is deleted. In this case, tokens can also be saved. Note that, in Example 2, text information in the character initial setting such as the character's role, name, conversation characteristics, or personality does not increase as much as the conversation history, so it is efficient and preferable to maintain the description of the text information in the initial setting instruction text of the character and reduce the number of tokens in the text information in the conversation history.
[0106] The processing described in Example 2 may be performed by a character operation program executed by the control unit 1110 controlling each unit.
[0107] Example 3 is a method for reducing the number of tokens by reducing the frequency of sending setting instruction sentences using an API. Specifically, even after the device is turned on or the display character is switched, or after the video setting and synthetic voice setting of the displayed character are completed, the setting instruction sentence is not sent in advance, and the setting instruction sentence is sent to the large-scale language model server 19001 only when the control unit 1110 determines that the natural language text information included in the user's voice picked up by the microphone 1139 is text information for which a large-scale language model of artificial intelligence should be used, thereby reducing the frequency of sending setting instruction sentences to the large-scale language model server 19001 and reducing the number of tokens.
[0108] Specifically, for example, after the device is powered on (ON) or after an operation input for switching the displayed character is made, a character 19051 (named "Koto") is displayed on the display unit 10011 as shown in FIG. 2H by a display process of the display unit 10011 controlled by a character operation program executed by the control unit 1110. At this time, for example, if a synthetic voice for the appearance corresponding to the character 19051 is stored and prepared in the storage unit 1170 or the like, a synthetic voice for the appearance such as "Good morning. I'm Koto," "Hello. I'm Koto," or "Good evening. I'm Koto" may be output from the speaker which is the audio output unit 1140. At this time, the image of the character 19051 is already set as the image of the character displayed on the display unit 10011, and the synthetic voice output from the speaker which is the audio output unit 1140 is set to the synthetic voice corresponding to the character 19051.
[0109] Here, the inference process of the large-scale language model of the artificial intelligence in the large-scale language model server 19001 described above also takes time if the instruction is long. In particular, when the setting instruction includes text information related to the past conversation history, the token amount of the instruction increases, and the inference process time is particularly long. The setting instruction itself and its response are not output to the user 230. From the response of the user instruction after the setting instruction, a synthetic voice as the "utterance" of the character is output from the speaker, which is the voice output unit 1140. In this case, it seems preferable at first glance to transmit the setting instruction from the artificial intelligence response output device 10010 to the large-scale language model server 19001 in advance and complete the inference process of the large-scale language model for the setting instruction in advance, because the response of the output of the synthetic voice of the "utterance" of the character 19051 after the user 230 speaks to the character 19051 is quicker.
[0110] However, if the setting instruction is transmitted to the large-scale language model server 19001 before the user 230 utters a word and the inference process of the large-scale language model for the setting instruction is completed in advance, for example, the user 230 may turn off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or the user 230 may switch the display character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. In this case, the number of tokens processed in the inference process of the large-scale language model by transmitting the setting instruction to the large-scale language model server 19001 in advance is the number of processed tokens that waste the usage fee. This is an obstacle to providing the character conversation device by the AI response output device 10010 and the character conversation service by the character conversation system by the AI response output device 10010 and the large-scale language model server 19001 at a lower cost to users.
[0111] Therefore, after the artificial intelligence response output device 10010 is powered on (ON) or after an operation to switch the displayed character is input, it is desirable to continue to not send a setting instruction sentence to the large-scale language model server 19001, even after the artificial intelligence response output device 10010 sets the image of character 19051 as the image of the character to be displayed on the display unit 10011 under the control of the character operation program executed by the control unit 1110 and sets the synthetic voice to be output from the speaker, which is the audio output unit 1140, to the synthetic voice corresponding to the character 19051, until the user 230 recognizes that he or she is speaking to the character 19051.
[0112] Here, the time when it is recognized that the user 230 is speaking to the character 19051 may be, for example, the time when the trigger keyword described in Fig. 2B is detected, or the time when the text of the words spoken by the user 230 is extracted. In this way, the number of processing tokens that waste the usage fee can be reduced, and the character conversation device by the artificial intelligence response output device 10010 and the character conversation system by the artificial intelligence response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.
[0113] Furthermore, even after the time when the user 230 recognizes that he / she is speaking to the character 19051, for example, if the text information extracted from the voice of the user 230 collected by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model, it is desirable to continue the state of not transmitting the setting instruction to the large-scale language model server 19001. Specifically, examples of the preset keywords include keywords such as "try jumping" and "try dancing" that request the user 230 to react to the character 19051 by an animation of the character 19051 moving or emitting a synthetic voice. In this case, the character operation program executed by the control unit 1110 reads out the motion data, animation video, and / or synthetic voice data corresponding to the reaction stored in the storage unit corresponding to the character 19051, and uses these data to generate the video to be displayed on the display unit 10011 and output the synthetic voice from the speaker that is the audio output unit 1140.
[0114] Such processing does not necessarily require inference processing of the large-scale language model of the large-scale language model server 19001. After the processing, if the user 230 turns off the power supply of the artificial intelligence response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or, for example, if the user 230 switches the display character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, if the setting instruction text is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be the number of processing tokens that wastes the usage fee.
[0115] Therefore, even after passing the point where it is recognized that the user 230 is speaking to the character 19051, it is desirable to continue the state where the setting instruction sentence is not transmitted to the large-scale language model server 19001 until it is determined whether or not the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Only when it is determined that inference processing of a large-scale language model is required, it is desirable to transmit the setting instruction sentence to the large-scale language model server 19001 and proceed with inference processing of the large-scale language model.
[0116] The processing described in Example 3 may be performed by a character operation program executed by the control unit 1110 controlling each unit.
[0117] According to the method for reducing (saving) the number of processed tokens in a large-scale language model using each example of Figure 2J described above, a character conversation device using the artificial intelligence response output device 10010, or a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001, can be provided to users at a lower cost.
[0118] Next, an example of the display of the character conversation device (artificial intelligence response output device 10010) of the second embodiment of the present invention will be described with reference to FIG. 2K. The example of FIG. 2K shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which has been described in each of FIGS. 2A to 2J, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063, which is a response from the large-scale language model, on the display unit 10011 together with an image of a character 19051. The text 10063, which is a response from the large-scale language model, may be displayed superimposed in front of the image of the character 19051 as shown in FIG. 2K. Also, the text 10063, which is a response from the large-scale language model, may be displayed together with the image of the character 19051 without being superimposed on the image of the character 19051.
[0119] The display in Figure 2K is an example, but for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by voice.
[0120] In this case, the control unit 1110 may control to start a display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, as shown in FIG. 2K. In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more suitably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051 may be manually switched ON / OFF by the user 230 operating via the operation input unit 1107 or the touch operation input sensor of the display unit 10011.
[0121] Next, an example of a fixed response phrase database (fixed response phrase DB) in a character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in FIG. 2H and FIG. 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition number and condition content are the same as those in FIG. 1C. In the example of FIG. 2L, individual fixed response phrases are set for each of the multiple characters for these conditions. For example, fixed response phrases for each condition are stored for each of the three characters, character 1: Koto, character 2: Tom, and character 3: Necco, as described in FIG. 2H and FIG. 2I. Output control of fixed response phrases is the same as that in FIG. 1C, and therefore will not be described again.
[0122] In the example of FIG. 2L, the control unit 1110 selects a corresponding fixed response phrase from the fixed response phrase database (fixed response phrase DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses it for output control as a response issued by the character. For example, in the example of the fixed response phrase database (fixed response phrase DB) in FIG. 2L, even if the conditions are the same, the fixed response phrases are changed to expressions or contents corresponding to the individuality of the character. This allows the character conversation device (artificial intelligence response output device 10010) to provide the user with a conversation corresponding to the individuality of the displayed character. The user can feel that each character has a more consistent individuality. This allows a character conversation device (artificial intelligence response output device 10010) to be realized that gives a greater sense of reality to multiple characters.
[0123] The fixed response phrase database (fixed response phrase DB) of FIG. 2L described above may be stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 may use it. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 2L may be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the artificial intelligence response output device 10010, instead of a response generated by a large-scale language model stored in each server. In this way, even if the artificial intelligence response output device 10010 does not have a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0124] According to the character conversation device and the character conversation system of the second embodiment described above, it is possible to reduce the sense of discomfort felt by the user from the conversation with the character displayed on the AI response output device 10010. Moreover, according to the character conversation device and the character conversation system of the second embodiment, it is possible to provide the character conversation service to the user at a lower cost.
[0125] In the above description of the second embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 19001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B, and the large-scale language model held by the local LLM processing unit 10028 may be used instead of the large-scale language model held by the large-scale language model server 19001. In this case, in the above description of the second embodiment, the large-scale language model held by the large-scale language model server 19001 may be read as the large-scale language model held by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).
[0126] In this case as well, it is possible to reduce the sense of discomfort felt by the user from the conversation with the character displayed on the artificial intelligence response output device 10010. When a large-scale language model held by the local LLM processing unit 10028 is used instead of the large-scale language model held by the large-scale language model server 19001, there is less need to consider the usage fee according to the number of processed tokens, but by reducing the number of processed tokens even in the case of a large-scale language model held by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power required for inference. In this case, it is possible to provide the user with a character conversation service that consumes less power.
[0127] In the above description of the second embodiment, an example has been described in which the conversation history with the character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). In contrast, the conversation history with the character may be recorded and stored in the second server 19002 or other cloud server connected to the Internet 19000. In this case, when a new conversation between a user and a character is started, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or other cloud server, acquires (downloads) the past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create an instruction for the large-scale language model. A specific method of using the past conversation history as an instruction for the large-scale language model is as described in each figure of the second embodiment, so a repeated description will be omitted.
[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character's conversation history up to a predetermined time point, such as each time a conversation between the user and the character occurs or when the conversation between the user and the character ends, to the above-mentioned second server 19002 or other cloud server. That is, the character conversation device (artificial intelligence response output device 10010) may upload the conversation history with the character to the second server 19002 or other cloud server at a predetermined timing, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or other cloud server and use it to generate an instruction sentence for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) that the user used the previous day and the character conversation device (artificial intelligence response output device 10010) that the user will use from now on are different individual devices and can display the same character.When the user has a conversation with the same character between the different individual devices multiple times at different times, it is possible to realize a conversation in which the character's memories are pseudo-carried over from the previous conversation, which is more convenient for the user.
[0129] The process described above in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with the character to the second server 19002 or other cloud servers to virtually take over the memory of the character is also effective when dealing with the database 19200 including the conversation history of multiple characters described in Fig. 2H and Fig. 2I. In other words, if the database 19200 described in Fig. 2I is configured to be uploaded and downloaded to the second server 19002 or other cloud servers, not only for one character but for multiple characters, when a user has multiple conversations with each of the multiple characters at different times between different individual devices, it is possible to realize a conversation in which the memory of each character is virtually taken over from the previous conversation, which is more convenient for the user.
[0130] <Example 3> Next, the third embodiment of the present invention is an improvement of the character conversation device (artificial intelligence response output device 10010) and the character conversation system described in each figure of the second embodiment. In this embodiment, the differences from the second embodiment will be described, and the repeated description of the same configurations as the second embodiment will be omitted.
[0131] As in the second embodiment, the character in the third embodiment can provide the user with the service of a large-scale language model, which is an artificial intelligence, and can be of assistance to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and the character conversation system in the present embodiment may be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, and an AI assistant response output system.
[0132] An example of a character conversation device and a character conversation system according to a third embodiment of the present invention will be described with reference to Fig. 3A. In the character conversation system of the third embodiment, a large-scale language model server 20001 is provided instead of the large-scale language model server 19001 in Fig. 2A, and is connected to the Internet 19000.
[0133] Here, the large-scale language model server 20001 is a server equipped with large-scale language model artificial intelligence, and is a multi-modal large-scale language model artificial intelligence capable of processing not only the natural language text information that could be processed by the large-scale language model server 19001, but also types of information other than natural language text information.
[0134] Moreover, the AI response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (AI response output device 10010) of the second embodiment, as an example.
[0135] In the third embodiment as well, the AI response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the Internet 19000 using an API.
[0136] The character conversation system of the third embodiment includes a mobile information processing terminal 20010 used by a user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.
[0137] 3B, an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011, which is a touch operation input panel, a control unit 20012, an external power supply input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.
[0138] The display panel 20011 is equipped with a touch operation input sensor and can receive touch operation input by the finger of the user 230. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel. The display panel 20011 may be called a display unit.
[0139] The communication unit 20020 may be configured with a Wi-Fi communication interface, a Bluetooth communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is provided with a control unit such as a CPU and a memory, and the control unit controls the display panel 20011 and the communication unit 20020. In addition, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000 by any of the communication methods of the communication unit 20020. This allows the mobile information processing terminal 20010 to communicate with various servers connected to the Internet 19000.
[0140] The power supply 20014 converts an AC current input from the outside via the external power supply input interface 20013 into a DC current, and supplies the required DC current to each unit of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Furthermore, the secondary battery 20015 supplies power to each unit requiring power via the external power supply input interface 20013 when power is not supplied from the outside.
[0141] The video signal input unit 20023 is connected to an external video output device and inputs video data. The video signal input unit 20023 may be implemented by various digital video input interfaces. For example, it may be implemented by a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may be implemented by various USB interfaces, etc.
[0142] The audio signal input unit 20024 is connected to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 20024 may be any of various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface in which a terminal and a cable are integrated.
[0143] The audio output unit 20021 can output audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 can also output audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured with a speaker. The audio output unit 20021 may also output built-in operation sounds and error warning sounds. Alternatively, the audio output unit 20021 may be configured to output a digital signal to an external device, like an Audio Return Channel function defined in the HDMI standard.
[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 20012, which will be described later, may perform voice recognition processing on the generated audio signal to obtain text information from the audio signal.
[0145] The imaging unit 20025 is a camera having an image sensor. The camera may be provided on the front side of the mobile information processing terminal 20010 facing the display panel 20011, or on the back side of the display panel 20011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.
[0146] The storage unit 20016 is a storage device that records various information such as various data including video data, image data, and audio data. The storage unit 20016 may be configured with a magnetic recording medium recording device such as a hard disk drive (HDD) or a semiconductor element memory such as a solid state drive (SSD). For example, various information such as various data including video data, image data, and audio data may be recorded in the storage unit 20016 in advance at the time of product shipment. The storage unit 20016 may also record various information such as various data including video data, image data, and audio data acquired from an external device, an external server, etc. via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may be output to an external device, an external server, etc. via the communication unit 20020.
[0147] The video control unit 20017 performs various controls related to the video signal input to the display panel 20011. The video control unit 20017 may be called a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 20017 may be called a video processing unit or an image processing unit. The video control unit 20017 performs control of video switching, such as which video signal is input to the display panel 20011 out of the video signal to be stored in the memory 20026 and the video signal (video data) input to the video signal input unit 20023. The video control unit 20017 may also perform control of image processing on the video signal input from the video signal input unit 20023 and the video signal to be stored in the memory 20026. Examples of image processing include scaling processing for enlarging, reducing, transforming, etc. an image, brightness adjustment processing for changing the brightness, contrast adjustment processing for changing the contrast curve of the image, and Retinex processing for decomposing an image into light components and changing the weighting of each component.
[0148] The attitude sensor 20018 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each unit connected thereto.
[0149] The non-volatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the non-volatile memory 20027 includes, for example, data for various operations to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data of objects to be operated by user operations, layout information, etc. The memory 20026 stores video data to be displayed on the display panel 20011, data for controlling the device, etc. The control unit 20012 may read out various software from the storage unit 20016 and expand and store it in the memory 20026.
[0150] The control unit 20012 controls the operation of each unit connected thereto. The control unit 20012 may also cooperate with a program stored in the memory 20026 to perform arithmetic processing based on information acquired from each unit in the mobile information processing terminal 20010.
[0151] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to Fig. 3C. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. In the third embodiment, the character conversation device (artificial intelligence response output device 10010) also expands a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing various processes described below.
[0152] In the second embodiment, the action performed by the user 230 with respect to the character conversation device (artificial intelligence response output device 10010) was mainly a call by the user 230's voice. In the character conversation device (artificial intelligence response output device 10010) of the second embodiment, a series of operations were performed starting from the process of collecting the voice of the user 230 with a microphone. In contrast, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the character conversation device (artificial intelligence response output device 10010) can also execute a series of operations performed starting from the process of collecting the voice of the user 230 with a microphone, as described in the second embodiment. In addition, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the user 230 can perform an action with respect to the character conversation device (artificial intelligence response output device 10010) by the user's operation via the operation input unit 1107 of FIG. 1B. Here, examples of the operation input unit 1107 of FIG. 1B include a mouse, a keyboard, and a touch panel.
[0153] In addition, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by a user touch operation that can be detected by the touch operation input sensor of the display unit 10011 of Figure 1B.
[0154] In addition, the user 230 can operate the mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting the user's 230 operation input to the character conversation device (artificial intelligence response output device 10010).
[0155] Also, an information storage image such as a two-dimensional code storing information that the user wants to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B of the character conversation device (artificial intelligence response output device 10010). The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as a two-dimensional code captured by the imaging unit 1180 to obtain the information. Also, an image that the user wants to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B of the character conversation device (artificial intelligence response output device 10010). The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180, and obtain the result of the image recognition processing.
[0156] In this way, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the types of actions that the user 230 can take against the character conversation device (artificial intelligence response output device 10010) are increased compared to the character conversation device (artificial intelligence response output device 10010) described in the second embodiment. As a result, the character conversation device (artificial intelligence response output device 10010) of the third embodiment can acquire the result of the action taken by the user 230 other than the user's voice, and generate an instruction sentence (prompt) to be transmitted to the large-scale language model server 20001 based on the result. As a result, the instruction sentence to be transmitted to the large-scale language model server 20001 can more suitably include information of a type other than the text information of the natural language extracted from the user's voice. The information of a type other than the text information of the natural language extracted from the user's voice is, for example, an image, a video, a sound, etc.
[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment transmits an instruction to the large-scale language model server 20001 using the API. In this embodiment, the instruction may be metadata or the like in which information described in a notation using tags such as a markup format of a markup language, a notation using predetermined symbols such as a Markdown format, or an object notation of a predetermined script such as JSON is stored. In this embodiment, the types of instruction include a setting instruction that stores an instruction such as an initial setting, and a user instruction that reflects an instruction from a user. Type identification information that identifies whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. At this time, the instruction includes text information in a natural language as the main message. Furthermore, in this embodiment, in addition to the text information in a natural language, the main message of the instruction may include a non-natural language information source such as an image, a video, or a sound as a type of information other than the text information in a natural language. A specific method of including a non-natural language information source in the instruction will be described later.
[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. The large-scale language model server 20001 receives an instruction sentence from a character conversation device (artificial intelligence response output device 10010). Based on the instruction sentence, the multimodal large-scale language model executes inference and generates a response including natural language text information that is the result of the inference. Here, since the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.
[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001, and extracts natural language text information stored as a main message in the response, and non-natural language information sources such as images, videos, or sounds. The character operation program of the character conversation device (artificial intelligence response output device 10010) may generate a natural language voice as a reply to the user using a voice synthesis technique based on the natural language text information extracted from the response, and output the voice from the voice output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.
[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display characters in a natural language that will be a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on text information in a natural language extracted from the above-mentioned response. At this time, the characters may be displayed together with the character 19051, may be displayed superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on information on the image of the non-natural language information source extracted from the above-mentioned response. At this time, the image may be displayed together with the character 19051, may be displayed superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video on the display screen of the character conversation device (artificial intelligence response output device 10010) based on the information of the video of the non-natural language information source extracted from the above-mentioned response in order to present the video to the user. At this time, the video may be displayed together with the character 19051, may be displayed superimposed on the video of the character 19051, or may be displayed in place of the video of the character 19051. These specific processes may be executed by the video control unit 1160.
[0163] In addition, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output a voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140, which is a speaker.
[0164] According to the character conversation device (artificial intelligence response output device 10010) of FIG. 3C or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 described above, it is not necessary to mount the large-scale language model itself, which requires a huge amount of data and computational resources for learning, in the character conversation device (artificial intelligence response output device 10010). In addition, it is possible to utilize the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model via the API. In response to an action from a user to a character, a response based on a non-natural language information source can be made in addition to a response based on a natural language text, making it possible to have a more suitable conversation.
[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to FIG. 3D. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3D shows an example of a non-natural language information source such as a natural language text of a main message of an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001 and an example of a non-natural language information source such as a natural language text of a main message of a server response which is a response thereto and an image. In this embodiment, the non-natural language information source can be an image, a video, a sound, etc., but FIG. 3D shows an example of an image as the non-natural language information source.
[0166] Also, Fig. 3D shows an exchange of instructions and responses in chronological order, from a setting instruction, a first round of user instructions and their responses, to a second round of user instructions and their responses. Here, the instructions and responses shown in Fig. 3D include a non-natural language information source 20061 and a non-natural language information source 20062, which were not shown in Fig. 2D of the second embodiment. In the example of Fig. 3D, both the non-natural language information source 20061 and the non-natural language information source 20062 are images.
[0167] Here, in Fig. 3D, for ease of explanation, an image of the non-natural language information source 20061 is shown pasted in the instruction text. However, there are a number of methods for transmitting data of the non-natural language information source 20061 or specifying data in an instruction text sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these methods or switch between them. An example of each method will be described below.
[0168] The first method for transmitting or specifying non-natural language information source data in an instruction sentence is used, for example, when the non-natural language information source to be specified is a non-natural language information source that exists in a location such as a server connected to a network such as the Internet. A specific example of the first method is a method in which a non-natural language information source file that exists on a network such as the Internet is specified by location information (such as a so-called URL) of the network such as the Internet and a file name using information such as tags and symbols in the instruction sentence.
[0169] For example, a tag that specifies an image in a markup language <img src=""****”"> By using the above, you can specify an image on a network such as the Internet by writing the location information and file name information of the image file in the **** part. Also, the tag that specifies a video in a markup language, etc. <video src=""****”">You can specify a video file on a network such as the Internet by writing the location information and file name information of the video file in the **** part using the tag. <audio src=""****”">and enter location information and file name information of the audio file in the **** portion to specify audio present on a network such as the Internet. In addition, in the case of JSON notation, an image present on a network such as the Internet may be specified by preparing a key such as img_src and entering location information and file name information of the image file in the value. For video files and audio files, the respective keys and values may be prepared. The specific example of this format is merely an example, and other unique formats may also be used. In either case, information specifying the location information and file name information of the non-natural language information source file may be stored in the directive.
[0170] When information specifying the location information and file name information of a non-natural language information source file is stored in a directive as in the first method, the directive itself does not need to store the data of the non-natural language information source file. Therefore, the amount of data in the directive can be reduced. In the first method, the large-scale language model server 20001 that receives a directive specifying non-natural language information source data can use the location information and file name information of the non-natural language information source file stored in the directive to acquire the non-natural language information source file located in a server or other location connected to a network such as the Internet.
[0171] Here, how to input location information and file name information when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence by the first method will be described. In FIG. 3C, it was explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are increased in addition to the voice of the user 230 compared to the second embodiment. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, etc., by user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 of FIG. 1B.
[0172] Also, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a WEB browser program and display the GUI of the WEB browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the WEB browser program may be accepted by a user operation (e.g., a mouse, a keyboard, a touch panel) via the operation input unit 1107 or a user's touch operation detectable by a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, a video, or a sound selected on the browser screen of the WEB browser program may be set as the data to be specified in the instruction sentence. In this case, the WEB browser program may obtain location information and file name information of the non-natural language information source data and pass them to the character operation program.
[0173] Also, the user 230 may operate the mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting location information such as a URL for specifying non-natural language information source data to the character conversation device (artificial intelligence response output device 10010). Also, location information such as a URL for specifying non-natural language information source data and file name information may be input in a manner in which an information storage image such as a two-dimensional code is displayed on the display panel 20011 of the mobile information processing terminal 20010, an image captured by the imaging unit 1180 of the character conversation device (artificial intelligence response output device 10010) is subjected to image recognition processing, and a result of the image recognition processing is acquired.
[0174] The use of the first method of transmitting or specifying non-natural language information source data in an instruction is not limited to the case where the non-natural language information source file is present in a location such as a server connected to a network such as the Internet in advance. For example, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) in an instruction, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to the second server 19002 via the Internet 19000, and include the location information (such as a so-called URL) and file name on the Internet of the non-natural language information source data of the uploaded second server 19002 in the instruction. In this case, the second server 19002 functions as a so-called intermediate server.
[0175] Similarly, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in the instruction, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit location information (such as a so-called URL) and a file name on the Internet of the non-natural language information source data of the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include the location information (such as a so-called URL) and a file name on the Internet of the non-natural language information source data uploaded to the second server 19002, which are acquired, in the instruction.
[0176] Furthermore, a character operation program of the character conversation device (artificial intelligence response output device 10010) may cooperate with the memory 1109 and the storage unit 1170 to construct a media server within the character conversation device (artificial intelligence response output device 10010) that can be accessed from other servers via the Internet 19000. In this case, when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction statement using the first method, the character conversation device (artificial intelligence response output device 10010) may store location information on the Internet (such as a so-called URL) indicating the inside of the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data in the instruction statement.
[0177] Next, the second method of specifying the transmission or designation of non-natural language information source data in the instruction sentence is, for example, a method of simply storing (attaching) the non-natural language information source data itself to the instruction sentence (prompt) and transmitting it. In general, non-natural language information source data such as images, videos, and sounds has a larger data amount than text information in natural language. Therefore, in this case, the data amount of the instruction sentence (prompt) itself is larger than that of the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) temporarily stores the non-natural language information source data to be stored (attached) in the instruction sentence (prompt) in the memory 1109, and when transmitting the instruction sentence (prompt), stores (attached) the non-natural language information source data in the memory 1109 via the communication unit 1132 and outputs it to the large-scale language model server 20001. The non-natural language information source data itself that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read out from the storage unit 1170 and stored in memory 1109.
[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data by instruction sentences.
[0179] Since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information, the first round of user instructions shown in the example of Figure 3D can obtain non-natural language information source 20061, i.e., an image of a swimming pool and poolside, and text information in natural language, and output the text information in natural language as shown in the figure as a response to the first round of user instructions as an inference result.
[0180] In addition, since the large-scale language model server 20001 is a multimodal large-scale language model capable of processing non-natural language information sources together with text information in natural language, as shown in the response to the second round of user instructions shown in the example of FIG. 3D, the large-scale language model server 20001 can include the non-natural language information source 20062 generated by inference of the multimodal large-scale language model in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). In FIG. 3D, the non-natural language information source 20062 shows an example of an image in which a circle image is added to the image of the swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in FIG. 3D, and may be a video or audio.
[0181] When the response from the large-scale language model server 20001 includes a non-natural language information source other than natural language text information, a method similar to the first method or the second method in which the above-mentioned character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence can be used.
[0182] Specifically, as a method similar to the above-mentioned first method, the large-scale language model server 20001 may store information specifying the location information and file name information of the non-natural language information source file in the instruction text in the response. The non-natural language information source 20062 such as an image, a video, or a sound may be stored in the large-scale language model server 20001 itself, or the non-natural language information source 20062 may be transferred to and stored in the second server 19002 functioning as an intermediate server. In either case, the large-scale language model server 20001 may store information specifying the location information and file name information of the non-natural language information source file in the instruction text in the response. The character conversation device (artificial intelligence response output device 10010) that has acquired the response may access the large-scale language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction text to acquire the non-natural language information source 20062.
[0183] Specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can obtain the data of the non-natural language information source 20062 stored (attached) in the instruction text and use it for various outputs to the user 230.
[0184] According to the operation of the character conversation device (artificial intelligence response output device 10010) and the character conversation system of the third embodiment described above with reference to Fig. 3D, instructions and responses are sent and received between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230 to realize a conversation using non-natural language information such as images, videos, and sounds. This makes it possible to realize a more advanced and natural conversation as shown in each message in Fig. 3D.
[0185] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to Fig. 3E. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3E shows an example of a main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and a main message of a server response which is the response.
[0186] FIG. 3E shows an example of a case where the user 230 talks to the character 19051 again to have a new conversation after the series of conversations shown in FIG. 3D has ended. In the example of FIG. 3E, the processing using the conversation history as described in FIG. 2F, FIG. 2G, and FIG. 2I of the second embodiment is not performed. Therefore, like FIG. 2E of the second embodiment, FIG. 3E shows a response with contents in a state where the name of the large-scale language model itself included in the setting instruction, the role to be played, the characteristics of the conversation, the name of the user, the conversation history, and the like are not remembered at all.
[0187] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to Fig. 3F. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3F shows an example of a main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and a main message of a server response which is the response.
[0188] FIG. 3F shows an example of a case where the user 230 talks to the character 19051 again to have a new conversation after the series of conversations shown in FIG. 3D has ended. Here, FIG. 3F shows a case where the method of storing a message explaining the history of past conversations in the setting instruction text, which was described in FIG. 2F in the second embodiment, is also applied to the character conversation device (artificial intelligence response output device 10010) in the third embodiment. Specifically, in FIG. 3F, the message that is the content of the setting instruction text in FIG. 3D is stored as a reset message, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.
[0189] Since the large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information, non-natural language information source data may be transmitted or specified in past instructions and responses. Therefore, in the example of FIG. 3F, the conversation history message reflects not only the natural language text information in past instructions and responses, but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method of transmitting or specifying non-natural language information source data in the instruction in FIG. 3F is the same as the transmission or specification of non-natural language information source data as described in FIG. 3D, so a repeated description will be omitted.
[0190] In the example of Fig. 3D, the method of transmitting or specifying the non-natural language information source data includes a case where the non-natural language information source data itself is stored (attached) in the instruction, and a case where the non-natural language information source data is not stored (attached) in the instruction. This point is also the same for the instruction in Fig. 3F.
[0191] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to Fig. 3G. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3G shows an example of a main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and a main message of a server response which is the response.
[0192] FIG. 3G shows an example of a series of conversations shown in FIG. 3F, from the first user instruction and its response to the third user instruction and its response, following the first setting instruction. FIG. 3G shows the exchange of instructions and responses in chronological order. The contents of the setting instruction are the same as those shown in FIG. 3F, so repeated description is omitted.
[0193] As described above, even when using the large-scale language model server 20001 having a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information of the third embodiment, even when the user 230 speaks to the character 19051 again to start a new conversation after a series of conversations has ended, if the generation process and transmission process of the setting instruction sentence in Fig. 3F are performed, the response of the user instruction sentence thereafter will reflect the settings and conversation history of the character's role, name, conversation characteristics, personality, and / or conversation characteristics at the time of the previous conversation, as shown in Fig. 3G. This is more preferable because it is recognized from the user's perspective that the identity of the settings and memories of the character's role, name, conversation characteristics, or personality at the time of the previous conversation is more assured.
[0194] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to FIG. 3H. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, FIG. 3H uses the example described in FIG. 2H of the second embodiment for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.
[0195] Also, database 20200 for managing character settings and character conversation history shown in Fig. 3H has the same format as database 19200 shown in Fig. 2I of the second embodiment, and Fig. 3H only describes the differences from database 19200 shown in Fig. 2I. Also, the content of the character "Koto" in the database will be explained, and the content of other characters will be omitted.
[0196] Here, as described above, the large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information, so that the instruction sentence from the character conversation device (artificial intelligence response output device 10010) and the response from the large-scale language model server 20001 contain not only natural language text information but also transmission or specification of non-natural language information source data. Therefore, in the database 20200 shown in FIG. 3H, in the data of the conversation history, not only the natural language text information contained in these instruction sentences and responses but also information on the transmission or specification of non-natural language information source data are recorded. The specific method of transmitting or specifying non-natural language information source data in the recording of the conversation history is the same as the transmission or specification of non-natural language information source data described in FIG. 3D, so repeated explanation will be omitted.
[0197] In the example of FIG. 3D, the method of transmitting or specifying non-natural language information source data includes a case where the non-natural language information source data itself is stored (attached) in the instruction text, and a case where the non-natural language information source data is not stored (attached) in the instruction text. This is also the case with the conversation history of FIG. 3H. However, in the conversation history of FIG. 3H, when the location information and file name information of a non-natural language information source file of a server (the second server 19002 functioning as an intermediate server or other cloud server) existing on a network such as the Internet is specified as a method of specifying non-natural language information source data, if the conversation history period is long, the non-natural language information source file on the server may be deleted. In this case, even if the location information and file name information are used, the non-natural language information source file cannot be obtained at a later date, and information in the conversation record may be lost.
[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts the instruction and response messages into a conversation history and records them, it may obtain the non-natural language information source file itself specified in the instruction and response from a server on the network using the location information and file name information, and store it in the storage unit 1170. Furthermore, the location information and file name of the non-natural language information source file may be rewritten by the character operation program of the character conversation device (artificial intelligence response output device 10010) into location information on the Internet (such as a so-called URL) indicating the inside of the media server of the media server constructed in the character conversation device (artificial intelligence response output device 10010), and then recorded in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself erases the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the information in the conversation record, which is more suitable for preserving the conversation record.
[0199] By using the database of Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between multiple character candidates to display on the display unit 10011, the user will feel less uncomfortable when conversing with each character, and will be able to share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience. This effect can be achieved as shown in Figure 2I of Example 2, and can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information.
[0200] In addition, in the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, the large-scale language model server 20001 uses a multimodal large-scale language model artificial intelligence that is capable of processing not only natural language text information but also non-natural language information other than natural language text information.
[0201] Here, an API is used for communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001. In a multimodal large-scale language model, the API usage fee may be charged according to the amount of data from a non-natural language information source in addition to the number of processing times for natural language text information in units of words that divide sentences, called tokens.
[0202] Therefore, in order to provide the character conversation service by the character conversation system according to this embodiment to the user at a lower cost, the following modified example may be used.
[0203] As a first modification, in the record of the conversation history of the database in FIG. 3H, the transmission or specification information of the non-natural language information source data is also recorded. However, the character and the user exchange conversations about the natural language information source data in text information in natural language, and the contents are recorded in text information in natural language. Then, even if the recording of the transmission or specification information of the natural language information source data is omitted in the record of the conversation history of the database in FIG. 3H, the conversation itself about the natural language information source data will be recorded as text information in natural language to some extent. Therefore, if a certain degree of information reduction is allowed, the recording of the transmission or specification information of the natural language information source data may be omitted in the record of the conversation history of the database in FIG. 3H. In this case, the transmission or specification information of the natural language information source data is also omitted from the conversation history message of the setting instruction sentence in FIG. 3F. This makes it possible to reduce the amount of data of the non-natural language information source communicated using the API.
[0204] Next, as a second modified example, in the recording of the conversation history of the database in FIG. 3H, instead of transmitting the non-natural language information source data or recording the specified information, natural language text information explaining the content of the non-natural language information source data is recorded. The natural language text information explaining the content of the non-natural language information source data may be acquired, for example, by starting a conversation between the large-scale language model of the large-scale language model server 20001 and a character conversation device (artificial intelligence response output device 10010) separately from the conversation as a character, and having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit. Also, the content of the non-natural language information source data may be acquired by having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit through a conversation with another large-scale language model of another server that can be used more inexpensively than the large-scale language model of the large-scale language model server 20001. Also, when alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be the natural language text information explaining the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is a tag of a markup language. <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">This is the text information written in the **** part of the text.
[0205] Furthermore, in the case of JSON format notation, in an object stored in correspondence with the location information and file name information of the non-natural language information source data, which are keys and values indicating the location information of the non-natural language information source data, a key corresponding to the alternative text can be further linked and stored as a value which is the alternative text data itself.
[0206] In this case, too, the transmission of the natural language information source data or the recording of the designated information can be omitted in the conversation history record of the database in Fig. 3H, and the transmission of the natural language information source data or the designated information is also omitted from the conversation history message of the setting instruction sentence in Fig. 3F, thereby reducing the amount of data of the non-natural language information source communicated using the API.
[0207] Next, as a third modified example, at the time of the first round of the user instruction in FIG. 3D, the information of the transmission or specification of the non-natural language information source data is not stored in the user instruction, but is replaced with text information in natural language explaining the contents of the non-natural language information source data. For example, in the first round of the user instruction in FIG. 3D, the information of the transmission or specification of the non-natural language information source data 20061 may be replaced with the user instruction, and an explanatory text such as "This image is an image of a swimming pool and a seat and a parasol by the pool. There is water in the swimming pool. There are drinks on the table beside the seat" may be stored as text information in natural language. In this case, the explanatory text may be obtained by having the content of the non-natural language information source data explained by specifying a predetermined character limit through a conversation with another large-scale language model of another server that can be used more cheaply than the large-scale language model of the large-scale language model server 20001. The explanatory text may also be obtained from a server of various other services that can obtain an overview or an explanation of the contents of non-natural language information source data such as images, videos, and audio. Furthermore, when alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be natural language text information explaining the contents of the non-natural language information source data.
[0208] Next, an example of a display example of the character conversation device (artificial intelligence response output device 10010) of the third embodiment of the present invention will be described with reference to FIG. 3I. In the example of FIG. 3I, an example of an example of displaying a response from a large-scale language model to an instruction sentence from a user described in each of FIG. 3A to FIG. 3H on a display unit 10011 of the character conversation device (artificial intelligence response output device 10010) is shown. Specifically, it is an example of displaying text 10063 of natural language information source data, image 10064 of non-natural language information source data, and / or video 10065 of non-natural language information source data, which are responses from the large-scale language model, together with a video of the character 19051 on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed superimposed in front of the video of the character 19051 as shown in FIG. 3I.
[0209] Also, the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of the character 19051 without being superimposed on the video of the character 19051. The display in FIG. 3I is an example, but for example, when the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to the minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 cannot confirm the response from the large-scale language model by audio. Therefore, in this case, the control unit 1110 may control to start a display mode in which the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of the character 19051, as shown in FIG. 3I.
[0210] In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more preferably use the character conversation device (artificial intelligence response output device 10010). Note that the user 230 may be configured to manually switch ON / OFF a display mode in which the text 10063, image 10064, and / or video 10065, which is a response from the large-scale language model, is displayed together with the image of the character 19051 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. According to the display example of FIG. 3I, it becomes possible to more preferably output a response from a large-scale language model in the character conversation device (artificial intelligence response output device 10010) that supports multimodal.
[0211] According to the character conversation device and the character conversation system of the embodiment 3 described above, in addition to the effects of the character conversation device and the character conversation system of the embodiment 2, it is possible to provide the user with a more advanced conversation experience including non-natural language information in addition to natural language information by using a multimodal large-scale language model. Moreover, according to the character conversation device and the character conversation system of the embodiment 3, it is possible to provide the user with a character conversation service at a lower cost.
[0212] In the above description of the third embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 20001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may include a local LLM processing unit 10028 shown in Fig. 1B, and may use a multimodal large-scale language model held by the local LLM processing unit 10028. In this case, the multimodal large-scale language model held by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model held by the large-scale language model server 20001.
[0213] In this case, in the above description of the third embodiment, the multimodal large-scale language model held by the large-scale language model server 20001 may be replaced with the multimodal large-scale language model held by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010). In this case, the multimodal large-scale language model can be used to provide the user with a more advanced conversation experience that includes non-natural language information in addition to natural language information. Note that, when the multimodal large-scale language model held by the local LLM processing unit 10028 is used instead of the multimodal large-scale language model held by the large-scale language model server 20001, there is less need to consider the usage fee according to the number of processed tokens and the amount of data of the non-natural language information source. However, even with the multimodal large-scale language model held by the local LLM processing unit 10028, the number of processed tokens and the amount of data of the non-natural language information source can be reduced, thereby reducing the consumption of resources such as power required for inference. In this case, a character conversation service with less power consumption can be provided to the user.
[0214] The configuration described in the second embodiment for uploading and downloading the conversation history with the character and the database data including the conversation history with the character to the second server 19002 or other cloud servers can also be used in the example using the multimodal large-scale language model described in the third embodiment. In this case, when a user has multiple conversations with each of the multiple characters at different times between different individual devices for one character or multiple characters, it is possible to realize a conversation in which the memory of each character is pseudo-carried over from the previous conversation, which is more convenient for the user.
[0215] <Example 4> Next, the fourth embodiment of the present invention is an improvement of the artificial intelligence response output device 10010, the character conversation device, or these systems described in the drawings of the second embodiment or the third embodiment. In this embodiment, the differences from the second embodiment or the third embodiment will be described, and the repeated description of the same configuration as these embodiments will be omitted.
[0216] As in the above-mentioned embodiment, the AI response output device 10010 may be called an AI response output device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be called an AI response output system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0217] An example of an operation using a database in a character conversation device (artificial intelligence response output device 10010) according to a fourth embodiment of the present invention will be described with reference to Fig. 4A. The database according to the fourth embodiment shown in Fig. 4A is an extension of the database described in Fig. 2I or Fig. 3I. Specifically, the database shown in Fig. 4A assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores in the database initial setting instruction statements and conversation histories corresponding to each user and character.
[0218] In the example of Fig. 4A, for user 1 with a user ID of 1, the initial setting command statements and conversation history of each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, the initial setting command statements and conversation history of each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored.
[0219] These initial setting instruction sentences and conversation history data are stored as separate data in different areas for each combination of user and character. For the sake of explanation, in Fig. 4A, the data stored in each area is indicated as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses the initial setting instruction sentences and conversation history stored in different areas for each combination of user and character based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or the system, thereby making it possible to more appropriately maintain the consistency of the character's personality and the continuity of memory for each different user.
[0220] Specifically, consider a situation in which user 1 has previously had a conversation with character Tom using a character conversation device (AI response output device 10010), and user 2 is unaware of that conversation, and then user 2 converses with character Tom. In this case, if the AI response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify the user, the response output from the AI response output device 10010 will be based on a conversation history that user 2 does not remember, and there is a possibility that the conversation between user 2 and the character of the AI response output device 10010 will become inconsistent.
[0221] In contrast, even in a similar situation, if the database shown in Fig. 4A is used, the control unit 1110 of the character conversation device (AI response output device 10010) identifies the user by ID, stores the initial setting instruction sentence and conversation history in a different area for each user, and uses the initial setting instruction sentence and conversation history stored in the different area for each user to generate an AI response. As a result, the initial setting instruction sentence and conversation history used to generate an AI response for each user are based on the operation or conversation history of that user, and are managed separately from the operation or conversation history of other users. This makes it possible to more appropriately maintain the consistency of the conversation history between each user and each character of the AI response output device 10010.
[0222] The database of the initial setting instruction and / or the conversation history described in FIG. 4A may be stored in the storage unit 1170 of the AI response output device 10010 and used by the control unit 1110. Also, without being limited to this, the database of the initial setting instruction and / or the conversation history may be stored in a server on the network. For example, when the AI response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 in generating the AI response, the database of the initial setting instruction and / or the conversation history described in FIG. 4A may be stored in these servers themselves. In this way, the process of including the initial setting instruction and the conversation history in the instruction again and sending it from the AI response output device 10010 to these servers can be omitted, and the number of transmission tokens for the use of the large-scale language model can be saved.
[0223] When storing the database of the initial setting instruction sentence and / or the conversation history described in FIG. 4A, the AI response output device 10010 may transmit the user ID, the character ID, and the user instruction sentence for the subsequent conversation to these servers. The large-scale language models on these servers may obtain the corresponding initial setting instruction sentence and the conversation history from the database of the initial setting instruction sentence and / or the conversation history of FIG. 4A using the user ID and the character ID obtained from the AI response output device 10010. The large-scale language models on these servers may perform inference using the initial setting instruction sentence and the conversation history, and the user instruction sentence for the subsequent conversation transmitted from the AI response output device 10010, generate an AI response, and transmit it to the AI response output device 10010. In this way, the effect of maintaining the consistency of the character's personality and the continuity of memory for each different user more appropriately can be obtained while saving the number of transmitted tokens for the use of the large-scale language model.
[0224] Next, an example of an operation using a database in the character conversation device (artificial intelligence response output device 10010) of the fourth embodiment of the present invention will be described with reference to Fig. 4B. The database according to the fourth embodiment shown in Fig. 4B is an extension of the database described in Fig. 1C or Fig. 2L. Specifically, the database shown in Fig. 4B assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data of standard response phrases corresponding to each user and character in the database.
[0225] In the example of Fig. 4B, for user 1 with a user ID of 1, response template data for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 is stored. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, response template data for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 is stored.
[0226] These standard response phrase data are stored as separate data in different areas for each combination of user and character. For the sake of explanation, in FIG. 4B, the data stored in each area is represented as standard response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, the standard response phrase data 101 is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 of character 1: Koto shown in FIG. 2L. The data 201 in FIG. 4B is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 of character 2: Tom shown in FIG. 2L.
[0227] Data 301 in FIG. 4B is stored as a database such as a table corresponding to the standard response phrases corresponding to condition numbers 1 to 7 of character 3: Necco shown in FIG. 2L. Data 102, 202, and 302 in FIG. 4B store standard response phrases in a similar format that have been modified for user 2. Data 103, 203, and 303 in FIG. 4B store standard response phrases in a similar format that have been modified for user 3. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses standard response phrase data stored in different areas for each combination of user and character, based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.
[0228] In this way, even if the character is the same, it is possible to respond with different standard response phrases for each user. That is, even if the character is the same, it may be more suitable to change the content of the standard response phrase depending on the relationship between the character and the user. For example, depending on the relationship between the age of the character and the age of the user registered in the AI response output device 10010 or the system, the user may be older than the character, the same age, or younger. In this case, the conversation between the user and the character will be more suitable or more natural if the content of the standard response phrase of the character for older users, the standard response phrase of the character for users of the same age, and the standard response phrase of the character for younger users are each different. That is, by performing the operation using the database of FIG. 4B, it is possible to produce a more suitable or more natural conversation by changing the content of the standard response phrase for each relationship between the character and the user.
[0229] The fixed response phrase database (fixed response phrase DB) of FIG. 4B described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the artificial intelligence response output device 10010. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 4B may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the artificial intelligence response output device 10010, instead of a response generated by a large-scale language model stored in each server. In this way, even if the artificial intelligence response output device 10010 is not provided with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0230] According to the character conversation device and the character conversation system of the fourth embodiment described above, it is possible to produce a more suitable or more natural conversation depending on the relationship between the character and the user, the conversation history, and the like.
[0231] <Example 5> Next, the fifth embodiment of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of the first, second, and third embodiments. Specifically, this is an example of switching the response generation process of the AI response output device 10010 from a response generation process using a large-scale language model on a network to a response generation process using a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, or a response generation process using a response template database. In this embodiment, differences from these embodiments will be described, and repeated explanations of configurations similar to those of these embodiments will be omitted.
[0232] As in the above-mentioned embodiment, the AI response output device 10010 may be called an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be called an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0233] An example of the switching process of the response generation process in the AI response output device 1001 of the fifth embodiment of the present invention will be described with reference to FIG. 5A. The table in FIG. 5A shows examples 1 to 9 of the switching process of the response generation process in the AI response output device 1001. In the table in FIG. 5A, the column "switching overview" shows an overview of the switching process of each example. The column "state before switching LLM (API connected LLM) on the network" shows a state before the response generation process by the large-scale language model on the network (large-scale language model connected using API) such as the large-scale language model provided in the large-scale language model server 19001 in FIG. 1 and the multimodal large-scale language model provided in the large-scale language model server 20001 is switched to another response generation process. The column "switching occurrence condition" shows a condition under which the switching process of the response generation process occurs. The column "switching destination from LLM on network (API-connected LLM)" shows the switching destination to which the response generation process of the AI response output device 10010 is switched from the large-scale language model on the network (large-scale language model connected using API), such as the large-scale language model provided in the large-scale language model server 19001 and the multimodal large-scale language model provided in the large-scale language model server 20001. The control unit 1110 of the AI response output device 1001 may perform control to switch to the large-scale language model, database, or correspondence shown in "switching destination from LLM on network (API-connected LLM)" when the condition shown in "switching occurrence condition" occurs in the state of "state before switching of LLM on network (API-connected LLM)" shown in FIG. 5A.
[0234] Below, each example shown in the table of FIG. 5A will be described. Example 1 is an example in which switching is performed according to the network connection status of the artificial intelligence response output device 1001, as shown in "Switching Overview". In Example 1, the "state before switching of LLM (API connection LLM) on the network" indicates that the network connection status of the artificial intelligence response output device 1001 is a connectable state. Here, in Example 1, the "switching occurrence condition" indicates "when network connection becomes impossible". That is, this means a case in which the connection via the network between the artificial intelligence response output device 1001 and the large-scale language model on the network (large-scale language model connected using API) becomes impossible. Specifically, the connection failure may be caused by a communication failure state on the connection path from the artificial intelligence response output device 1001 to the Internet 19000. Alternatively, the connection failure may be caused by a communication failure state on the Internet 19000. Alternatively, the connection failure may be caused by a situation in which the large-scale language model on the network (large-scale language model connected using API) itself cannot connect to the Internet 19000. Also, in Example 1, "local LLM" is shown as the "switching destination from LLM on the network (API-connected LLM)". This specifically means that switching processing to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010 is performed. That is, in Example 1, even if connection to a large-scale language model on the network (large-scale language model connected using an API) becomes impossible for some reason and response generation processing by a large-scale language model on the network (large-scale language model connected using an API) cannot be used, the response generation processing is switched to the local LLM processing unit 10028 of the AI response output device 10010. As a result, it is possible to continue response generation processing using the large-scale language model, even if there is a performance difference as a large-scale language model.
[0235] Next, Example 2 of FIG. 5A will be described. Example 2 is an example in which the "switching destination from the LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "predefined response phrase DB (database)". The response generation process using the "predefined response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 2, when a connection to a large-scale language model on the network (large-scale language model connected using an API) becomes impossible for some reason and the response generation process using the large-scale language model on the network (large-scale language model connected using an API) cannot be used, it is possible to generate a response by simpler processing and output the response to the user by switching to a response generation process using a predefined response phrase database.
[0236] Next, Example 3 of FIG. 5A will be described. Example 3 is an example in which the "switching destination from LLM on the network (API connection LLM)" in Example 1 is changed from "local LLM" to "non-response handling". The "non-response handling" means that even if there is a user input from the user via the touch panel, microphone 1139, or operation input unit 1107 requesting a response from a large-scale language model, no response to this input is generated, or even if there is a user input requesting a response from a large-scale language model, no response to this is output. That is, Example 3 makes it possible to more easily handle a case in which connection to a large-scale language model on the network (large-scale language model connected using an API) becomes impossible for some reason, and response generation processing by a large-scale language model on the network (large-scale language model connected using an API) cannot be used.
[0237] Next, Example 4 of FIG. 5A will be described. Example 4 is an example in which switching is performed due to a response delay of the LLM on the network, as shown in "Switching Overview". In Example 4, the "state before switching of LLM on the network (API connection LLM)" shows a state in which a response from the LLM on the network is obtained within a predetermined time. Here, in Example 4, the "switching occurrence condition" shows a case in which a response from the LLM on the network is not obtained within a predetermined time and exceeds the predetermined time. Also, in Example 4, "local LLM" is shown as "switching destination from LLM on the network (API connection LLM)". The "local LLM" of the switching destination is the same as in Example 1, so repeated explanation will be omitted. That is, in Example 4, even if the response from the LLM on the network (large-scale language model connected using API) exceeds a predetermined time for some reason and the response generation process by the LLM on the network (large-scale language model connected using API) cannot be used smoothly, the response generation process is switched to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This makes it possible to continue response generation processing using a large-scale language model, even if there are differences in performance as a large-scale language model.
[0238] Next, Example 5 of FIG. 5A will be described. Example 2 is an example in which the "switching destination from the LLM on the network (API-connected LLM)" in Example 4 is changed from "local LLM" to "predefined response phrase DB (database)". The response generation process using the "predefined response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 5, if a response from the LLM on the network (large-scale language model connected using an API) exceeds a predetermined time for some reason and the response generation process using the LLM on the network (large-scale language model connected using an API) cannot be used smoothly, it is possible to generate a response by simpler processing and output the response to the user by switching to a response generation process using a predefined response phrase database.
[0239] Next, examples 6 to 9 of FIG. 5A will be described. As shown in "Switching Overview", examples 6 to 9 are examples in which switching is performed when the API usage or usage fee reaches the upper limit. Here, as described in the second embodiment, the provider of the large-scale language model often collects the cost used in learning the large-scale language model from the user of the terminal as the API usage fee for the terminal. In this case, in the natural language model, the API usage fee is often charged based on the number of processing of units of words that divide sentences, called tokens. Here, various methods of charging and limiting the API usage fee are conceivable. As one of such ideas, an example can be considered in which the upper limit of the amount of the service of using the large-scale language model that a user can receive in a normal state is specified using the number of tokens processed.
[0240] In this case, the user may be able to receive services using large-scale language models at a specified API usage fee until the usage volume (or corresponding usage fee) is reached, and once the upper limit of the usage volume (or corresponding usage fee) is reached, certain restrictions may be imposed, such as the user being unable to receive services using large-scale language models in their normal state (performance or frequency).
[0241] Examples 6 to 9 in FIG. 5A are examples of switching control of the response generation process by the control unit 1110 of the artificial intelligence response output device 1001 when such restrictions occur in a service using a large-scale language model. Specifically, in Example 6, the "state before switching LLM on the network (API-connected LLM)" is a state in which the API usage amount and API usage fee have not reached a predetermined upper limit. This means a state in which the usage amount of the LLM on the network (API-connected LLM) has not reached a predetermined upper limit. At this time, the user can use the LLM on the network (API-connected LLM) in a normal state.
[0242] Here, in Example 6, the "switching occurrence condition" is when the API usage amount or API usage fee reaches a predetermined upper limit. This means that the usage amount of the LLM on the network (API connection LLM) reaches a predetermined upper limit. Also, in Example 6, the "switching destination from the LLM on the network (API connection LLM)" is a second LLM on the network different from the LLM (which may be called the first LLM) used in the normal state. An example of the second LLM on the network is an LLM with a lower fee than the first LLM used in the normal state. Since it is a lower-fee service, the performance of the second LLM is considered to be lower than that of the first LLM. Even in this case, there is a sufficient advantage if a large-scale language model can be used cheaply even after the usage amount / fee limit of the first LLM is reached.
[0243] Next, Example 7 in Fig. 5A will be described. In Example 7, the "switching destination from the LLM on the network (API-connected LLM)" in Example 6 is changed from a second LLM on a network different from the LLM (which may be called the first LLM) used in the normal state to a "local LLM." In Example 7, even if the usage amount of the API or the usage fee of the API reaches a predetermined upper limit, that is, even if the usage amount of the LLM on the network (API-connected LLM) reaches a predetermined upper limit, it is possible to continue the response generation processing using a large-scale language model by switching to the response generation processing using the local LLM, which is not restricted by the usage amount of the LLM on the network, the usage amount of the API, or the usage fee of the API.
[0244] Next, Example 8 of FIG. 5A will be described. Example 8 is an example in which the "switching destination from the LLM on the network (API connection LLM)" in Example 7 is changed from "local LLM" to "predefined response phrase DB (database)". The response generation process using the "predefined response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so repeated description will be omitted. In Example 8, even if the usage amount of the API or the usage fee of the API reaches a predetermined upper limit, that is, even if the usage amount of the LLM on the network (API connection LLM) reaches a predetermined upper limit, the process switches to a response generation process using a predefined response phrase database that is not limited by the usage amount of the LLM on the network, the usage amount of the API, or the usage fee of the API. This makes it possible to generate a response by simpler processing and output the response to the user.
[0245] Next, Example 9 in FIG. 5A will be described. Example 9 is an example in which the "switching destination from LLM on the network (API-connected LLM)" in Example 7 is changed from "local LLM" to "non-response response." The "non-response response" refers to a response in which a response to the user is not generated or a response to the user is not output. Example 9 makes it easier to respond to a case in which the API usage or API usage fee has reached a predetermined upper limit, that is, the usage of the LLM on the network (API-connected LLM) has reached a predetermined upper limit, and thus response generation processing by a large-scale language model on the network (large-scale language model connected using an API) cannot be used.
[0246] According to the switching control of the response generation process of the artificial intelligence response output device 10010 shown in Examples 1 to 9 of Figure 5A described above, even in a situation where the response generation process using an LLM (large-scale language model connected using an API) on the network cannot be used as usual, more appropriate switching or response can be performed according to each situation.
[0247] Note that the switching controls of Examples 1 to 9 in Fig. 5A may be performed by combining a plurality of examples. For example, the switching controls of Examples 1 to 3 may be combined with any of the controls of Examples 4 to 9. Similarly, the control of Example 4 or Example 5 may be combined with any of the controls of Examples 1 to 3 or Examples 6 to 9. Similarly, the control of Examples 6 to 9 may be combined with any of the controls of Examples 1 to 5.
[0248] Next, an example of a display of an AI assistant or a character when the artificial intelligence response output device 10010 of the fifth embodiment is configured as an AI assistant device or a character conversation device will be described with reference to Figs. 5B to 5D.
[0249] First, Fig. 5B is a display example of an AI assistant or character in the AI response output device 10010 when performing the switching control of Example 3 of Fig. 5A. In the example of Fig. 5B, the display state of the AI assistant or character is changed depending on whether the network connection state of the AI response output device 10010 is network connectable or network unconnectable. The network connectable and network unconnectable states of the AI response output device 10010 are as described in Fig. 5A, so a repeated description will not be given.
[0250] In the example of FIG. 5B, the AI response output device 10010 (1) displays the AI assistant or character in a normal awake state when a network connection is possible, but (2) displays the AI assistant or character in a "sleeping" state when a network connection is not possible. In the switching control of example 3 of FIG. 5A, when the AI response output device 10010 cannot connect to the network, it does not generate a response even if an instruction sentence is input from the user, or does not output a response. In this case, if the AI assistant or character displayed by the AI response output device 10010 is in a normal awake state, the user feels uncomfortable, but if the AI assistant or character displayed by the AI response output device 10010 is displayed in a sleeping state, the user can understand that "the AI assistant or character is not responding because it is sleeping," and it is possible to further reduce the discomfort felt by the user.
[0251] In the case of Fig. 5B(2), it is desirable that the user understands that "the AI assistant or character is not responding because it is asleep" before making a user input requesting a response by a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 1001. Therefore, it is desirable that the start timing of the state of displaying the AI assistant or character in the "sleeping" state when network connection is not possible in Fig. 5B(2) is immediately after the control unit 1110 of the artificial intelligence response output device 1001 determines that network connection is not possible, before the user input requesting a response by a large-scale language model.
[0252] Next, as another display example, the display example of FIG. 5C will be described. The display example of FIG. 5C is an example in which the display state of the AI assistant or character is changed according to the state of "switching destination from LLM on the network (API connection LLM)" in the table in the switching control of FIG. 5A. Specifically, FIG. 5C shows (1) a display example of the AI assistant or character in a state in which the AI response output device 10010 can connect to a large-scale language model on the network (large-scale language model connected using API) and can use a response generation process by the large-scale language model on the network (referred to as a normal state in this figure), (2) a display example of the AI assistant or character in a state in which the AI response output device 10010 switches to a response generation process by an LLM or a response template database with lower performance than the large-scale language model on the network (large-scale language model connected using API), and (3) a display example of the AI assistant or character in a state in which the AI response output device 10010 switches to the no-response response described in FIG. 5A.
[0253] In the example of FIG. 5C, for example, (1) when the artificial intelligence response output device 10010 is in a "normal state", the artificial intelligence response output device 10010 displays the AI assistant or character in a state where there is no particular problem. Note that the "normal state" in FIG. 5C may be considered to be a state other than the states (2) and (3). Also, for example, in a state (2) where the artificial intelligence response output device 10010 switches to a response generation process using an LLM or a response template database with lower performance than a large-scale language model on the network (a large-scale language model connected using an API), the artificial intelligence response output device 10010 displays the AI assistant or character in a "sleepy" state. Note that "displaying the AI assistant or character in a "sleepy" state" may be expressed as "a display showing that the AI assistant or character is feeling sleepy".
[0254] The response generation process of (2) has lower performance than the response generation process by a large-scale language model on the network (large-scale language model connected using an API) in the normal state of (1). Therefore, by displaying the AI assistant or character in a "sleepy" state, it is possible to implicitly inform the user that the response performance of the AI assistant or character is low. This makes it possible to further reduce the discomfort felt by the user in response to a low-performance response. Note that the switching conditions under which the artificial intelligence response output device 10010 switches to the response generation process by an LLM or a response template database, which has lower performance than the large-scale language model on the network (large-scale language model connected using an API), are as described in FIG. 5A, so repeated explanation will be omitted.
[0255] In the case of Fig. 5C (2), it is desirable to inform the user that the response performance of the AI assistant or character is low before the user inputs a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 1001. Therefore, it is desirable that the timing of starting the state of displaying the AI assistant or character in a "sleepy" state in Fig. 5C (2) is immediately after the artificial intelligence response output device 10010 switches to a response generation process using an LLM or a response template database with lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user inputs a response from a large-scale language model.
[0256] Also, for example, (3) when the AI response output device 10010 switches to the no-response response described in FIG. 5A, the AI response output device 10010 displays the AI assistant or character in a "sleeping" state. As described in FIG. 5B, by displaying the AI assistant or character displayed by the AI response output device 10010 in a "sleeping" state, the user can understand that "the AI assistant or character is not responding because it is sleeping," and it is possible to further reduce the sense of discomfort felt by the user. Note that the switching occurrence conditions for the AI response output device 10010 to switch to the no-response response described in FIG. 5A are as described in Example 3 or Example 9 of FIG. 5A, and therefore repeated explanations will be omitted. Note that in the case of FIG. 5C(3), it is desirable for the user to understand that "the AI assistant or character is not responding because it is sleeping" before making a user input requesting a response by a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 1001. Therefore, it is desirable that the start timing of the state (3) in Figure 5C in which the AI assistant or character is displayed in a "sleeping" state be immediately after the artificial intelligence response output device 10010 switches to the no-response response described in Figure 5A, prior to any user input requesting a response from a large-scale language model.
[0257] In the display example of FIG. 5C, the AI response output device 10010 displays a technical explanation of the state of the AI response output device 10010 related to the response generation process to the user, without directly explaining the state of the AI assistant or the character, which is implicitly reflected as a change in the state of the AI assistant or the character. This can reduce the sense of discomfort felt by the user more than when the technical explanation of the state of the AI response output device 10010 related to the response generation process is directly given to the user. In addition, it can reduce the sense of discomfort felt by the user more than when the display state of the AI assistant or the character remains the same as the normal state despite the change in the state of the AI response output device 10010 related to the response generation process.
[0258] However, some users may want to know a more accurate explanation of the technical state in each state. Therefore, a display example for such users will be described with reference to FIG. 5D. Among the rows of the table shown in FIG. 5D, the rows of the device state and display state explanation are exactly the same as those in FIG. 5C, so repeated explanations will be omitted. In addition, the display example of the AI assistant or character shown in the row of the display example of the AI assistant or character is almost the same as that in FIG. 5C, but differs in that a question mark (?) is displayed in the display example. The question mark (?) is a mark that the user operates when requesting an explanation from the AI response output device 10010, and may be called a help mark.
[0259] In the example of FIG. 5D, when the user selects the question mark (?) by user operation via the touch panel of the operation input unit 1107 or the display unit 10011 of FIG. 1B, the display of the AI assistant or character of the artificial intelligence response output device 10010 is changed to the display example shown in the row of the display example after the user operation. Specifically, regardless of whether the device state is (1), (2), or (3), an explanation of the technical state in each state is displayed. For example, in the example of FIG. 5D, if the device state is (1) normal state, a display may be displayed to explain that the device is in a normal state with no particular technical restrictions, saying "normal state." Also, if the device state is (2) a state in which a low-performance LLM or a response template database is being used, a display may be displayed to technically explain that the device is in a low-performance state, saying "low-performance mode." The display may be considered to be a display that explains the cause of the AI assistant or character displaying in a "sleepy" state.
[0260] In this case, a more technically detailed explanation may be provided. Specifically, a display such as "Low performance LLM usage mode" or "Canned response mode" may be provided. If the state of the device is (3) unresponsive, a display such as "Network connection unavailable" may be provided to technically explain the cause of switching to unresponsive mode. If the cause of switching to unresponsive mode is that the response from the LLM (large-scale language model connected using an API) on the network exceeds a specified time, a display such as "Response from the LLM is delayed" may be provided. If the cause of switching to unresponsive mode is that the usage amount, API usage amount, or API usage fee on the network has reached its upper limit, a display such as "LLM usage has reached its upper limit," "API usage has reached its upper limit," or "API usage fee has reached a specified amount" may be provided. These displays may be considered to be displays that explain the cause of the AI assistant or character display being "asleep."
[0261] According to the display example of FIG. 5D described above, even if there is a technical constraint in the response generation process in the AI response output device 10010, first, a direct explanation is not given to the user, but the state of the device is implicitly indicated by a change in the display state of the AI assistant or character, thereby further reducing the discomfort felt by the user. This display is more suitable for users who do not need a technical explanation. Furthermore, by displaying an operation mark for explaining the technical state, a display is provided that technically explains the state (normal state or state with technical constraints) of the response generation process in the AI response output device 10010 to the user who operates the mark. This makes it possible to provide a more suitable display for users who want to know the technical state accurately.
[0262] In the examples of Figures 5B, 5C, and 5D, the "sleeping" state is shown as an example of the display state of the AI assistant or character when "non-response compatible", but this is only an example, and the embodiment of this embodiment is not limited to this. Instead of the "sleeping" state, other display states that imply a situation in which the AI assistant or character cannot respond, such as "taking a break", may be used. In addition, in the examples of Figures 5C and 5D, the "sleepy" state is shown as an example of the display state of the AI assistant or character when a low-performance LLM or a response template database is used, but this is only an example, and the embodiment of this embodiment is not limited to this. Other display states that imply that the response performance of the AI assistant or character is low, such as "hungry", may be used.
[0263] According to the artificial intelligence response output device and the artificial intelligence response output system according to the fifth embodiment described above, it is possible to more appropriately switch the response generation process used by the artificial intelligence response output device depending on the connection state between the large-scale language model on the network and the artificial intelligence response output device, the response delay state from the large-scale language model on the network, or the usage amount of the large-scale language model on the network, etc. Also, when the artificial intelligence response output device according to the fifth embodiment is configured as an AI assistant device or a character conversation device, it is possible to perform a display that is less strange to the user.
[0264] <Example 6> Next, the sixth embodiment of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in each of the drawings of the first to fifth embodiments. Specifically, this is an example in which the response output is generated by more suitably combining the response generation process of the AI response output device 10010 with the response generation process by a large-scale language model on a network or the response generation process by a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010 and the response generation process by a response template database. In this embodiment, the differences from these embodiments will be described, and repeated explanations of the same configurations as these embodiments will be omitted.
[0265] As in the above-mentioned embodiment, the AI response output device 10010 may be called an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be called an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0266] An example of a response generation process in the AI response output device 1001 according to the sixth embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 shows an example of a flowchart of the response generation process in the AI response output device 1001 according to the sixth embodiment of the present invention. Specifically, a time axis in which time progresses from top to bottom, a process flow, and a response output example are shown. The response shown in the response output example may be output via display by the display unit 10011 of the AI response output device 1001 or audio output by the audio output unit 1140.
[0267] In the example of FIG. 6, first, at time t0, a user input is made via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 1001, requesting a response using a large-scale language model, and the control unit 1110 of the AI response output device 1001 acquires the user input (step 600). Next, at time t1, the control unit 1110 starts preparation for response output using the standard response phrase database stored in the storage unit 1170, and starts response output using the standard response phrase database (step 601). In the example of FIG. 6, at time t2, response output using the standard response phrase database is started, and as shown in the figure, the standard response is being output and has not been completed. "Good morning" in the figure indicates the output of the sentence "Good morning..." followed by the sentence up to the middle.
[0268] At time t3, before the response output using the response template database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, transmits the generated instruction sentence to a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the artificial intelligence response output device 10010, and starts requesting a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the response template database is completed, the control unit 1110 starts acquiring a response from the large-scale language model (step 603).
[0269] An example of a response output in which the response output using the fixed response phrase database is completed at time t5 is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is XXth month, isn't it?" is completed at time t5 using the fixed phrase stored in the fixed response phrase database and the date information stored in the memory. Here, at time t4 before the response output using the fixed response phrase database is completed, the control unit 1110 has already started acquiring a response from the large-scale language model. Therefore, at time t6 following time t5 when the response output display using the fixed response phrase database is completed, the control unit 1110 starts the response output from the large-scale language model following the response output using the fixed response phrase database (step 604). After that, at time t7, the response from the large-scale language model is output following the response output using the fixed response phrase database. When the response output from the large-scale language model is completed, the response output according to the process flow shown in FIG. 6 is completed (step 605).
[0270] Next, the effect of the processing flow shown in FIG. 6 of the present invention will be described. Processing of a large-scale language model requires a lot of computational resources. In general, even if inference, which requires less computational resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several to several tens of seconds from when the control unit starts a response request to the large-scale language model until it can obtain a response from the large-scale language model. This period corresponds to the period from time t3 to time t4 shown in FIG. 6. In addition, from time t0 when a user input is made to time t4, the control unit 1110 cannot obtain a response output from the large-scale language model, and therefore cannot output a response from the large-scale language model to the user.
[0271] 6, there is no start of preparation for response output using the fixed response phrase database, and no start of response output using the fixed response phrase database, the user may continue to wait for a period of several to more than ten seconds from time t0 when the user input is made to time t4 without receiving a response from the AI response output device 1001. For example, when the AI response output device 1001 is configured as an AI assistant device or a character conversation device, the waiting time may give the user a sense of discomfort.
[0272] In contrast, in the process flow according to the sixth embodiment of the present invention shown in Fig. 6, the control unit 1110 starts the process of outputting a response using a template response database, which requires less computational resources than the process of a large-scale language model, before the response acquisition from the large-scale language model is started. This prevents the user from having to wait from time t0 to time t4 without receiving a response from the AI response output device 1001. For the user, whether the response output is from the template response database or from the large-scale language model, it is the same as if it is a response from the AI response output device 1001.
[0273] Therefore, in the process flow shown in Fig. 6, by providing step 601 before step 603, the response of the AI response output device 1001 to the user can be artificially accelerated. This makes it possible to further reduce the sense of discomfort felt by the user due to the long waiting time. In addition, by outputting the response from the large-scale language model following the response using the response template database in step 604, the user can recognize these outputs as if they were a series of more natural outputs.
[0274] According to the artificial intelligence response output device and the artificial intelligence response output system of Example 6 described above, the waiting time for a user to receive a response from the artificial intelligence response output device can be shortened, and the sense of discomfort felt by the user can be further reduced.
[0275] Furthermore, the technology according to the present embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into higher quality, more reliable infrastructure. The introduction of this technology into infrastructure can contribute to economic development and support for human welfare with an emphasis on affordable and fair access for all people. This will contribute to "build resilient infrastructure, promote inclusive and sustainable development," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0276] In addition, the technology according to the present embodiment makes it possible to provide a more suitable AI response output technology. Such AI response output technology is expected to be introduced into public transportation facilities in order to improve access to the transportation system for vulnerable people. The introduction of this technology into public transportation facilities can contribute to improving traffic safety through the expansion of public transportation facilities, and realizing access to a sustainable transportation system that is safe, inexpensive, and easy to use for all people. This contributes to "Sustainable cities and communities" in the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0277] Although various embodiments have been described above in detail, the present invention is not limited to the above-described embodiments, and various modified examples are included. For example, the above-described embodiments are detailed descriptions of the entire system in order to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the configurations described. In addition, it is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. In addition, it is possible to add, delete, or replace a part of the configuration of each embodiment with another configuration. [Explanation of symbols]
[0278] 10010... artificial intelligence response output device, 10010... display unit, 10028... local LLM processing unit, 1107... operation input unit, 1110... control unit, 1132... communication unit, 1140... audio output unit, 1139... microphone, 1160... video control unit, 1170... storage unit, 1180... imaging unit< / audio> < / video> < / audio> < / video>
Claims
1. A control unit that obtains responses from a large-scale language model to instructions given to a large-scale language model, Display unit and Audio output section, Equipped with, The control state by the control unit includes a state in which the response from the large-scale language model is output via the display unit or the audio output unit. Response output device.
2. A response output device according to claim 1, Equipped with a storage unit, The aforementioned storage unit stores a database containing multiple standard phrases that serve as the basis for responses. The control state of the control unit includes a state in which it outputs a response generated based on a standard phrase stored in the database, rather than a response from the large-scale language model. Response output device.
3. A response output device according to claim 2, The control input unit or microphone, Equipped with, The control unit, When input is received from the user via the operation input unit or the microphone requesting a response from the large-scale language model, preparation for outputting a response based on a standard phrase stored in the database is initiated, and output of the response generated based on the standard phrase stored in the database is initiated via the display unit or the audio output unit. Before the output of the response generated based on the standard phrases stored in the database is completed, an instruction is sent to the large-scale language model, the acquisition of a response from the large-scale language model is started, and the acquired response from the large-scale language model is output via the display unit or the audio output unit, following the response generated based on the standard phrases stored in the database. Response output device.
4. A response output device according to claim 2, The display unit is capable of displaying multiple different AI assistants. The database in the storage unit stores predefined text data that can generate different responses corresponding to each of the multiple different AI assistants. Response output device.
5. A response output device according to claim 2, Equipped with a communications department, The control unit can send an instruction to a large-scale language model on a server on the network via the communication unit, and can obtain a response to the instruction from the large-scale language model on the server on the network. The control state by the control unit includes: When connection to the network via the communication unit is possible, a first state is reached in which a response obtained from a large-scale language model on a server on the network is output via the display unit or the audio output unit. In the event that connection to the network via the communication unit is not possible, a second state is reached in which a response generated based on a standard phrase stored in the database is output via the display unit or the audio output unit. There is, Response output device.
6. A response output device according to claim 2, Equipped with a communications department, The control unit can send an instruction to a large-scale language model on a server on the network via the communication unit, and can obtain a response to the instruction from the large-scale language model on the server on the network. The control state by the control unit includes: A first state in which, if a response from a large-scale language model on a server on the network connected via the communication unit can be obtained within a predetermined time, the response obtained from the large-scale language model on the network server is output via the display unit or the audio output unit. If a response from a large-scale language model on a server on the network connected via the communication unit cannot be obtained within the predetermined time, a second state is reached in which a response generated based on a standard phrase stored in the database is output via the display unit or the audio output unit. There is, Response output device.
7. A response output device according to claim 2, Equipped with a communications department, The control unit can send an instruction to a large-scale language model on a server on the network via the communication unit, and can obtain a response to the instruction from the large-scale language model on the server on the network. The control state by the control unit includes: A first state in which, when the usage of a large-scale language model on a server on the network connected via the communication unit has not reached a predetermined upper limit, a response obtained from the large-scale language model on the server on the network is output via the display unit or the audio output unit, A second state occurs in which, when the usage of a large-scale language model on a server on the network connected via the communication unit reaches a predetermined upper limit, a response generated based on a standard phrase stored in the database is output via the display unit or the audio output unit. There is, Response output device.
8. A response output device according to claim 1, The response output device includes a local large-scale language model processing unit capable of processing large-scale language models, The control state by the control unit includes: There is a state in which the response obtained from the large-scale language model of the local large-scale language model processing unit is output via the display unit or the audio output unit. Response output device.
9. A response output device according to claim 1, Communications Department and, It includes a local large-scale language model processing unit, The control state by the control unit includes: A first state in which an instruction is sent via the communication unit to a large-scale language model on a server on the network, and a response to the instruction is obtained from the large-scale language model on the server on the network and output via the display unit or the audio output unit, A second state in which the response obtained from the large-scale language model of the local large-scale language model processing unit is output via the display unit or the audio output unit, There is, Response output device.
10. A response output device according to claim 1, Communications Department and, It includes a local large-scale language model processing unit, The control state by the control unit includes: When connection to the network via the communication unit is possible, a first state is reached in which a response obtained from a large-scale language model on a server on the network is output via the display unit or the audio output unit. In the event that connection to the network via the communication unit is not possible, a second state is reached in which the response obtained from the large-scale language model of the local large-scale language model processing unit is output via the display unit or the audio output unit. There is, Response output device.
11. A response output device according to claim 1, Communications Department and, It includes a local large-scale language model processing unit, The control state by the control unit includes: A first state in which, if a response from a large-scale language model on a server on the network connected via the communication unit can be obtained within a predetermined time, the response obtained from the large-scale language model on the network server is output via the display unit or the audio output unit. If a response from the large-scale language model on the server on the network connected via the communication unit cannot be obtained within the predetermined time, a second state is reached in which the response obtained from the large-scale language model of the local large-scale language model processing unit is output via the display unit or the audio output unit. There is, Response output device.
12. A response output device according to claim 1, Communications Department and, It includes a local large-scale language model processing unit, The control state by the control unit includes: A first state in which, when the usage of a large-scale language model on a server on the network connected via the communication unit has not reached a predetermined upper limit, a response obtained from the large-scale language model on the server on the network is output via the display unit or the audio output unit, When the usage of the large-scale language model on the server on the network connected via the communication unit reaches a predetermined upper limit, a second state is reached in which the response obtained from the large-scale language model of the local large-scale language model processing unit is output via the display unit or the audio output unit. There is, Response output device.
13. A response output device according to claim 1, Equipped with a communications department, The control unit, The communication unit transmits an instruction to a first large-scale language model located on a server on the network, and the response to the instruction can be obtained from the first large-scale language model located on the server on the network. Furthermore, the communication unit transmits an instruction to a second large-scale language model located on a server on the network, and the response to the instruction is obtained from the second large-scale language model located on the server on the network. Response output device.
14. A response output device according to claim 13, The control state by the control unit includes: A first state in which a response obtained from the first large-scale language model on the server on the network via the communication unit is output via the display unit or the audio output unit, A second state in which a response obtained from the second large-scale language model on the server on the network via the communication unit is output via the display unit or the audio output unit, There is, Response output device.
15. A response output device according to claim 13, The control state by the control unit includes: A first state in which, when the usage of the first large-scale language model on the server on the network connected via the communication unit has not reached a predetermined upper limit, the response obtained from the first large-scale language model on the server on the network is output via the display unit or the audio output unit, When the usage of the first large-scale language model on the server on the network connected via the communication unit reaches a predetermined upper limit, a second state is reached in which a response obtained from the second large-scale language model on the server on the network is output via the display unit or the audio output unit. There is Response output device.
16. A response output device according to claim 1, Communications Department and, The control input unit or microphone, Equipped with, The control unit can send an instruction to a large-scale language model on a server on the network via the communication unit, and can obtain a response to the instruction from the large-scale language model on the server on the network. The control state by the control unit includes: When connection to the network via the communication unit is possible, a first state is reached in which a response obtained from a large-scale language model on a server on the network is output via the display unit or the audio output unit. In the event that connection to the network via the communication unit is not possible, even if the user requests a response from the large-scale language model via the operation input unit or the microphone, the display unit or the audio output unit will not output a response to that input, resulting in a second state in which, There is, Response output device.
17. A response output device according to claim 1, Communications Department and, The control input unit or microphone, Equipped with, The control unit can send an instruction to a large-scale language model on a server on the network via the communication unit, and can obtain a response to the instruction from the large-scale language model on the server on the network. The control state by the control unit includes: A first state in which, when the usage of a large-scale language model on a server on the network connected via the communication unit has not reached a predetermined upper limit, a response obtained from the large-scale language model on the server on the network is output via the display unit or the audio output unit, When the usage of the large-scale language model on the server on the network connected via the communication unit reaches a predetermined upper limit, even if the user requests a response from the large-scale language model via the operation input unit or the microphone, the display unit or the audio output unit will not output a response to that input, resulting in a second state. There is, Response output device.
18. A response output device according to claim 2, The aforementioned display unit is capable of displaying an AI assistant. Equipped with a communications department, The control unit can send an instruction to a large-scale language model on a server on the network via the communication unit, and can obtain a response to the instruction from the large-scale language model on the server on the network. The display control state of the AI assistant on the display unit by the control unit is as follows: A first display control state when connection to the network via the communication unit is possible, There is a second display control state when connection to the network via the communication unit is not possible, The display state of the AI assistant in the second display control state differs from the display state of the AI assistant in the first display control state, as it is a state in which the AI assistant is displayed in a dormant state. Response output device.
19. A response output device according to claim 18, The control input unit or microphone, Equipped with, In the second display control state described above, the display state when the AI assistant is asleep starts after the connection to the network via the communication unit becomes impossible, and before input is received from the user via the operation input unit or the microphone requesting a response from the large-scale language model. Response output device.
20. A response output device according to claim 2, The aforementioned display unit is capable of displaying an AI assistant. Equipped with an operation input unit or microphone, The display control state of the AI assistant on the display unit by the control unit is as follows: A first display control state when output using the response from the large-scale language model is possible via the display unit or the audio output unit, There is a second display control state in which, even if the user inputs a request for a response from the large-scale language model via the operation input unit or the microphone, the display unit or the audio output unit does not output a response to that input. The display state of the AI assistant in the second display control state is different from the display state of the AI assistant in the first display control state. Response output device.
21. A response output device according to claim 20, In the second display control state, the display state in which the AI assistant is displayed in a state different from the display state of the AI assistant in the first display control state is the state in which the AI assistant is displayed in a sleeping state. Response output device.
22. A response output device according to claim 20, In the second display control state, the display state in which the AI assistant is displayed in a state different from the display state of the AI assistant in the first display control state starts from a point in time when, even if there is input from the user via the operation input unit or the microphone requesting a response from the large-scale language model, the display unit or the audio output unit does not output a response to said input, but before the input requesting a response from the large-scale language model is received from the user via the operation input unit or the microphone. Response output device.
23. A response output device according to claim 20, The operation input unit or the microphone, of which at least the operation input unit is provided, The display unit displays a predetermined mark in the second display control state. The control unit, When a user selects a predetermined mark via the operation input unit, the system controls the display unit to display a message explaining the reason why the AI assistant is displayed in a state different from the display state of the AI assistant in the first display control state. Response output device.
24. A response output device according to claim 2, The aforementioned display unit is capable of displaying an AI assistant. The display control state of the AI assistant on the display unit by the control unit is as follows: A first display control state when output using the response from the large-scale language model is possible via the display unit or the audio output unit, There is a second display control state when the response that can be output from the response output device is a response that is of lower performance than the output using the response from the large-scale language model, The display state of the AI assistant in the second display control state is different from the display state of the AI assistant in the first display control state. Response output device.
25. A response output device according to claim 24, In the second display control state, the display state in which the AI assistant is displayed in a state different from the display state of the AI assistant in the first display control state is the display state in which the AI assistant is feeling sleepy. Response output device.
26. A response output device according to claim 24, Equipped with an operation input unit or microphone, In the second display control state, the display state in which the AI assistant is displayed in a state different from the display state of the AI assistant in the first display control state begins after the response output device has reached a state in which the output of the response is of lower performance than the output using the response from the large-scale language model, and before an input requesting a response from the large-scale language model is received from the user via the operation input unit or the microphone. Response output device.
27. A response output device according to claim 24, Equipped with an operation input section, The display unit displays a predetermined mark in the second display control state. The control unit, When a user selects a predetermined mark via the operation input unit, the system controls the display unit to display a message explaining the reason why the AI assistant is displayed in a state different from the display state of the AI assistant in the first display control state. Response output device.