Response output device

The response output device enhances user interaction and response generation by integrating a client application and large-scale language model, addressing the limitations of existing AI response output technologies.

JP2025187344APending Publication Date: 2025-12-25MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024096046
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing response output technologies using artificial intelligence do not adequately consider user interaction and configuration for optimal response delivery.

Method used

A response output device equipped with an input interface, control unit, storage unit, and output interface, capable of interacting with a large-scale language model to generate and output responses based on user input, utilizing a client application and a large-scale language model application to enhance user interaction.

Benefits of technology

Provides a more suitable and effective response output technique by improving user interaction and response generation through a comprehensive AI-driven system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025187344000001_ABST
    Figure 2025187344000001_ABST
Patent Text Reader

Abstract

To provide a suitable artificial intelligence response output technology.SOLUTION: A response output device includes an input interface, a control part, a storage part, and an output interface, the control part can execute a client application that can transmit and receive information with a large-scale language model application for controlling a large-scale language model stored in a server outside the device or inside the device, the client application can generate an instruction sentence to the large-scale language model on the basis of a user input, can transmit control information different from the instruction sentence to the large-scale language model application, can transmit the instruction sentence to the large-scale language model application, can receive a response sentence being a result of inference executed by the large-scale language model from the large-scale language model application, and can output a response based on the response sentence to a user through the output interface, and the storage part stores setting about a feature of conversations of a character.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a response output device. [Background technology]

[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special table 2019-528512 publication Summary of the Invention [Problem to be solved by the invention]

[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.

[0005] An object of the present invention is to provide a more suitable response output technique. [Means for solving the problem]

[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes multiple means for solving the above problem, but one example thereof may be a response output device including an input interface that accepts user input, a control unit, a storage unit, and an output interface that outputs a response to a user, wherein the control unit is capable of executing a client application that is capable of sending and receiving information with a server external to the response output device or a large-scale language model application that controls a large-scale language model stored in the response output device, wherein the client application is capable of generating instruction statements for the large-scale language model based on the user input accepted via the input interface, transmitting control information different from the instruction statements to the large-scale language model application, transmitting the instruction statements to the large-scale language model application, receiving response statements that are the results of inference performed by the large-scale language model from the large-scale language model application, and outputting a response based on the response statements to the user via the output interface, and wherein the storage unit stores settings related to characteristics of a character's conversation. [Effects of the Invention]

[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments. [Brief explanation of the drawings]

[0008] [Figure 1A] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 1B] 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 1C] 1 is a diagram showing an example of the operation of an artificial intelligence response output device and system according to an embodiment of the present invention; [Figure 2A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2B]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2C] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2I] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2J] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2K] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2L] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3C]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3I] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4A] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 5A] FIG. 2 is an explanatory diagram of an example of the operation of the artificial intelligence response output device according to an embodiment of the present invention. [Figure 5B] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5C] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5D] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 6] FIG. 10 is an explanatory diagram of an example of a response generation process of the AI ​​response output device according to an embodiment of the present invention. [Figure 7]1 is an explanatory diagram of an example of an artificial intelligence response output device and system according to an embodiment of the present invention; [Figure 8A] 1 is an explanatory diagram illustrating an example of the configuration of an artificial intelligence response output system according to an embodiment of the present invention; [Figure 8B] 1 is an explanatory diagram illustrating an example of the configuration of an artificial intelligence response output system according to an embodiment of the present invention; [Figure 8C] 1 is an explanatory diagram illustrating an example of the configuration of an artificial intelligence response output system according to an embodiment of the present invention; [Figure 9A] FIG. 1 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 9B] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. [Figure 9C] FIG. 1 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 9D] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 9E] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. [Figure 9F] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 9G] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 10A] FIG. 1 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 10B] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. [Figure 10C] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 10D] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 11A] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. [Figure 11B] FIG. 1 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 11C] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 12A] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 12B] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 12C] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 13A] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. [Figure 13B] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. [Figure 13C] FIG. 10 is an explanatory diagram of an example of a conversation used to explain an embodiment of the present invention. [Figure 13D] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 13E] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 13F] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 13G] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 13H] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 13I] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 13J] FIG. 10 is an explanatory diagram of an example of a reading speed adjustment according to an embodiment of the present invention. [Figure 14] FIG. 10 is an explanatory diagram of an example of table information according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.

[0010] Note that if the AI ​​response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI ​​response output device has an audio output function, it may be referred to as an audio output device. The AI ​​response output device may simply be referred to as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI ​​response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI ​​response output device or the display output of the AI ​​response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI ​​response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI ​​response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.

[0011] Example 1 As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.

[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI ​​response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system in which the AI ​​response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.

[0013] In the example of FIG. 1A, the AI ​​response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Furthermore, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which the pixels emit light themselves. Furthermore, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.

[0014] 1A, the audio output unit 1140 provided in the AI ​​response output device 10010 is configured with a speaker. The AI ​​response output device 10010 also has a microphone 1139, which can pick up the user's voice. By audio input from the microphone 1139 or operation input from the user via an operation input unit (described later), the AI ​​response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.

[0015] The AI ​​response output device 10010 may be provided with a local large-scale language model in the AI ​​response output device 10010. In this case, the response of the large-scale language model may be output as a display output of the display unit 10011 and / or an audio output of the audio output unit 1140.

[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.

[0017] Alternatively, the AI ​​response output device 10010 may also include a local large-scale language model, and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI ​​response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.

[0018] The configuration when the AI ​​response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions, or may go via a router or repeater. The AI ​​response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI ​​response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI ​​response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as one system.

[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI ​​response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.

[0020] 1A shows an example in which a display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of FIG. 1A, the prompt display area 10051 displays an icon 10052 representing a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of FIG. 1A, the artificial intelligence response display area 10061 displays an icon 10062 representing an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI ​​response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI ​​response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.

[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment. These large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in these artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called Transformer. Reference 1, for example, has been published as an example of the learning of these models.

[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf

[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI ​​response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI ​​response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001 and the AI ​​response output device 10010. In this way, instead of performing pre-learning for each application or terminal, replicating the large-scale language model, which is the base model generated by large-scale pre-learning, and using it on individual servers and terminals allows the resources used for learning to be shared, resulting in good resource efficiency.

[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.

[0025] Furthermore, large-scale language models can pre-train natural languages ​​and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. FIG. 1A illustrates a large-scale language model server 20001 having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are AI models generated by large-scale pre-training on natural language and types of information other than natural language text information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. In addition, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.

[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf

[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.

[0028] The AI ​​response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a nonvolatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, and an imaging unit 1180. The AI ​​response output device 10010 may have a large screen, such as a monitor or television.

[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire the user input that serves as the basis for a prompt to the large-scale language model, which is the artificial intelligence.

[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN-type communication method. This allows the AI ​​response output device 10010 to communicate with various servers connected to the Internet 19000.

[0031] The AI ​​response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.

[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI ​​response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.

[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller or the like, and inputs a signal regarding an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. In addition to a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may also be used by, for example, an administrator to operate the AI ​​response output device 10010. By the user's operation input via the operation input unit 1107, the AI ​​response output device 10010 can acquire a user input that serves as the basis for a command sentence (prompt) to a large-scale language model, which is an AI. Note that a modified configuration in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107 is also possible.

[0034] The video signal input unit 1131 is connected to an external video output device and inputs video data. The video signal input unit 1131 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces, etc.

[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.

[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.

[0037] The microphone 1039 is a microphone that picks up sounds around the AI ​​response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI ​​response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.

[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI ​​response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.

[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.

[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 controls video switching, such as which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also control image processing of the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling, which enlarges, reduces, or deforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.

[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI ​​response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.

[0042] The nonvolatile memory 1108 stores various data used by the AI ​​response output device 10010. The data stored in the nonvolatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI ​​response output device 10010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.

[0043] The local LLM processing unit 10028 has a memory capable of holding a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI ​​response output device 10010.

[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI ​​response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, which is a speaker or the like.

[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 of the AI ​​response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and responses are obtained from these large-scale language models. All of this control can be performed by the control unit 1110.

[0046] The storage unit 1170 may also store a fixed response phrase database (which may be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction sentences from the AI ​​response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI ​​response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI ​​response output device 10010.

[0047] Furthermore, in the example of the fixed response phrases in the database shown in FIG. 1C, if multiple fixed response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the fixed response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the fixed response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.

[0048] Next, an example of condition number 5 shown in Fig. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I might not know about that." By responding in this way, the user can be prompted to input again, and the corrected user input can be waited for.

[0049] Next, an example of condition number 6 shown in Fig. 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI ​​response output device 10010 shown in Fig. 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the fixed response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI ​​response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.

[0050] The AI ​​response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI ​​response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).

[0051] 1C described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI ​​response output device 10010. However, the response template database (response template DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI ​​response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).

[0052] In the above explanation, it has been explained that the AI ​​response output device 10010 has a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.

[0053] 1A and 1B, an example has been described in which the AI ​​response output device 10010 includes the display unit 10011. However, the AI ​​response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI ​​may be configured to accept input from a user to the AI ​​via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the input from the user via the voice output unit 1140.

[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference of the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.

[0055] <Example 2> Next, as a second embodiment of the present invention, an example will be described in which the AI ​​response output device 10010 described in the first embodiment is connected to the Internet, and is connected to a server equipped with a large-scale language model AI via the Internet to perform operation. In this embodiment, differences from the first embodiment will be described, and repeated explanations of the same configurations as those of these embodiments will be omitted.

[0056] An example of a connection state between an AI response output device 10010 and a large-scale language model server 19001 according to the second embodiment of the present invention will be described with reference to FIG. 2A. The AI ​​response output device 10010 according to the second embodiment may be called a character conversation device. A system including the AI ​​response output device 10010 and the large-scale language model server 19001 according to the second embodiment may be called a character conversation system. A video of a character 19051 is displayed on a display unit 10011 displayed by the AI ​​response output device 10010. The video of the character 19051 is generated by rendering a 3D model of the character in a virtual space.

[0057] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can serve as an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this embodiment may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.

[0058] In the example of FIG. 2A, the audio output unit 1140 provided in the AI ​​response output device 10010 is composed of a speaker. The AI ​​response output device 10010 also has a microphone 1139, which can pick up the user's voice. The AI ​​response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A, an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication is also acceptable. The communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions. The AI ​​response output device 10010 can communicate with a large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI ​​response output device 10010 can communicate with a second server 19002 different from the large-scale language model server 19001 via a communication device 19011 and the Internet 19000. A configuration including the AI ​​response output device 10010 and the large-scale language model server 19001 may be considered as one system.

[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2B. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that Figure 2B does not illustrate communication paths such as the Internet 19000 shown in Figure 2A. Figure 2B also illustrates a user 230 of the artificial intelligence response output device 10010.

[0060] Here, we will explain the sequence of operations of the AI ​​response output device 10010. The AI ​​response output device 10010 loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.

[0061] First, the AI ​​response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is in natural language. Note that the extraction of the text of the words spoken by the user 230 may be performed continuously for all words, or may start when the user utters a word within a predetermined period after a trigger keyword. For example, the trigger keyword may be when the user utters "hello" followed by the character's name. For example, if the name of the character 19051 is "Koto," then "Hello, Koto!" may be the trigger keyword.

[0062] The character operation program of the AI ​​response output device 10010 creates a prompt based on the text of the words spoken by the user 230 and transmits the prompt to the large-scale language model server 19001 using an API. Here, the prompt may be metadata containing information written in a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON. The prompt stores text information in a natural language as a main message. The prompts transmitted from the AI ​​response output device 10010 to the large-scale language model server 19001 include setting prompts that store instructions such as initial settings, and user prompts that reflect instructions from the user. Type identification information identifying whether the prompt is a setting prompt or a user prompt may be stored in a portion of the prompt other than the main message. When the character operation program of the AI ​​response output device 10010 creates a prompt based on the text of the words spoken by the user 230, it creates the user prompt and sends it to the large-scale language model server 19001.

[0063] Next, the large-scale language model of the artificial intelligence of the large-scale language model server 19001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response including natural language text information based on the result. The large-scale language model server 19001 uses an API to send the response to the artificial intelligence response output device 10010. The response stores the natural language text information as a main message. Here, the response may be metadata storing information written in the same format as the instruction (e.g., a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON). If the response uses the same format as the instruction, type identification information may be stored outside the main message to indicate that the information is a different type from the initial setting instruction and the user instruction. For example, information indicating that the response is a response from the large-scale language model may be stored.

[0064] Next, the AI ​​response output device 10010 receives a response from the large-scale language model server 19001 and extracts the natural language text information stored as the main message in the response. The character operation program of the AI ​​response output device 10010 uses speech synthesis technology to generate a natural language voice as a response to the user based on the natural language text information extracted from the response, and outputs the voice from the speaker, i.e., the voice output unit 1140, so that it sounds as if it were the voice of the character 19051. This process may also be referred to as the character's "speaking."

[0065] Conversation examples 1 to 5 in Fig. 2C show specific examples of response voices from character 19051 in response to words from user 230, which are generated by the above-described processing of AI response output device 10010 and large-scale language model server 19001. In this way, user 230 can converse with character 19051 as if it were a real person.

[0066] 2B or a system including the AI ​​response output device 10010, there is no need to install a large-scale language model, which requires a huge amount of data and computational resources for learning, in the AI ​​response output device 10010. In addition, the advanced natural language processing capabilities of the large-scale language model can be utilized via an API, and when a user speaks to a character, a more appropriate response can be given to the user, enabling a more appropriate conversation to be held.

[0067] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2D. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2D shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.

[0068] FIG. 2D also shows the exchange of instructions and responses in chronological order, from the display setting instruction, the first round of user instructions and their responses, to the fourth round of user instructions and their responses.

[0069] As shown in FIG. 2D , the setting directive can instruct the large-scale language model of the artificial intelligence of the large-scale language model server 19001 as initial settings, such as the name of the large-scale language model itself, the role to be played, and conversation characteristics. The user's name can also be understood as the initial setting. This allows the large-scale language model to generate responses from the first round onward while respecting the role. Then, to a user who hears the voice of the character 19051 based on the responses from the first round onward, the character 19051 will seem to have the setting and personality of the person described in the setting directive. Furthermore, the large-scale language model server 19001 according to this embodiment includes a memory that stores the content of the conversation until the end of the series of conversations, and is configured to store a series of user directives and their responses and then generate responses. This allows a conversation like that shown in FIG. 2D to be realized.

[0070] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2E shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.

[0071] 2E shows an example of a case where, after the series of conversations shown in FIG. 2D has ended, the user 230 speaks to the character 19051 again to start a new conversation. In FIG. 2E, the exchange of instructions and responses is shown in chronological order, from the first round of user instructions and their responses to the third round of user instructions and their responses.

[0072] Here, "termination" of the "continuation of a series of conversations" refers to a process in which, when a predetermined condition is met, the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversation that the large-scale language model server 19001 has maintained while the series of conversations was continuing. An example of the predetermined condition is, for example, when the AI ​​response output device 10010 issues an instruction to the large-scale language model server 19001 to "terminate" the "continuation of a series of conversations." Another example of the predetermined condition is, for example, when a predetermined time or more has passed since the AI ​​response output device 10010 stopped sending instruction statements to the large-scale language model server 19001 regarding the series of conversations (timeout). Another example of the predetermined condition is when, after authentication processing has been performed in the connection between the AI ​​response output device 10010 and the large-scale language model server 19001, the authentication processing is terminated due to factors such as communication disconnection or the AI ​​response output device 10010 being powered off while exchanging the instruction statements and responses.

[0073] Note that when the "continuation of a series of conversations" "ends," the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversations that it had maintained while the series of conversations was continuing. Therefore, even though the conversation shown in FIG. 2E takes place after the series of conversations shown in FIG. 2D, the server's response to the user's instruction is a response with content that does not include any memory of the character's name set in the large-scale language model, the role to be played, the characteristics of the conversation, or the user's name, which were included in the setting instruction shown in FIG. 2D. Similarly, the conversation shown in FIG. 2E is a response with content that does not include any memory of the series of conversations shown in FIG. 2D. In other words, with the "end" of the "continuation of a series of conversations" shown in FIG. 2D, the conversation in FIG. 2E starts from an initialized state of the large-scale language model of the artificial intelligence of the large-scale language model server 19001.

[0074] This causes the user 230 to feel as if the character 19051 has lost its memory or is a completely different person. From the user 230's perspective, the character's response feels very strange, resulting in a feeling of loneliness and disappointment. Such behavior poses a problem in that it is not possible to ensure the consistency of the settings and memories of the character 19051, such as its name, role, conversational characteristics, and personality, displayed on the AI ​​response output device 10010.

[0075] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2F. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2F shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.

[0076] FIG. 2F illustrates an example of a case in which the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations illustrated in FIG. 2D has ended. Unlike the process illustrated in FIG. 2E, in the process illustrated in FIG. 2F, when starting a new conversation, the AI ​​response output device 10010 transmits a setting instruction as the first instruction to the large-scale language model server 19001. The setting instruction stores the same natural language text as the initial setting instruction of FIG. 2D. This may be referred to as a reset text. The setting instruction is followed by natural language text describing the history of past conversations. This may be referred to as a conversation history text. The AI ​​response output device 10010 may record the history of past conversations as natural language text information in the storage unit 1170 while the series of conversations illustrated in FIG. 2D is continuing, linking the history of the conversations to information on the date and time of the conversations. If there are conversations on different dates, each conversation may be recorded linked to date and time information, and the conversation history may be accumulated. When generating a setting instruction for the first instruction of a later conversation, such as that shown in Figure 2F, the natural language text information of the conversation recorded in storage unit 1170 and the information on the date and time when the conversation took place can be read out and used to generate the setting instruction.

[0077] When natural language text information of past conversation history is used to generate the setting instruction sentence, the format can be determined freely to a certain extent because it is data to be sent to a large-scale language model, but as shown in Figure 2F, it is sufficient to prepare prefixes and suffixes in natural language, such as "I talked about the following on ____ day of ____ month," or "You talked about the following on ____ day of ____ month," and combine these with the natural language text information of the recorded conversation to generate the text of the setting instruction sentence. Also, information on the date and time of the conversation read from storage unit 1170 may be combined with the above-mentioned "____ day of ____ month" portion to form part of the text of the setting instruction sentence.

[0078] Even if user 230 speaks to character 19051 again to start a new conversation after a series of conversations has ended, by performing the above-described generation process and transmission process of the setting instruction sentence in Fig. 2F, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation. This is preferable because it is perceived by the user as ensuring the consistency of the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation.

[0079] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2G shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.

[0080] Figure 2G shows an example of a series of conversations shown in Figure 2F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 2G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 2F, so repeated description is omitted.

[0081] As shown in the natural language text of the server response in the table of Figure 2F, by using the setting directive shown in Figure 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 reflects the settings and conversation history of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation. This is more preferable because it allows the user to recognize that the settings and memories of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation are more closely matched. Note that this allows the characters to be viewed as the same from the user's perspective, and may therefore be referred to as a pseudo-identity of the characters from the user's perspective.

[0082] Furthermore, from the user's perspective, they can share memories with the character, providing a more enjoyable character conversation experience.

[0083] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, FIG. 2H shows an example of the operation in which a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. A character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input input to the operation input unit 1107 or an operation detected by a touch operation input sensor of the display unit 10011.

[0084] In the example of FIG. 2H, in addition to character 19051 (named "Koto") used in the description of FIGS. 2A to 2G, character 19052 (named "Tom") and character 19053 (named "Necco") are shown. Character 19051 (named "Koto") and character 19052 (named "Tom") are human characters, and character 19053 (named "Necco") is a cat character. The display of characters displayed on display unit 10011 can be switched by switching the display on display unit 10011 to display images generated by rendering the characters in different virtual 3D spaces for each character.

[0085] Furthermore, it is preferable that the character operation program executed by control unit 1110 also changes the synthetic voice used for the "speech" of each character when switching the display of the characters displayed on display unit 10011. This can be achieved by storing synthetic voice data of voice tones associated with each character in storage unit 1170 in advance, and performing synthetic voice change processing when switching the display of the character.

[0086] In the example of Fig. 2H, the user 230 is configured to be able to converse with any of the characters. In the AI ​​response output device 10010 of Fig. 2H, each of these characters is set with a different role, name, conversational characteristics, personality, etc. Also, the memories of each character based on the conversation history are managed as different for each character.

[0087] Therefore, the AI ​​response output device 10010 constructs a database shown in FIG. 2I in the storage unit 1170, and manages the character settings and the character conversation history using this database.

[0088] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Fig. 2I. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.

[0089] A character operation program executed by the control unit 1110 of the AI ​​response output device 10010 constructs the database 19200 in, for example, the storage unit 1170. The character ID is an identification number that identifies each of multiple characters that can be displayed on the AI ​​response output device 10010, and may be a natural number or may use alphabets, etc. The name is data of the name of each of multiple characters that can be displayed on the AI ​​response output device 10010.

[0090] The initial setting instruction is text information in a natural language that explains the settings such as the role, name, conversational features, or personality of each of multiple characters that can be displayed on the AI ​​response output device 10010. The initial setting instruction is natural language text information that is the main data of the setting instruction sent from the AI ​​response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the description can be read as is by the large-scale language model of the AI ​​in the large-scale language model server 19001.

[0091] The conversation histories, which continue as conversation histories 1, 2, ..., are records of conversations between each character and the user, and are recorded separately for each character. The conversation histories will be included in the natural language text information, which is the main data of the setting instruction statement sent from the AI ​​response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the conversation histories be readable as is by the large-scale language model of the AI ​​in the large-scale language model server 19001.

[0092] When the character displayed on the display unit 10011 of the AI ​​response output device 10010 is switched, the character operation program executed by the control unit 1110 of the AI ​​response output device 10010 uses the database 19200 of Figure 2I to select and switch the initial setting instruction statement and conversation history used for the natural language text information that is the main data of the setting instruction statement transmitted from the AI ​​response output device 10010 to the large-scale language model server 19001 so as to correspond to the character displayed on the display unit 10011 of the AI ​​response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the conversation history in the conversation history area of ​​the database 19200 of Figure 2I that corresponds to the character displayed on the display unit 10011.

[0093] By using the database 19200 in this manner, the character operation program executed by the control unit 1110 of the AI ​​response output device 10010 establishes a conversation between the user 230 and the character using character utterances that utilize responses from the same AI large-scale language model of the same large-scale language model server 19001. From the user's perspective, the uniqueness of each character's personality and other settings is preserved, and it appears as if each character's unique conversational memories continue. This is more preferable because it appears as if the identity of each character's settings and memories, such as their role, name, conversational characteristics, or personality, from the time of the previous conversation has been more consistently maintained. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.

[0094] Therefore, even when the AI ​​response output device 10010 is configured to switch the character displayed on the display unit 10011 from among multiple character candidates, the operation using the database 19200 described above reduces the sense of discomfort felt by the user when conversing with each character, and the user can share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.

[0095] Note that if the initial setting instructions for multiple characters cannot be edited by the user, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained close to the intentions of the provider of the AI ​​response output device 10010 or the creator of the character content. Alternatively, the initial setting instructions for each character may be edited by the user in response to input from the operation input unit 1107 or the like. In this case, the character's role, name, conversational characteristics, or personality can be set to a preferred setting, allowing the user to converse with a character that they have individually set. In this case, the character's 3D model, its rendered image, and the type of synthesized voice for the character may be replaced accordingly.

[0096] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2J. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service at a lower cost using the character conversation device based on the artificial intelligence response output device 10010 and the character conversation system based on the artificial intelligence response output device 10010 and the large-scale language model server 19001 will be described.

[0097] As explained in Figure 2B, it is extremely resource inefficient to train a large-scale language model at this level of artificial intelligence by limiting it to a specific application. Therefore, it is more resource efficient to perform large-scale training to generate a foundation model that can be applied to a variety of applications, and then use it on various devices via an API (Application Programming Interface). In such cases, providers of large-scale language models often recover the costs incurred in training the large-scale language model from device users as API usage fees. In natural language models, API usage fees are often charged based on the number of tokens, which are units of words that separate sentences, processed.

[0098] Therefore, in the artificial intelligence response output device 10010 of the second embodiment of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010 and a character conversation service using a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.

[0099] For example, by using the processing and configuration of Examples 1 to 3 shown in the table of Figure 2J, it is technically possible to reduce the number of tokens in the natural language text information transmitted between the AI ​​response output device 10010 and the large-scale language model server 19001 using an API.

[0100] Example 1 is an example of a method for reducing the number of tokens in the conversation history text stored in the API setting directive and transmitted, in which the conversation history text is shortened using a document summarization process to reduce the number of tokens. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. The text summarization may be performed at the start of the next conversation, but it is more time-efficient to perform it at the end of the "series of conversations."

[0101] Furthermore, the text summarization process may be requested from the large-scale language model itself of the large-scale language model server 19001. However, in this case, the effect of saving the number of tokens is small. Therefore, for example, if the second server 19002 provides text summarization process for natural language via an API at a lower cost than the large-scale language model of the large-scale language model server 19001, the text summarization process can be requested from the second server 19002 via the API, and the text summary of the conversation history can be stored in a setting instruction statement for the large-scale language model server 19001 and transmitted.

[0102] Furthermore, if it is only text summarization processing, it can be performed on the terminal side, and text summarization can be performed by the control unit 1110 executing a document summarization program loaded in the memory 1109 of the AI ​​response output device 10010. In this case, the effect of saving the number of tokens is high. Furthermore, even if the conversation history becomes long, if an upper limit on the number of characters after summary is specified in the text summarization processing, the upper limit on the text length of the conversation history is determined, so that an upper limit on the number of tokens can be set and tokens can be saved.

[0103] In addition, since the text information of the character's initial settings, such as the character's role, name, conversational characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information of the character's initial setting instructions and reduce the number of tokens of the text information in the conversation history.

[0104] The processing described in Example 1 may be performed by the character movement program executed by the control unit 1110 controlling each unit.

[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored in the API setting instruction and transmitted. For example, the number of tokens is reduced by deleting the oldest conversation history among the conversation histories with characters recorded in the storage unit 1170. Specifying an upper limit on the number of characters in the conversation history determines the upper limit on the length of the conversation history text, thereby setting an upper limit on the number of tokens and enabling token conservation. Alternatively, a predetermined period of the conversation history may be specified and conversation history exceeding that period may be deleted. This also enables token conservation. Note that, in Example 2, the text information for the character's initial setting, such as the character's role, name, conversation characteristics, or personality, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the text information in the character's initial setting instruction and reduce the number of tokens in the text information for the conversation history.

[0106] The processing described in Example 2 may be performed by the character operation program executed by the control unit 1110 controlling each unit.

[0107] Example 3 is a method of reducing the number of tokens by reducing the frequency of sending setting instruction sentences using an API. Specifically, even after the device is powered on, after the displayed character is switched, or after the video settings and synthetic voice settings of the displayed character are completed, setting instruction sentences are not sent in advance, and only when the control unit 1110 determines that natural language text information included in the user's speech picked up by the microphone 1139 is text information that should use a large-scale language model of artificial intelligence is the setting instruction sentence sent to the large-scale language model server 19001, thereby reducing the frequency of sending setting instruction sentences to the large-scale language model server 19001 and reducing the number of tokens.

[0108] Specifically, for example, after the device is powered on or after an operation input to switch the displayed character is made, character 19051 (named "Koto") is displayed on display unit 10011 as shown in FIG. 2H by display processing of display unit 10011 controlled by a character operation program executed by control unit 1110. At this time, for example, if a synthetic voice for the character 19051 to appear is stored and prepared in storage unit 1170 or the like, a synthetic voice for the character's appearance such as "Good morning, I'm Koto," "Hello, I'm Koto," or "Good evening, I'm Koto" may be output from the speaker, which is audio output unit 1140. At this time, the image of character 19051 has already been set as the image of the character to be displayed on display unit 10011, and the synthetic voice output from the speaker, which is audio output unit 1140, is set to the synthetic voice corresponding to character 19051.

[0109] Here, the inference process of the large-scale language model of the artificial intelligence in the large-scale language model server 19001 already explained also takes time as the instruction sentence becomes longer. In particular, if the setting instruction sentence includes text information related to past conversation history, the instruction sentence will have a large number of tokens, which increases the inference process time. The setting instruction sentence itself and its response are not output to the user 230. Based on the response of the user instruction sentence following the setting instruction sentence, synthesized speech as the character's "utterance" is output from the speaker, which is the voice output unit 1140. Therefore, it may seem preferable at first glance to send the setting instruction sentence from the AI ​​response output device 10010 to the large-scale language model server 19001 in advance and complete the inference process of the large-scale language model for the setting instruction sentence in advance, because this would result in a faster response in the output of synthesized speech of the character's 19051's "utterance" after the user 230 speaks to the character 19051.

[0110] However, if the setting instruction is transmitted to the large-scale language model server 19001 before the user 230 speaks, and the inference processing of the large-scale language model for the setting instruction is completed in advance, for example, the user 230 may turn off the power of the AI ​​response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or the user 230 may switch the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. In this case, the number of tokens processed in the inference processing of the large-scale language model by transmitting the setting instruction to the large-scale language model server 19001 in advance becomes the number of processed tokens for which the usage fee is wasted. This hinders the provision of a character conversation device using the AI ​​response output device 10010 and a character conversation service using a character conversation system using the AI ​​response output device 10010 and the large-scale language model server 19001 at a lower cost to users.

[0111] Therefore, after the AI ​​response output device 10010 is powered on or after an operation to switch the displayed character is input, it sets the image of character 19051 as the image of the character to be displayed on the display unit 10011 under the control of the character operation program executed by the control unit 1110, and sets the synthetic voice output from the speaker, which is the audio output unit 1140, to the synthetic voice corresponding to character 19051, and it is desirable that the AI ​​response output device 10010 continue not to send a setting instruction statement to the large-scale language model server 19001 until the user 230 recognizes that he or she is speaking to character 19051.

[0112] Here, the point in time at which it is recognized that the user 230 is speaking to the character 19051 may be, for example, the point in time at which the trigger keyword described in Fig. 2B is detected, or the point in time at which the text of the words spoken by the user 230 is extracted. In this way, the number of processing tokens that waste usage fees can be reduced, and a character conversation service by a character conversation device using the AI ​​response output device 10010 or a character conversation system using the AI ​​response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.

[0113] Furthermore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051 has passed, it is desirable to continue not sending setting instructions to the large-scale language model server 19001, for example, if the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Specifically, examples of preset keywords include keywords such as "try jumping" and "try dancing," which are keywords by which the user 230 requests the character 19051 to react, such as by animating the character 19051 to move or uttering a synthesized voice. In this case, the character operation program executed by the control unit 1110 reads out motion data, animation video, and / or synthesized voice data corresponding to the character 19051 stored in the storage unit, and uses these data to generate video to be displayed on the display unit 10011 and output synthesized voice from the speaker, which is the audio output unit 1140.

[0114] Such processing does not necessarily require inference processing of the large-scale language model of the large-scale language model server 19001. After the processing, if the user 230 turns off the power of the AI ​​response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or if the user 230 switches the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, if the setting instruction text is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be the number of processing tokens that unnecessarily consumes the usage fee.

[0115] Therefore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051, it is desirable to continue not sending the setting instruction statement to the large-scale language model server 19001 until it is determined whether or not the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Only when it is determined that inference processing of a large-scale language model is required is it desirable to send the setting instruction statement to the large-scale language model server 19001 and proceed with the inference processing of the large-scale language model.

[0116] The processing described in Example 3 may be performed by the character operation program executed by the control unit 1110 controlling each unit.

[0117] According to the method for reducing (saving) the number of processing tokens for a large-scale language model using the examples of Figure 2J described above, a character conversation device using the AI ​​response output device 10010 and a character conversation system using the AI ​​response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.

[0118] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2K. The example of FIG. 2K shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which has been described in each of FIGS. 2A to 2J, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063, which is a response from the large-scale language model, together with an image of a character 19051 on the display unit 10011. The text 10063, which is a response from the large-scale language model, may be displayed superimposed on top of the image of the character 19051, as shown in FIG. 2K. Alternatively, the text 10063, which is a response from the large-scale language model, may be displayed together with the image of the character 19051 without being superimposed on the image of the character 19051.

[0119] The display in Figure 2K is an example, but for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by audio.

[0120] Therefore, in this case, the control unit 1110 may perform control to start a display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, as shown in Fig. 2K. In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more suitably use the character conversation device (artificial intelligence response output device 10010). Note that the user 230 may be configured to manually switch ON / OFF the display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011.

[0121] Next, an example of a fixed response phrase database (fixed response phrase DB) in the character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in FIGS. 2H and 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition numbers and condition contents are the same as those in FIG. 1C. In the example of FIG. 2L, individual fixed response phrases are set for each of the multiple characters in response to these conditions. For example, fixed response phrases for each condition are stored for each of the three characters, Character 1: Koto, Character 2: Tom, and Character 3: Necco, as described in FIGS. 2H and 2I. The output control of fixed response phrases is the same as that in FIG. 1C, and therefore will not be repeated.

[0122] In the example of FIG. 2L, the control unit 1110 selects a corresponding fixed response phrase from a fixed response phrase database (fixed response phrase DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses the selected fixed response phrase for output control as a response uttered by the character. For example, in the example of the fixed response phrase database (fixed response phrase DB) in FIG. 2L, even under the same conditions, the fixed response phrases are changed to expressions or contents that correspond to the individuality of the characters. This allows the character conversation device (artificial intelligence response output device 10010) to provide the user with conversations that correspond to the individuality of the displayed characters. The user can feel that each character has a more consistent personality. This makes it possible to realize a character conversation device (artificial intelligence response output device 10010) that gives multiple characters a more realistic presence.

[0123] The fixed response phrase database (fixed response phrase DB) of FIG. 2L described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI ​​response output device 10010. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 2L may also be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the AI ​​response output device 10010, instead of a response generated using a large-scale language model stored in each server. In this way, even if the AI ​​response output device 10010 does not have a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).

[0124] The character conversation device and character conversation system according to the second embodiment described above can reduce the sense of discomfort felt by the user from the conversation with the character displayed on the AI ​​response output device 10010. Furthermore, the character conversation device and character conversation system according to the second embodiment can provide the character conversation service to the user at a lower cost.

[0125] In the above description of the second embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 19001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B, and the large-scale language model held by the local LLM processing unit 10028 may be used instead of the large-scale language model held by the large-scale language model server 19001. In this case, in the above description of the second embodiment, the large-scale language model held by the large-scale language model server 19001 may be read as the large-scale language model held by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).

[0126] In this case, too, it is possible to further reduce the sense of discomfort felt by the user from conversations with characters displayed on the AI ​​response output device 10010. Note that when a large-scale language model held by the local LLM processing unit 10028 is used instead of the large-scale language model held by the large-scale language model server 19001, there is less need to consider usage fees according to the number of processed tokens, but by reducing the number of processed tokens even for the large-scale language model held by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power required for inference. In this case, it is possible to provide users with a character conversation service that consumes less power.

[0127] In the above description of the second embodiment, an example has been described in which a conversation history with a character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). Alternatively, a conversation history with a character may be recorded and stored in a second server 19002 or another cloud server connected to the Internet 19000. In this case, when a user and a character start a new conversation, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or another cloud server, acquires (downloads) a past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create an instruction for a large-scale language model. Specific methods for using past conversation history to create an instruction for a large-scale language model are as described in the figures of the second embodiment, and therefore repeated description will be omitted.

[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character's conversation history up to a predetermined point in time, such as every time a conversation between the user and the character takes place or when the conversation between the user and the character ends, to the second server 19002 or another cloud server. That is, the character conversation device (artificial intelligence response output device 10010) may upload the conversation history with the character to the second server 19002 or another cloud server at a predetermined timing, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or another cloud server and use it to generate an instruction sentence for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) that the user used the previous day and the character conversation device (artificial intelligence response output device 10010) that the user is about to use are different individual devices and can display the same character, and when the user has multiple conversations with the same character between the different individual devices at different times, it is possible to realize a conversation in which the character's memories are pseudo-carried over from the previous conversation, which is more convenient for the user.

[0129] The process described above in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with a character to the second server 19002 or another cloud server, thereby pseudo-taking over the character's memory, is also effective when handling the database 19200 including the conversation histories of multiple characters described in Figures 2H and 2I. In other words, if the database 19200 described in Figure 2I is configured to be uploaded and downloaded to the second server 19002 or another cloud server, not only for one character but for multiple characters, and between different individual devices, when a user has multiple conversations with each of the multiple characters at different times, it is possible to realize a conversation in which the memory of each character is pseudo-taken over from the previous conversation, which is more convenient for the user.

[0130] Example 3 Next, the third embodiment of the present invention is an improvement of the character conversation device (artificial intelligence response output device 10010) and the character conversation system explained in the drawings of the second embodiment. In this embodiment, differences from the second embodiment will be explained, and repeated explanations of the same configurations as those of the second embodiment will be omitted.

[0131] As in Example 2, the character in Example 3 can provide the user with the service of a large-scale language model, which is an artificial intelligence, and can be of assistance to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this example may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.

[0132] An example of a character conversation device and a character conversation system according to a third embodiment of the present invention will be described with reference to Fig. 3A. The character conversation system of the third embodiment is provided with a large-scale language model server 20001 instead of the large-scale language model server 19001 in Fig. 2A, and is connected to the Internet 19000.

[0133] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence, and is a multimodal large-scale language model artificial intelligence that can process not only the natural language text information that the large-scale language model server 19001 could process, but also types of information other than natural language text information.

[0134] Moreover, the AI ​​response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (AI response output device 10010) of the second embodiment, as an example.

[0135] In the third embodiment as well, the AI ​​response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the Internet 19000 using an API.

[0136] The character conversation system of the third embodiment includes a mobile information processing terminal 20010 used by a user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.

[0137] 3B, an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011, which is a touch operation input panel, a control unit 20012, an external power supply input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.

[0138] The display panel 20011 is equipped with a touch operation input sensor and can accept touch operation input by the finger of the user 230. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel. The display panel 20011 may also be referred to as a display unit.

[0139] The communication unit 20020 may be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and a memory, and the control unit controls the display panel 20011 and the communication unit 20020. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000. This allows the mobile information processing terminal 20010 to communicate with various servers connected to the Internet 19000.

[0140] The power supply 20014 converts AC current input from the outside via the external power supply input interface 20013 into DC current and supplies the required DC current to each unit of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Furthermore, the secondary battery 20015 supplies power to each unit requiring power via the external power supply input interface 20013 when power is not supplied from the outside.

[0141] The video signal input unit 20023 is connected to an external video output device and inputs video data. The video signal input unit 20023 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be various USB interfaces.

[0142] The audio signal input unit 20024 is connected to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 20024 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface in which a terminal and a cable are integrated.

[0143] The audio output unit 20021 can output audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 can also output audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured with a speaker. The audio output unit 20021 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard.

[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 20012, which will be described later, may perform voice recognition processing on the generated audio signal to obtain text information from the audio signal.

[0145] The imaging unit 20025 is a camera having an image sensor. A camera may be provided on the front side of the mobile information processing terminal 20010, on the display panel 20011 side, or on the back side of the display panel 20011 side. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.

[0146] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 20016 before product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices, external servers, etc., via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may also be output to external devices, external servers, etc., via the communication unit 20020.

[0147] The video control unit 20017 performs various controls related to the video signal input to the display panel 20011. The video control unit 20017 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be referred to as a video processing unit or image processing unit. The video control unit 20017 controls video switching, such as which video signal to input to the display panel 20011 between the video signal to be stored in the memory 20026 and the video signal (video data) input to the video signal input unit 20023. The video control unit 20017 may also control image processing of the video signal input from the video signal input unit 20023 and the video signal to be stored in the memory 20026. Examples of image processing include scaling, which enlarges, reduces, or transforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.

[0148] The attitude sensor 20018 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each unit connected thereto.

[0149] The nonvolatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the nonvolatile memory 20027 includes, for example, data for various operations to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 20026 stores video data to be displayed on the display panel 20011, data for controlling the device, etc. The control unit 20012 may read various software from the storage unit 20016 and expand and store it in the memory 20026.

[0150] The control unit 20012 controls the operation of each unit connected to it. The control unit 20012 may also work in conjunction with a program stored in the memory 20026 to perform arithmetic processing based on information acquired from each unit within the mobile information processing terminal 20010.

[0151] Next, an example of the operation of the character conversation device (AI response output device 10010) according to the third embodiment of the present invention will be described with reference to Figure 3C. This can also be said to be an example of the operation of a character conversation system including the AI ​​response output device 10010 and the large-scale language model server 20001. In the third embodiment as well, the character conversation device (AI response output device 10010) loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.

[0152] In the second embodiment, the actions performed by the user 230 with respect to the character conversation device (artificial intelligence response output device 10010) were mainly calls made by the user 230 using his / her voice. The character conversation device (artificial intelligence response output device 10010) of the second embodiment performed a series of operations starting from the process of collecting the voice of the user 230 with a microphone. In contrast, the character conversation device (artificial intelligence response output device 10010) of the third embodiment is also capable of performing the series of operations performed by the character conversation device (artificial intelligence response output device 10010) starting from the process of collecting the voice of the user 230 with a microphone, as described in the second embodiment. In addition, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the user 230 can perform an action with respect to the character conversation device (artificial intelligence response output device 10010) by user operation via the operation input unit 1107 of FIG. 1B. Here, examples of the operation input unit 1107 of FIG. 1B include a mouse, a keyboard, and a touch panel.

[0153] Furthermore, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by performing a touch operation by the user that can be detected by the touch operation input sensor of the display unit 10011 of Figure 1B.

[0154] In addition, the user 230 can operate the mobile information processing terminal 20010 and communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting the user's 230 operation input to the character conversation device (artificial intelligence response output device 10010).

[0155] Alternatively, an information storage image such as a two-dimensional code storing information that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as a two-dimensional code captured by the imaging unit 1180, and obtain the information. Alternatively, an image that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180, and obtain the results of the image recognition processing.

[0156] As described above, the character conversation device (artificial intelligence response output device 10010) of the third embodiment has a greater variety of actions that the user 230 can take toward the character conversation device (artificial intelligence response output device 10010) than the character conversation device (artificial intelligence response output device 10010) described in the second embodiment. As a result, the character conversation device (artificial intelligence response output device 10010) of the third embodiment can acquire the results of actions taken by the user 230 other than the user's voice, and generate an instruction sentence (prompt) to be sent to the large-scale language model server 20001 based on the results. As a result, the instruction sentence to be sent to the large-scale language model server 20001 can more preferably include information of a type other than text information in a natural language extracted from the user's voice. Examples of information of a type other than text information in a natural language extracted from the user's voice include images, videos, and sounds.

[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment uses an API to send an instruction to the large-scale language model server 20001. In this embodiment, the instruction may be metadata containing information written in a notation using tags, such as the markup format of a markup language, a notation using predetermined symbols, such as Markdown, or an object notation of a predetermined script, such as JSON. In this embodiment, the instruction may be classified into two types: a setting instruction that stores instructions, such as initial settings, and a user instruction that reflects instructions from a user. Type identification information identifying whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. In this case, the instruction includes text information in a natural language as the main message. Furthermore, in this embodiment, the main message of the instruction may include, in addition to the text information in natural language, a non-natural language information source, such as an image, video, or audio, as information of a type other than the natural language text information. A specific method for including a non-natural language information source in the instruction will be described later.

[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. The large-scale language model server 20001 receives an instruction from a character conversation device (artificial intelligence response output device 10010). Based on the instruction, the multimodal large-scale language model performs inference and generates a response including natural language text information that is the result of the inference. Here, because the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.

[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or sounds stored as the main message in the response. The character operation program of the character conversation device (artificial intelligence response output device 10010) may use a voice synthesis technology to generate a natural language voice as a reply to the user based on the natural language text information extracted from the response, and output the voice from the voice output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.

[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display characters in a natural language that serve as a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on text information in a natural language extracted from the above-mentioned response. At this time, the characters may be displayed together with the character 19051, may be displayed superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.

[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the information of the image of the non-natural language information source extracted from the above-mentioned response. At this time, the image may be displayed together with the character 19051, may be superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.

[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the video information of the non-natural language information source extracted from the above-mentioned response. At this time, the video may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed in place of the video of the character 19051. These specific processes may be executed by the video control unit 1160.

[0163] In addition, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output a voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140, which is a speaker.

[0164] As described above, the character conversation device (artificial intelligence response output device 10010) of FIG. 3C or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 does not require the large-scale language model itself, which requires vast amounts of data and computational resources for learning, to be installed in the character conversation device (artificial intelligence response output device 10010). Furthermore, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via an API. In response to a user's action toward a character, a response based on a non-natural language information source can be provided in addition to a response based on natural language text, enabling a more appropriate conversation to be held.

[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3D . This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3D shows examples of non-natural language information sources, such as natural language text and images, of the main message of an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001, and examples of non-natural language information sources, such as natural language text and images, of the main message of a server response that is the response. In this embodiment, the non-natural language information source can be an image, video, audio, or the like, but FIG. 3D shows an example of an image as the non-natural language information source.

[0166] 3D also shows an exchange of instructions and responses in chronological order, from a setting instruction, a first round of user instructions and their responses, to a second round of user instructions and their responses. The instructions and responses shown in FIG. 3D include non-natural language information source 20061 and non-natural language information source 20062, which were not shown in FIG. 2D of Example 2. In the example of FIG. 3D, both non-natural language information source 20061 and non-natural language information source 20062 are images.

[0167] Here, for ease of explanation, FIG. 3D shows an image of the non-natural language information source 20061 pasted into the instruction. However, there are multiple methods for transmitting or specifying data from the non-natural language information source 20061 in an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these multiple methods or switch between them. An example of each method will be described below.

[0168] The first method for transmitting or specifying non-natural language information source data in a directive is used, for example, when the non-natural language information source to be specified is a non-natural language information source located in a location such as a server connected to a network such as the Internet. A specific example of the first method is to use information such as tags and symbols in the directive to specify a non-natural language information source file located on a network such as the Internet by using location information (such as a URL) on the network such as the Internet and a file name.

[0169] For example, it is a tag that specifies an image in a markup language. <img src=""****”"> You can also specify an image on a network such as the Internet by writing the location information and file name information of the image file in the **** part using the tag. <video src=""****”">You can also specify a video file on a network such as the Internet by writing the location information and file name information of the video file in the **** part using the tag. <audio src=""****”">By using the above and entering the location information and file name information of the audio file in the **** part, audio that exists on a network such as the Internet can be specified. Furthermore, if the notation is JSON, an image that exists on a network such as the Internet can be specified by preparing a key such as img_src and entering the location information and file name information of the image file as the value. For video files and audio files, it is sufficient to prepare the respective keys and values. This specific example of a format is just one example, and other unique formats may also be used. In either case, information that specifies the location information and file name information of the non-natural language information source file can be stored in the directive.

[0170] When information specifying the location information and file name information of a non-natural language information source file is stored in a directive, as in the first method, the directive itself does not need to store the data of the non-natural language information source file itself. Therefore, the amount of data in the directive can be reduced. In the first method, the large-scale language model server 20001 that receives a directive specifying non-natural language information source data simply uses the location information and file name information of the non-natural language information source file stored in the directive to obtain the non-natural language information source file located in a location such as a server connected to a network such as the Internet.

[0171] Here, how location information and file name information are input when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence using the first method will be described. In FIG. 3C, it has been explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are increased beyond the voice of the user 230 compared to the second embodiment. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, and the like, by user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in FIG. 1B.

[0172] Furthermore, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a web browser program and display a GUI of the web browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the web browser program may be accepted by a user operation via the operation input unit 1107 (e.g., a mouse, keyboard, or touch panel) or a user touch operation detectable by a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, video, or audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction. In this case, the web browser program may acquire location information and file name information of the non-natural language information source data and pass them to the character operation program.

[0173] Furthermore, user 230 may operate mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from mobile information processing terminal 20010, thereby inputting location information such as a URL for specifying non-natural language information source data to the character conversation device (artificial intelligence response output device 10010). Alternatively, location information such as a URL for specifying non-natural language information source data, file name information, and the like may be input in a manner such as displaying an information storage image such as a two-dimensional code on display panel 20011 of mobile information processing terminal 20010, performing image recognition processing on an image captured by imaging unit 1180 of character conversation device (artificial intelligence response output device 10010), and acquiring the results of the image recognition processing, as described in FIG.

[0174] Note that the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file is previously present in a location such as a server connected to a network such as the Internet. For example, if non-natural language information source data such as images, videos, or audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) is to be included in the instruction statement, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to the second server 19002 via the Internet 19000, and include in the instruction statement the location information on the Internet (such as a so-called URL) and file name of the non-natural language information source data on the uploaded second server 19002. In this case, the second server 19002 functions as a so-called intermediate server.

[0175] Similarly, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in an instruction statement, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the location information on the Internet (such as a URL) and the file name of the non-natural language information source data on the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include the acquired location information on the Internet (such as a URL) and the file name of the non-natural language information source data uploaded to the second server 19002 in the instruction statement.

[0176] Furthermore, a media server may be constructed within the character conversation device (artificial intelligence response output device 10010) so that the character operation program of the character conversation device (artificial intelligence response output device 10010) can cooperate with memory 1109 and storage unit 1170 to be accessible from other servers via the Internet 19000. In this case, when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction statement using the first method, the character conversation device (artificial intelligence response output device 10010) may store location information on the Internet (such as a URL) indicating the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data in the instruction statement.

[0177] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction sentence is, for example, a method in which the non-natural language information source data itself is simply stored (attached) to the instruction sentence (prompt) and transmitted. Generally, non-natural language information source data such as images, videos, and audio has a larger data volume than text information in natural language. Therefore, in this case, the data volume of the instruction sentence (prompt) itself is larger than that in the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) temporarily stores the non-natural language information source data to be stored (attached) to the instruction sentence (prompt) in the memory 1109, and when transmitting the instruction sentence (prompt), stores (attaches) the non-natural language information source data in the memory 1109 via the communication unit 1132 and outputs the data to the large-scale language model server 20001. The non-natural language information source data itself that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read from the storage unit 1170 and stored in memory 1109.

[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data using instruction sentences.

[0179] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. Therefore, in the first pass of the user instruction shown in the example of Figure 3D, it can obtain images of the swimming pool and poolside, which are non-natural language information source 20061, and text information in natural language, and output the text information in natural language as a response to the first pass of the user instruction as an inference result, as shown in the figure.

[0180] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example of the response to the second round of user instructions in Figure 3D, the large-scale language model server 20001 can include in the response a non-natural language information source 20062 generated by inference from the multimodal large-scale language model and send it to the character conversation device (artificial intelligence response output device 10010). In Figure 3D, the non-natural language information source 20062 is an example of an image in which a circle image is added to the image of the swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 3D, but may be video or audio.

[0181] When the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, a method similar to the first method or the second method in which the character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence can be used.

[0182] Specifically, as a method similar to the first method described above, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The non-natural language information source 20062, such as an image, video, or audio, itself may be stored in the large-scale language model server 20001, or the non-natural language information source 20062 may be transferred to and stored in a second server 19002 functioning as an intermediate server. In either case, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The character conversation device (artificial intelligence response output device 10010) that receives the response may access the large-scale language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction to acquire the non-natural language information source 20062.

[0183] Specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction sentence and use it for various outputs to the user 230.

[0184] According to the operation of the character conversation device (artificial intelligence response output device 10010) and the character conversation system of the third embodiment described above with reference to Fig. 3D, instructions and responses are sent and received between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230 to realize a conversation using non-natural language information such as images, videos, and sounds. This makes it possible to realize a more advanced and natural conversation as shown in each message in Fig. 3D.

[0185] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3E shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.

[0186] 3E shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in FIG. 3D has ended. In the example of FIG. 3E, processing using the conversation history as described in FIGS. 2F, 2G, and 2I of Example 2 is not performed. Therefore, like FIG. 2E of Example 2, FIG. 3E shows a response with content that does not remember at all the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, the user's name, the conversation history, and the like that were included in the setting directive.

[0187] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3F. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3F shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.

[0188] Figure 3F shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in Figure 3D has ended. Here, Figure 3F shows an example in which the method of storing a message explaining the history of past conversations in the setting instruction sentence, which was explained in Figure 2F of the second embodiment, is also applied to the character conversation device (artificial intelligence response output device 10010) of the third embodiment. Specifically, in Figure 3F, the message that is the content of the setting instruction sentence in Figure 3D is stored as a reset message, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.

[0189] The large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. Therefore, non-natural language information source data may have been transmitted or specified in past instructions and responses. Therefore, in the example of Figure 3F, the conversation history message reflects not only the natural language text information in past instructions and responses, but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method for transmitting or specifying non-natural language information source data in the instruction in Figure 3F is similar to the transmission or specification of non-natural language information source data described in Figure 3D, and therefore a repeated description will be omitted.

[0190] In the example of Figure 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself in the instruction statement, and storing (attaching) the non-natural language information source data in the instruction statement without storing (attaching) the non-natural language information source data in the instruction statement. This also applies to the instruction statement of Figure 3F.

[0191] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3G shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.

[0192] Figure 3G shows an example of a series of conversations shown in Figure 3F, from the first user instruction and its response following the first setting instruction to the third user instruction and its response. Figure 3G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 3F, so repeated description is omitted.

[0193] As described above, even when using the large-scale language model server 20001 having a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information according to the third embodiment, even when the user 230 speaks to the character 19051 again to start a new conversation after a series of conversations has ended, if the generation and transmission processes of the setting instruction sentence in Fig. 3F are performed, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation, as shown in Fig. 3G. This is preferable because it allows the user to recognize that the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation are more closely matched.

[0194] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, FIG. 3H uses the example described in FIG. 2H of the second embodiment for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.

[0195] Furthermore, database 20200 for managing character settings and character conversation history, shown in Figure 3H, has the same format as database 19200 shown in Figure 2I of Example 2, and in Figure 3H, only the differences from database 19200 shown in Figure 2I will be explained. Also, the content of the character "Koto" in the database will be explained, and the content of other characters will be omitted.

[0196] As described above, the large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information. Therefore, both the instruction statements from the character conversation device (artificial intelligence response output device 10010) and the responses from the large-scale language model server 20001 include not only natural language text information but also the transmission or specification of non-natural language information source data. Therefore, the database 20200 shown in FIG. 3H records, in the conversation history data, not only the natural language text information included in these instruction statements and responses, but also the transmission or specification of non-natural language information source data. The specific method for transmitting or specifying non-natural language information source data in the conversation history recording is the same as the transmission or specification of non-natural language information source data described in FIG. 3D, and therefore a repeated description will be omitted.

[0197] In the example of FIG. 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself to the instruction statement and not storing (attaching) the non-natural language information source data to the instruction statement. This is also true for the conversation history of FIG. 3H. However, in the conversation history of FIG. 3H, if the method of specifying non-natural language information source data is to specify location information and file name information of a non-natural language information source file on a server on a network such as the Internet (such as the second server 19002 functioning as an intermediate server or another cloud server), there is a possibility that the non-natural language information source file on the server may be deleted over a long period of time in the conversation history. In this case, even if the location information and file name information are used, the non-natural language information source file cannot be retrieved at a later date, and information in the conversation record may be lost.

[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts the instruction sentence and the response message into a conversation history and records them, it can use the location information and file name information to obtain the non-natural language information source file itself specified in the instruction sentence and the response from a server on the network and store it in the storage unit 1170. Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) can rewrite the location information and file name of the non-natural language information source file to Internet location information (such as a URL) indicating the location of the non-natural language information source file within the media server built within the character conversation device (artificial intelligence response output device 10010), and then record the non-natural language information source in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself deletes the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the conversation record information, which is more suitable for preserving the conversation record.

[0199] By using the database of Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between multiple character candidates to display on the display unit 10011, the user will feel less uncomfortable when conversing with each character, and will be able to share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.This effect can be achieved as shown in Figure 2I of Example 2, and can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information.

[0200] In addition, in the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, the large-scale language model server 20001 uses a multimodal large-scale language model artificial intelligence that can process not only natural language text information but also non-natural language information other than natural language text information.

[0201] Here, an API is used for communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001. In a multimodal large-scale language model, the API usage fee may be charged according to the number of processed natural language text information, which are units of words that separate sentences called tokens, as well as the amount of data from non-natural language information sources.

[0202] Therefore, in order to provide users with a character conversation service by the character conversation system according to this embodiment at a lower cost, the following modified example may be used.

[0203] In a first variation, the database conversation history record of FIG. 3H also records the transmission or specification of non-natural language information source data. However, the character and the user exchange conversations about the non-natural language information source data using text information in natural language, and the content of the conversations is recorded as text information in natural language. Even if the transmission or specification of the non-natural language information source data is omitted from the database conversation history record of FIG. 3H, the conversation itself about the non-natural language information source data will still be recorded as text information in natural language to some extent. Therefore, if a certain amount of information reduction is acceptable, the transmission or specification of the non-natural language information source data may be omitted from the database conversation history record of FIG. 3H. In this case, the transmission or specification of the non-natural language information source data is also omitted from the conversation history message of the setting instruction sentence of FIG. 3F. This reduces the amount of data from non-natural language information sources communicated using the API.

[0204] Next, as a second modification, in the recording of the conversation history of the database in FIG. 3H, natural language text information explaining the content of the non-natural language information source data is recorded instead of transmitting the non-natural language information source data or recording specified information. The natural language text information explaining the content of the non-natural language information source data may be obtained, for example, by starting a conversation between the large-scale language model of the large-scale language model server 20001 and a character conversation device (artificial intelligence response output device 10010) separately from the conversation as a character, and having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the natural language text information may be obtained by having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit through a conversation with another large-scale language model of another server that is available at a lower cost than the large-scale language model of the large-scale language model server 20001. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be used as natural language text information explaining the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is a tag in a markup language. <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">This is the text information written in the **** part of the above.

[0205] Furthermore, in the case of JSON notation, in an object that is stored in association with the location information and file name information of the non-natural language information source data, which are keys and values ​​indicating the location information of the non-natural language information source data, a key corresponding to the alternative text can be further linked and stored as a value that is the alternative text data itself.

[0206] In this case, too, the transmission of the non-natural language information source data or the recording of the specified information can be omitted in the recording of the conversation history in the database of Fig. 3H, and the transmission of the non-natural language information source data or the specified information is also omitted from the conversation history message of the setting instruction sentence of Fig. 3F, thereby reducing the amount of data from the non-natural language information source communicated using the API.

[0207] Next, a third variation is an example in which, at the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data is not stored in the user instructions, but is replaced with natural language text information explaining the content of the non-natural language information source data. For example, in the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data 20061 may be replaced with the user instructions, and a description such as "This image is of a swimming pool with seats and parasols by the pool. There is water in the swimming pool. There are drinks on the table next to the seats" may be stored as natural language text information. In this case, the description may be obtained by having another large-scale language model on another server, which is more inexpensive to use than the large-scale language model on the large-scale language model server 20001, explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the description may be obtained from a server of various other services that can obtain summaries and descriptions of non-natural language information source data such as images, videos, and audio. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be natural language text information that explains the content of the non-natural language information source data.

[0208] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3I. The example of FIG. 3I shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which is described in each of FIGS. 3A to 3H, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063 of natural language information source data, image 10064 of non-natural language information source data, and / or video 10065 of non-natural language information source data, which are responses from the large-scale language model, together with a video of a character 19051, on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed superimposed on top of the video of the character 19051, as shown in FIG. 3I.

[0209] Furthermore, text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on the video of character 19051. The display in FIG. 3I is an example, but for example, if user 230 adjusts the volume of the audio output of audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating operation input unit 1107 or the touch operation input sensor of display unit 10011, user 230 will not be able to hear the response from the large-scale language model by audio. Therefore, in this case, control unit 1110 may control to start a display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in FIG. 3I.

[0210] In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more preferably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the image of the character 19051, may be manually switched on / off by the user 230 operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. According to the display example of FIG. 3I, the character conversation device (artificial intelligence response output device 10010) that supports multimodal display can more preferably output a response from the large-scale language model.

[0211] According to the character conversation device and character conversation system of Example 3 described above, in addition to the effects of the character conversation device and character conversation system of Example 2, it is possible to provide users with a more advanced conversation experience that includes information in non-natural languages ​​in addition to information in natural languages ​​by using a multimodal large-scale language model. Furthermore, according to the character conversation device and character conversation system of Example 3, it is possible to provide users with a character conversation service at a lower cost.

[0212] In the above description of the third embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 20001 is used as the large-scale language model. In contrast to this, the character conversation device (artificial intelligence response output device 10010) may be equipped with the local LLM processing unit 10028 shown in FIG. 1B and may use the multimodal large-scale language model held by the local LLM processing unit 10028. In this case, the multimodal large-scale language model held by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model held by the large-scale language model server 20001.

[0213] In this case, in the above description of the third embodiment, the multimodal large-scale language model held by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model held by the local LLM processing unit 10028 of the character conversation device (the AI ​​response output device 10010). In this case, too, a more sophisticated conversation experience that includes non-natural language information in addition to natural language information can be provided to the user using the multimodal large-scale language model. Note that, when the multimodal large-scale language model held by the local LLM processing unit 10028 is used instead of the multimodal large-scale language model held by the large-scale language model server 20001, there is less need to consider usage fees based on the number of processed tokens and the amount of data in the non-natural language information source. However, even with the multimodal large-scale language model held by the local LLM processing unit 10028, the number of processed tokens and the amount of data in the non-natural language information source can be reduced, thereby reducing the consumption of resources such as power required for inference. In this case, a character conversation service that consumes less power can be provided to the user.

[0214] The configuration described in Example 2 for uploading and downloading the conversation history with a character or the database data including the conversation history with a character to the second server 19002 or another cloud server can also be used in the example using a multimodal large-scale language model described in Example 3. In this case, too, when a user has multiple conversations with each of the multiple characters at different times between different individual devices for one character or multiple characters, it is possible to realize a conversation in which the memory of each character is pseudo-carried over from the previous conversation, which is more convenient for the user.

[0215] Example 4 Next, Example 4 of the present invention is an improvement of the AI ​​response output device 10010, the character conversation device, or these systems described in the drawings of Example 2 or Example 3. In this example, differences from Example 2 or Example 3 will be described, and repeated explanations of the same configurations as those examples will be omitted.

[0216] As in the above-described embodiments, the AI ​​response output device 10010 may be referred to as an AI response output device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI ​​response output device 10010 and the large-scale language model server may be referred to as an AI response output system, an AI assistant system, an AI assistant display system, or an AI interface system.

[0217] An example of operation using a database in a character conversation device (artificial intelligence response output device 10010) according to a fourth embodiment of the present invention will be described with reference to Fig. 4A. The database according to the fourth embodiment shown in Fig. 4A is an extension of the database described with reference to Fig. 2I or Fig. 3I. Specifically, the database shown in Fig. 4A assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores initial setting instructions and conversation histories corresponding to each user and character in the database.

[0218] In the example of Fig. 4A, for user 1 with a user ID of 1, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored.

[0219] These initial setting instruction sentences and conversation history data are stored as separate data in different areas for each combination of user and character. For the sake of explanation, in Fig. 4A, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (AI response output device 10010) uses the initial setting instruction sentences and conversation history stored in different areas for each combination of user and character based on the user currently using (logging in to) the character conversation device (AI response output device 10010) or its system, thereby making it possible to more appropriately maintain the consistency of the character's personality and the continuity of memory for each different user.

[0220] Specifically, consider a situation in which User 1 has previously had a conversation with character Tom using a character conversation device (AI response output device 10010), and User 2 is unaware of that conversation, and then User 2 subsequently has a conversation with character Tom. In this case, if AI response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify users, the response output from AI response output device 10010 will be based on a conversation history that User 2 does not remember, and the conversation between User 2 and the character of AI response output device 10010 may become inconsistent.

[0221] In contrast, even in a similar situation, if the database shown in Figure 4A is used, the control unit 1110 of the character conversation device (AI response output device 10010) identifies users by ID, stores the initial setting instructions and conversation history in a different area for each user, and uses the initial setting instructions and conversation history stored in a different area for each user to generate an AI response. As a result, the initial setting instructions and conversation history used to generate an AI response for each user are based on the operation or conversation history of that user and are managed separately from the operation or conversation history of other users. This makes it possible to more appropriately maintain consistency in the conversation history between each user and each character of the AI ​​response output device 10010.

[0222] The database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in the storage unit 1170 of the AI ​​response output device 10010 and used by the control unit 1110. Furthermore, without being limited to this, the database of initial setting directives and / or conversation histories may be stored in a server on the network. For example, if the AI ​​response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 in generating an AI response, the database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in these servers themselves. In this way, it is possible to omit the process of transmitting the initial setting directives and conversation histories again from the AI ​​response output device 10010 to these servers by including them in directives, thereby reducing the number of transmission tokens for use of the large-scale language model.

[0223] When storing the database of the initial setting instruction sentence and / or the conversation history described in FIG. 4A, the AI ​​response output device 10010 can transmit the user ID, the character ID, and the user instruction sentence for the subsequent conversation to these servers. The large-scale language models on these servers can use the user ID and character ID obtained from the AI ​​response output device 10010 to obtain the corresponding initial setting instruction sentence and the conversation history from the database of the initial setting instruction sentence and / or the conversation history of FIG. 4A. The large-scale language models on these servers can perform inference using the initial setting instruction sentence and the conversation history, and the user instruction sentence for the subsequent conversation transmitted from the AI ​​response output device 10010, to generate an AI response and transmit it to the AI ​​response output device 10010. In this way, the effect of more appropriately maintaining the consistency of the character's personality and the continuity of memory for each different user can be obtained while saving the number of transmitted tokens when using the large-scale language model.

[0224] Next, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) according to the fourth embodiment of the present invention will be described with reference to Fig. 4B. The database according to the fourth embodiment shown in Fig. 4B is an extension of the database described in Fig. 1C or Fig. 2L. Specifically, the database shown in Fig. 4B assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data on standard response phrases corresponding to each user and character in the database.

[0225] In the example of Fig. 4B, for user 1, whose user ID is 1, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3. In addition, for user 2, whose user ID is 2, and user 3, whose user ID is 3, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3.

[0226] These standard response phrase data are stored as separate data in different areas for each user-character combination. For ease of explanation, in FIG. 4B, the data stored in each area is represented as standard response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, standard response phrase data 101 is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 1: Koto shown in FIG. 2L. Data 201 in FIG. 4B is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 2: Tom shown in FIG. 2L.

[0227] Data 301 in FIG. 4B is stored as a database such as a table corresponding to the fixed response phrases corresponding to condition numbers 1 to 7 of character 3: Necco shown in FIG. 2L. Data 102, 202, and 302 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 2. Data 103, 203, and 303 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 3. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses fixed response phrase data stored in different areas for each combination of user and character, based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.

[0228] In this way, even the same character can respond with different template responses for each user. That is, even for the same character, it may be more appropriate to vary the content of the template responses depending on the relationship between the character and the user. For example, depending on the relationship between the character's set age and the age of the user registered in the AI ​​response output device 10010 or the system, the user may be older, the same age, or younger than the character. In this case, different content for the template responses of the character to older users, the template responses of the character to users of the same age, and the template responses of the character to younger users can make the conversation between the user and the character more appropriate or natural. That is, by performing operations using the database of FIG. 4B and varying the content of the template responses depending on the relationship between the character and the user, it is possible to create a more appropriate or natural conversation.

[0229] 4B described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI ​​response output device 10010. However, the response template database (response template DB) shown in FIG. 4B may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI ​​response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).

[0230] According to the character conversation device and character conversation system of Example 4 described above, it is possible to produce more suitable or more natural conversation depending on the relationship between the character and the user, the conversation history, etc.

[0231] <Example 5> Next, a fifth embodiment of the present invention is an improvement of the AI ​​response output device 10010 or the AI ​​response output system described in the drawings of the first, second, and third embodiments. Specifically, this is an example in which the response generation process of the AI ​​response output device 10010 is switched from response generation process using a large-scale language model on a network to response generation process using a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI ​​response output device 10010, or response generation process using a fixed response phrase database. In this embodiment, differences from these embodiments will be described, and repeated description of configurations similar to those of these embodiments will be omitted.

[0232] As in the above-described embodiments, the AI ​​response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI ​​response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.

[0233] An example of response generation process switching processing in the AI ​​response output device 10010 of the fifth embodiment of the present invention will be described using FIG. 5A. The table in FIG. 5A shows examples 1 to 9 of response generation process switching processing in the AI ​​response output device 10010. In the table in FIG. 5A, the column "Switching Overview" shows an overview of the switching processing for each example. The column "State before switching of LLM on the network (API connected LLM)" shows the state before response generation processing by a large-scale language model on the network (large-scale language model connected using an API), such as the large-scale language model provided in the large-scale language model server 19001 in FIG. 1 and the multimodal large-scale language model provided in the large-scale language model server 20001, is switched to another response generation process. The column "Switching Occurrence Condition" shows the conditions under which switching processing of the response generation process occurs. The column "Switching destination from LLM on the network (API-connected LLM)" indicates the switching destination to which the response generation process of the AI ​​response output device 10010 is switched from a large-scale language model on the network (a large-scale language model connected using an API), such as a large-scale language model provided in the large-scale language model server 19001 and a multimodal large-scale language model provided in the large-scale language model server 20001. When the condition indicated in "Switching occurrence condition" occurs in the state of "State before switching of LLM on the network (API-connected LLM)" shown in Figure 5A, the control unit 1110 of the AI ​​response output device 10010 may perform control to switch to the large-scale language model, database, or correspondence indicated in "Switching destination from LLM on the network (API-connected LLM)."

[0234] Below, each example shown in the table of FIG. 5A will be described. Example 1, as shown in the "Switching Overview," is an example in which switching is performed depending on the network connection status of the AI ​​response output device 10010. In Example 1, the "state before switching of the LLM (API-connected LLM) on the network" indicates that the network connection status of the AI ​​response output device 10010 is connectable. Here, in Example 1, the "switching occurrence condition" indicates "when network connection becomes unavailable." That is, this means when the connection via the network between the AI ​​response output device 10010 and a large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable. Specifically, the connection failure may be due to a communication failure state on the connection path from the AI ​​response output device 10010 to the Internet 19000. Alternatively, the connection failure may be due to a communication failure state on the Internet 19000. Alternatively, the connection failure may be due to a situation in which the large-scale language model on the network (a large-scale language model connected using an API) itself cannot connect to the Internet 19000. Also, in Example 1, "local LLM" is shown as the "switching destination from LLM on the network (API-connected LLM)." This specifically means that switching processing is performed to response generation processing by the local LLM processing unit 10028 of the AI ​​response output device 10010. That is, in Example 1, even if connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and response generation processing by the large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable, response generation processing is switched to response generation processing by the local LLM processing unit 10028 of the AI ​​response output device 10010. This makes it possible to continue response generation processing using the large-scale language model, despite differences in performance as a large-scale language model.

[0235] Next, Example 2 in FIG. 5A will be described. In Example 2, the "switching destination from the networked LLM (API-connected LLM)" in Example 1 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 2, if a connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and the response generation process using the large-scale language model on the network (a large-scale language model connected using an API) cannot be used, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response through simpler processing and output the response to the user.

[0236] Next, Example 3 of FIG. 5A will be described. In Example 3, the "switching destination from LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "no-response handling." The "no-response handling" means that even if a user input requesting a response from a large-scale language model is received from the user via the touch panel, microphone 1139, or operation input unit 1107, no response to this input is generated, or even if a user input requesting a response from a large-scale language model is received, no response to this is output. In other words, Example 3 makes it easier to handle a situation where connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason, making it impossible to use the response generation process by the large-scale language model on the network (a large-scale language model connected using an API).

[0237] Next, Example 4 of FIG. 5A will be described. As shown in the "Switching Overview," Example 4 is an example in which switching is performed due to a response delay of an LLM on the network. In Example 4, the "state before switching an LLM on the network (API-connected LLM)" indicates a state in which a response from an LLM on the network is obtained within a predetermined time. Here, in Example 4, the "switching occurrence condition" indicates a case in which a response from an LLM on the network is not obtained within the predetermined time and exceeds the predetermined time. Also, in Example 4, the "switching destination from an LLM on the network (API-connected LLM)" indicates a "local LLM." The "local LLM" of the switching destination is the same as in Example 1, so repeated explanation will be omitted. That is, in Example 4, even if the response from an LLM on the network (a large-scale language model connected using an API) exceeds the predetermined time for some reason and the response generation process by the LLM on the network (a large-scale language model connected using an API) cannot be used smoothly, the response generation process is switched to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This makes it possible to continue the response generation process using a large-scale language model, even if there are differences in performance as a large-scale language model.

[0238] Next, Example 5 in FIG. 5A will be described. Example 2 is an example in which the "switching destination from the networked LLM (API-connected LLM)" in Example 4 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 5, if for some reason a response from the networked LLM (large-scale language model connected using an API) exceeds a predetermined time and the response generation process using the networked LLM (large-scale language model connected using an API) cannot be used smoothly, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response using simpler processing and output the response to the user.

[0239] Next, Examples 6 to 9 of FIG. 5A will be described. As shown in "Switching Overview," Examples 6 to 9 are examples in which switching is performed when the upper limit of API usage or usage fee is reached. As described in Example 2, providers of large-scale language models often collect the costs used to train the large-scale language model from terminal users as API usage fees for the terminal. In such cases, with natural language models, API usage fees are often charged based on the number of processing of units of words that separate sentences, called tokens. Here, various methods of charging and limiting API usage fees are conceivable. One possible example is to define the upper limit of the amount of large-scale language model usage services that a user can receive under normal circumstances using the number of tokens processed.

[0240] In this case, the user may be able to receive services using large-scale language models at a specified API usage fee until the usage volume (or corresponding usage fee) is reached, and once the upper limit of the usage volume (or corresponding usage fee) is reached, certain restrictions may be imposed, such as the user being unable to receive services using large-scale language models in their normal state (performance or frequency).

[0241] Examples 6 to 9 in FIG. 5A are examples of response generation process switching control by the control unit 1110 of the AI ​​response output device 10010 when such restrictions occur in a service using a large-scale language model. Specifically, in Example 6, the "state before switching the on-network LLM (API-connected LLM)" is a state in which the API usage volume and API usage fee have not reached a predetermined upper limit. This means that the usage volume of the on-network LLM (API-connected LLM) has not reached a predetermined upper limit. In this case, the user can use the on-network LLM (API-connected LLM) in a normal state.

[0242] Here, in Example 6, the "switching trigger condition" is when the API usage volume or API usage fee reaches a predetermined upper limit. This means when the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit. Also, in Example 6, the "switching destination from the network-based LLM (API-connected LLM)" is a second LLM on the network that is different from the LLM (which may be referred to as the first LLM) used in normal operation. An example of a second LLM on the network is an LLM with a lower fee than the first LLM used in normal operation. Since it is a lower-fee service, the performance of the second LLM may be lower than that of the first LLM. Even in this case, there is still a significant advantage if large-scale language models can be used inexpensively even after the usage volume / fee limit of the first LLM is reached.

[0243] Next, Example 7 in Figure 5A will be described. In Example 7, the "switching destination from the network-based LLM (API-connected LLM)" in Example 6 is changed from a second LLM on a network different from the LLM (which may be referred to as the first LLM) used in the normal state to a "local LLM." In Example 7, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit, it is possible to continue performing response generation processing using a large-scale language model by switching to response generation processing using a local LLM that is not subject to restrictions such as the usage volume of the network-based LLM, the API usage volume, or the API usage fee.

[0244] Next, Example 8 in FIG. 5A will be described. In Example 8, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is similar to the process described in FIG. 1C, FIG. 2L, or FIG. 4B, and therefore will not be described again. In Example 8, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the networked LLM (API-connected LLM) reaches a predetermined upper limit, the process switches to response generation using a preparatory response phrase database, which is not subject to restrictions such as the usage volume of the networked LLM, the API usage volume, or the API usage fee. This makes it possible to generate a response using simpler processing and output the response to the user.

[0245] Next, Example 9 in FIG. 5A will be described. In Example 9, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "no-response handling." The "no-response handling" refers to a handling in which a response to the user is not generated or a response to the user is not output. Example 9 makes it easier to handle a situation in which response generation processing using a large-scale language model on the network (a large-scale language model connected using an API) is unavailable because the API usage or API usage fee has reached a predetermined upper limit, i.e., the usage of the networked LLM (API-connected LLM) has reached a predetermined upper limit.

[0246] According to the switching control of the response generation process of the AI ​​response output device 10010 shown in Examples 1 to 9 of Figure 5A as described above, even in situations where the response generation process using an LLM (large-scale language model connected using an API) on the network cannot be used as usual, more appropriate switching or response can be made according to each situation.

[0247] 5A may be performed by combining a plurality of examples. For example, the switching control of Examples 1 to 3 may be combined with any of the controls of Examples 4 to 9. Similarly, the control of Example 4 or Example 5 may be combined with any of the controls of Examples 1 to 3 or Examples 6 to 9. Similarly, the control of Examples 6 to 9 may be combined with any of the controls of Examples 1 to 5.

[0248] Next, an example of display of an AI assistant or a character when the AI ​​response output device 10010 of the fifth embodiment is configured as an AI assistant device or a character conversation device will be described with reference to FIGS. 5B to 5D.

[0249] First, Figure 5B is an example of the display of an AI assistant or character on the AI ​​response output device 10010 when performing the switching control of Example 3 in Figure 5A. In the example of Figure 5B, the display state of the AI ​​assistant or character is changed depending on whether the network connection status of the AI ​​response output device 10010 is network connectable or network unconnectable. The network connectable and network unconnectable states of the AI ​​response output device 10010 are as explained in Figure 5A, so a repeated explanation will not be given.

[0250] In the example of FIG. 5B, the AI ​​response output device 10010 (1) displays the AI ​​assistant or character in a normal, awake state when a network connection is possible, but (2) displays the AI ​​assistant or character in a "sleeping" state when a network connection is not possible. In the switching control of example 3 of FIG. 5A, when the AI ​​response output device 10010 cannot connect to the network, it does not generate or output a response even if a command is input from the user. In this case, if the AI ​​assistant or character displayed by the AI ​​response output device 10010 is in a normal, awake state, the user will feel uncomfortable. However, if the AI ​​assistant or character displayed by the AI ​​response output device 10010 is displayed in a sleeping state, the user will understand that "the AI ​​assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user.

[0251] In the case of Figure 5B(2), it is desirable that the user understand that "the reason the AI ​​assistant or character is not responding is because it is asleep" before making a user input requesting a response using a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI ​​response output device 10010. Therefore, it is desirable that the start timing of the state in Figure 5B(2) where the AI ​​assistant or character is displayed in a "sleeping" state when network connection is not possible is immediately after the control unit 1110 of the AI ​​response output device 10010 determines that network connection is not possible, before the user makes a user input requesting a response using a large-scale language model.

[0252] Next, as another display example, the display example of FIG. 5C will be described. The display example of FIG. 5C is an example in which the display state of the AI ​​assistant or character is changed depending on the state of the "switching destination from the LLM on the network (API-connected LLM)" in the table during the switching control of FIG. 5A. Specifically, FIG. 5C shows (1) a display example of the AI ​​assistant or character in a state in which the AI ​​response output device 10010 can connect to a large-scale language model on the network (a large-scale language model connected using an API) and is able to use a response generation process using the large-scale language model on the network (referred to as the normal state in this figure); (2) a display example of the AI ​​assistant or character in a state in which the AI ​​response output device 10010 has switched to a response generation process using an LLM or a template response database with lower performance than the large-scale language model on the network (a large-scale language model connected using an API); and (3) a display example of the AI ​​assistant or character in a state in which the AI ​​response output device 10010 has switched to the no-response mode described in FIG. 5A.

[0253] In the example of FIG. 5C , for example, (1) when the AI ​​response output device 10010 is in a "normal state," the AI ​​response output device 10010 displays the AI ​​assistant or character in a state where there are no particular problems. Note that the "normal state" in FIG. 5C may be considered a state other than states (2) and (3). Also, for example, (2) when the AI ​​response output device 10010 has switched to a response generation process using an LLM or a response template database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), the AI ​​response output device 10010 displays the AI ​​assistant or character in a "sleepy" state. Note that "displaying the AI ​​assistant or character in a "sleepy" state" may also be expressed as "a display indicating that the AI ​​assistant or character is feeling drowsy."

[0254] The response generation process (2) has lower performance than the response generation process using a large-scale language model on the network (a large-scale language model connected using an API) in the normal state (1). Therefore, by displaying the AI ​​assistant or character in a "sleepy" state, it is possible to implicitly convey to the user that the response performance of the AI ​​assistant or character is low. This makes it possible to further reduce the sense of discomfort felt by the user due to a low-performance response. Note that the switching conditions under which the AI ​​response output device 10010 switches to response generation processing using an LLM or a fixed response phrase database, which has lower performance than the large-scale language model on the network (a large-scale language model connected using an API), are as described in FIG. 5A, and therefore a repeated explanation will be omitted.

[0255] 5C(2), it is desirable to implicitly inform the user that the response performance of the AI ​​assistant or character is low before the user makes a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI ​​response output device 10010. Therefore, it is desirable that the start timing of the state in FIG. 5C(2) where the AI ​​assistant or character is displayed in a "sleepy" state be immediately after the AI ​​response output device 10010 switches to a response generation process using an LLM or a template response database, which has lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user makes a user input requesting a response from a large-scale language model.

[0256] Also, for example, in (3) the state in which the AI ​​response output device 10010 has switched to the no-response mode described in FIG. 5A, the AI ​​response output device 10010 displays the AI ​​assistant or character in a "sleeping" state. As also described in FIG. 5B, by displaying the AI ​​assistant or character displayed by the AI ​​response output device 10010 in a "sleeping" state, the user can understand that "the AI ​​assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user. Note that the conditions under which the AI ​​response output device 10010 switches to the no-response mode described in FIG. 5A are the same as those described in Example 3 or Example 9 of FIG. 5A, and therefore a repeated explanation will be omitted. Note that in the case of FIG. 5C(3), it is desirable for the user to understand that "the AI ​​assistant or character is not responding because it is sleeping" before making a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI ​​response output device 10010. Therefore, it is desirable that the timing for starting the state (3) in Figure 5C in which the AI ​​assistant or character is displayed in a "sleeping" state be immediately after the AI ​​response output device 10010 switches to the no-response response described in Figure 5A, before the user input requesting a response from a large-scale language model.

[0257] 5C, the AI ​​response output device 10010 displays a display that implicitly reflects a change in the state of the AI ​​assistant or character without directly providing the user with a technical explanation of the state of the AI ​​response output device 10010 related to the response generation process. This can reduce the sense of discomfort felt by the user compared to when a technical explanation of the state of the AI ​​response output device 10010 related to the response generation process is directly provided to the user. Furthermore, this can reduce the sense of discomfort felt by the user compared to when the display state of the AI ​​assistant or character remains the same as its normal state despite a change in the state of the AI ​​response output device 10010 related to the response generation process.

[0258] However, some users may wish to know a more precise explanation of the technical state of each state. Therefore, a display example for such users will be described with reference to FIG. 5D. Among the rows of the table shown in FIG. 5D, the rows for explaining the device state and display state are identical to those in FIG. 5C, and therefore, repeated explanations will be omitted. Furthermore, the display example of the AI ​​assistant or character shown in the row for the display example of the AI ​​assistant or character is almost identical to that in FIG. 5C, except that a question mark (?) is displayed in the display example. The question mark (?) is a mark that the user operates when requesting an explanation from the AI ​​response output device 10010, and may also be referred to as a help mark.

[0259] In the example of FIG. 5D , when a user selects the question mark (?) through a user operation, such as via the touch panel of the operation input unit 1107 or the display unit 10011 in FIG. 1B , the display of the AI ​​assistant or character on the AI ​​response output device 10010 changes to the display example shown in the row of the display example after user operation. Specifically, regardless of whether the device state is (1), (2), or (3), an explanation of the technical state of each state is displayed. For example, in the example of FIG. 5D , if the device state is (1) normal, a display explaining that the device is in a normal state with no particular technical limitations can be displayed, such as "Normal state." Furthermore, if the device state is (2) using a low-performance LLM or a standard response phrase database, a display technically explaining the low-performance state can be displayed, such as "Low-performance mode." This display can also be considered a display explaining the reason why the AI ​​assistant or character is displaying a "sleepy" state.

[0260] In this case, a more detailed technical explanation may be provided. Specifically, a message such as "Low-performance LLM usage mode" or "Canned response mode" may be displayed. If the device status is (3) unresponsive, a message such as "Network connection unavailable" may be displayed, providing a technical explanation of the reason for switching to unresponsive mode. If the reason for switching to unresponsive mode is that a response from an LLM (a large-scale language model connected via an API) on the network exceeds a specified time, a message such as "Response from LLM is delayed" may be displayed. If the reason for switching to unresponsive mode is that the network LLM usage, API usage, or API usage fee has reached its limit, a message such as "LLM usage limit reached," "API usage limit reached," or "API usage fee has reached a specified amount" may be displayed. These messages may be considered to explain the reason why the AI ​​assistant or character is displayed in a "sleeping" state.

[0261] According to the display example of FIG. 5D described above, even if there are technical constraints in the response generation process in the AI ​​response output device 10010, first, instead of providing a direct explanation to the user, the state of the device is implicitly indicated by a change in the display state of the AI ​​assistant or character, thereby further reducing the sense of discomfort felt by the user. This display is more suitable for users who do not need technical explanations. Furthermore, by displaying an operation mark to explain the technical state, a display is provided to users who operate the mark that technically explains the state of the response generation process in the AI ​​response output device 10010 (normal state or state with technical constraints). This makes it possible to provide a more suitable display for users who want to know the technical state accurately.

[0262] In the examples of Figures 5B, 5C, and 5D, a "sleeping" state is shown as an example of the display state of the AI ​​assistant or character when the AI ​​assistant or character is "unresponsive," but this is only an example and the embodiment is not limited to this. Instead of the "sleeping" state, another display state that implies a situation where the AI ​​assistant or character is unable to respond, such as "taking a break," may be used. In the examples of Figures 5C and 5D, a "sleepy" state is shown as an example of the display state of the AI ​​assistant or character when a low-performance LLM or a standard response phrase database is being used, but this is only an example and the embodiment is not limited to this. Alternatively, another display state that implies that the AI ​​assistant or character has low response performance, such as "hungry," may be used.

[0263] According to the AI ​​response output device and the AI ​​response output system according to the fifth embodiment described above, it is possible to more appropriately switch the response generation process used by the AI ​​response output device depending on the connection state between the large-scale language model on the network and the AI ​​response output device, the response delay state from the large-scale language model on the network, the usage amount of the large-scale language model on the network, etc. Furthermore, when the AI ​​response output device according to the fifth embodiment is configured as an AI assistant device or a character conversation device, it is possible to perform a display that is less strange to the user.

[0264] Example 6 Next, Example 6 of the present invention is an improvement of the AI ​​response output device 10010 or the AI ​​response output system described in the drawings of Examples 1 to 5. Specifically, this is an example in which the response generation process of the AI ​​response output device 10010 is more suitably combined with a response generation process using a large-scale language model on a network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI ​​response output device 10010, and a response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated explanations of configurations similar to those examples will be omitted.

[0265] As in the above-described embodiments, the AI ​​response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI ​​response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.

[0266] An example of a response generation process in the AI ​​response output device 10010 according to the sixth embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 shows an example of a flowchart of the response generation process in the AI ​​response output device 10010 according to the sixth embodiment of the present invention. Specifically, a time axis in which time progresses from top to bottom, a processing flow, and a response output example are shown. The response shown in the response output example may be output via display on the display unit 10011 of the AI ​​response output device 10010 or audio output by the audio output unit 1140.

[0267] In the example of FIG. 6, first, at time t0, a user input is made via the touch panel, microphone 1139, or operation input unit 1107 of the AI ​​response output device 10010, requesting a response based on a large-scale language model, and the control unit 1110 of the AI ​​response output device 10010 acquires the user input (step 600). Next, at time t1, the control unit 1110 starts preparations for response output using a fixed response phrase database stored in the storage unit 1170, and starts response output using the fixed response phrase database (step 601). In the example of FIG. 6, response output using the fixed response phrase database starts at time t2, and as shown in the figure, the fixed response is being output but has not yet been completed. "Good morning" in the figure indicates the output of part of the sentence that continues "Good morning..."

[0268] At time t3, before the response output using the fixed response phrase database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and transmits the generated instruction sentence to a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI ​​response output device 10010, thereby starting a request for a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 starts acquiring a response from the large-scale language model (step 603).

[0269] An example of a response output in which the response output using the fixed response phrase database is completed at time t5 is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is the ____ day of the month, isn't it?" is completed at time t5 using the fixed phrases stored in the fixed response phrase database and date information stored in memory. Here, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 has already started acquiring a response from the large-scale language model. Therefore, at time t6, which follows time t5 when the response output display using the fixed response phrase database is completed, the control unit 1110 starts outputting a response from the large-scale language model following the response output using the fixed response phrase database (step 604). Thereafter, at time t7, a response from the large-scale language model is output following the response output using the fixed response phrase database. When the response output from the large-scale language model is completed, the response output according to the processing flow shown in FIG. 6 is completed (step 605).

[0270] Next, the effect of the processing flow shown in Fig. 6 of the present invention will be described. Processing a large-scale language model requires a large amount of computational resources. Generally, even if inference, which requires fewer computational resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several to several tens of seconds from when the control unit starts to request a response from the large-scale language model until it can obtain a response from the large-scale language model. This period corresponds to the period from time t3 to time t4 shown in Fig. 6. Furthermore, from time t0, when a user input is made, until time t4, the control unit 1110 is unable to obtain a response output from the large-scale language model, and therefore is unable to output a response from the large-scale language model to the user.

[0271] 6, there is no start of preparation for response output using the fixed response phrase database and no start of response output using the fixed response phrase database, the user may continue to wait for a period of several seconds to more than ten seconds from time t0 when the user input is made to time t4 without receiving a response from the AI ​​response output device 10010. For example, when the AI ​​response output device 10010 is configured as an AI assistant device or a character conversation device, the waiting time may give the user a sense of discomfort.

[0272] In contrast, in the processing flow according to the sixth embodiment of the present invention shown in Fig. 6, the control unit 1110 starts the process of outputting a response using a fixed response phrase database, which requires fewer computational resources than the process of a large-scale language model, before starting to acquire a response from the large-scale language model. This prevents the user from having to wait from time t0 to time t4 without receiving a response from the AI ​​response output device 10010. From the user's perspective, whether the response output is using a fixed response phrase database or a response output from a large-scale language model, it is the same as receiving a response from the AI ​​response output device 10010.

[0273] 6, by providing step 601 before step 603, the response of the AI ​​response output device 10010 to the user can be artificially accelerated. This can further reduce the sense of discomfort felt by the user due to long waiting times. Furthermore, by outputting a response from a large-scale language model following a response using the template response database in step 604, the user can perceive these outputs as if they were a series of more natural outputs.

[0274] According to the AI ​​response output device and AI response output system of Example 6 described above, the waiting time for a response from the AI ​​response output device can be shortened, and the sense of discomfort felt by the user can be further reduced.

[0275] Example 7 Next, Example 7 of the present invention is an improvement of the AI ​​response output device 10010 or the AI ​​response output system described in the drawings of Examples 1 to 6. Specifically, in the response generation process of the AI ​​response output device 10010, a response output is generated by more suitably combining a response generation process using a large-scale language model on a network or a local large-scale language model (such as the local LLM processing unit 10028) provided inside the AI ​​response output device 10010, and a response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated explanations of configurations similar to those examples will be omitted.

[0276] As in the above-described embodiments, the AI ​​response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI ​​response output device 10010 and a large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system. The same applies to the AI ​​response output devices in the following embodiments.

[0277] First, an example of an AI response output device and an AI response output system according to a seventh embodiment of the present invention will be described with reference to Fig. 7. The configuration of the AI ​​response output system in Fig. 7 is an improved version of the character conversation system shown in Fig. 3C. Components with the same reference numerals as those in Fig. 3C are the same as those in Fig. 3C, and therefore repeated explanations will be omitted.

[0278] In the example of FIG. 7, as shown in explanation 7010, the AI ​​response output device 10010 includes a client application as software that is loaded into the memory 1109 shown in FIG. 1B and executed by the control unit 1110, for example.

[0279] The client application can control the input of user instructions described in each of the figures in the first to sixth embodiments. The client application also uses the user instructions to send and receive information to and from the large-scale language model. The client application acquires a response from the large-scale language model. Based on the user instructions, the client application can control the various components shown in FIG. 1B, including the display unit 10011 and the audio output unit 1140, to output a response to the user. The user instructions may be generated by the client application in response to an input from the user. As with the description in FIG. 3C, examples of the user input interface that accepts input from the user include a mouse, keyboard, and touch panel, which are examples of the operation input unit 1107 in FIG. 1B. The microphone 1139 that picks up the user's voice can also be considered a user input interface. The communication unit 1132 that communicates with the mobile information processing terminal 20010, such as a smartphone or tablet information processing terminal, used by the user can also be considered a user input interface. In addition, the output interface through which the client application outputs a response to the user includes a display unit 10011 that outputs the response from the large-scale language model in media such as text, images, or video, and an audio output unit 1140 that outputs the response from the large-scale language model in audio.

[0280] Here, the client application can control the insertion of fixed phrases into the response output by the AI ​​response output device 10010. This includes the output control of fixed phrases described in Figures 1C, 2L, 4B, 5A, and 6. Examples of the insertion of fixed phrases include the insertion of fixed phrases for greetings and fixed phrases for backchanneling. These fixed phrases can be stored in the storage unit 1170 of the AI ​​response output device 10010 in Figure 1B, for example, and then expanded into memory 1109 for use by the client application.

[0281] Also, in the example of FIG. 7 , as shown in explanation 7020, the large-scale language model has an LLM (large-scale language model) application that controls the input and output of information to the large-scale language model. As in the example of FIG. 7 , when the large-scale language model server 20001 has an LLM application (short for large-scale language model application; the same applies below), the LLM application is loaded into the memory of the large-scale language model server 20001 and executed by a control unit of the large-scale language model server 20001. The LLM application can accept preset instructions for adjusting the output of the large-scale language model to specific specifications. While user instructions are instructions whose content changes with each user input, preset instructions for adjusting the output of the large-scale language model to specific specifications are instructions whose content does not change with each user input and are given constantly. Therefore, these preset instructions may be referred to as constant instructions. They may also be referred to as instructions for customizing the output of the large-scale language model. For example, a preset instruction may be used to set the ending of a natural language response sentence generated by the large-scale language model to a specific catchphrase. Another preset instruction may be used to insert a specific backchannel into a natural language response sentence generated by the large-scale language model. The preset instructions for the LLM application may be set by sending instructions to the LLM application from a client application of the AI ​​response output device 10010. Alternatively, the preset instructions for the LLM application may be set by communicating with the LLM application via a communication unit of an information terminal such as a smartphone or personal computer used separately by the user.

[0282] 7 shows an example in which an LLM application that controls the input and output of information to the large-scale language model is present on the large-scale language model side. In contrast, as another variation, the AI ​​response output device 10010 may be configured to include an LLM application that controls the input and output of information to the local LLM processing unit 10028 of the AI ​​response output device 10010. In this case, the LLM application may be software that is deployed in memory 1109 and executed by the control unit 1110. In this case, the LLM application that controls the input and output of information to the local LLM processing unit 10028 and the above-mentioned client application may be different applications, or may be the same application.

[0283] In the example of Fig. 7, not only instruction statements but also control information can be transmitted from the AI ​​response output device 10010 to the large-scale language model server 20001. The control information may be transmitted together with the instruction statements. The control information may be transmitted prior to the instruction statements. The transmission of the control information and instruction statements may be controlled by the client application.

[0284] An example of the control information will be described in detail below. First, the control information may store authentication information for logging in to the LLM application. For example, a user may create an account in the LLM application in advance, and generate authentication information including identification information, a password, and the like. The account may be created by communicating with the LLM application via the communication unit 1132 of the AI ​​response output device 10010. Alternatively, the account may be created by communicating with the LLM application via a communication unit of an information terminal, such as a smartphone or personal computer, used separately by the user. The authentication information is sent from the communication unit 1132 of the AI ​​response output device 10010 to the LLM application, and login is performed through authentication processing. If authentication is successful, the AI ​​response output device 10010 and the LLM application establish communication as the user corresponding to the authentication information. This allows the user to use information available to the user's account from the information stored in the storage area of ​​the LLM application. In the case of an LLM application on the large-scale language model server 20001 side, the storage area of ​​the LLM application may be provided in the memory or storage unit of the large-scale language model server 20001. In addition, in the case of an LLM application that controls the input and output of information to the local LLM processing unit 10028 of the AI ​​response output device 10010, the storage area of ​​the LLM application may be provided in the storage unit 1170 or memory 1109 in Figure 1B.

[0285] Information available to the user's account includes the preset instructions described above. Furthermore, as shown in Figures 2H, 2I, 2L, 4A, and 4B, when the AI ​​response output device 10010 can switch between multiple characters, preset instruction information corresponding to each of the multiple characters may be stored in a storage area of ​​the LLM application. In this case, if the control information transmitted from the AI ​​response output device 10010 to the LLM application includes a character ID identifying the character, the LLM application can determine which character the user wants to apply the preset instruction to. In other words, this is nothing more than including preset instruction switching information for each character in the control information, and switching the preset instructions in the LLM application according to the switching information. Furthermore, when a user sets a preset instruction for each character in the LLM application, the AI ​​response output device 10010 may store and transmit setting information for the preset instruction in the control information transmitted to the LLM application. Furthermore, the preset instruction setting information may be transmitted to the LLM application via a communication unit of an information terminal, such as a smartphone or personal computer, separately used by the user. The LLM application that receives the setting information of the preset instruction stores the preset instruction information based on the setting information in a storage area. As described above, the preset instruction information may be stored for each character. Alternatively, it may be stored for each user account. Once the setting of the preset instruction is complete, a character ID identifying the character desired by the user is transmitted from the AI ​​response output device 10010 to the LLM application, and the LLM application can select the preset instruction to be applied to the character and apply it to the large-scale language model. In this case, there is no need to transmit information corresponding to the preset instruction to the LLM application each time an instruction sentence is sent, thereby reducing the amount of communication data.

[0286] As another example, each time an instruction is sent, information corresponding to a preset instruction may be stored and transmitted in the control information sent from the AI ​​response output device 10010 to the LLM application. In this case, the frequency of transmission of information corresponding to a preset instruction increases, but it is possible to omit preparations such as advance user account registration and advance preset instruction setting. Also, each time an instruction is sent, information corresponding to the preset instruction may be stored and transmitted in the setting instruction section of the instruction rather than the user instruction section. In this case, it is also possible to omit preparations such as advance user account registration and advance preset instruction setting.

[0287] Next, several examples of deployment and execution of client applications and LLM applications in an AI response output system will be described using Figures 8A, 8B, and 8C. The examples in Figures 8A, 8B, and 8C are schematic diagrams that focus on the deployment, execution, and communication of client applications and LLM applications in the configuration of an AI response output system that includes an AI response output device 10010 and a large-scale language model server 20001. For simplicity, the notation and description of configurations other than the client application and LLM application will be omitted.

[0288] First, in the example of FIG. 8A, a client application 8010 is executed in the AI ​​response output device 10010. Specifically, the client application 8010 is loaded into the memory 1109 shown in FIG. 1B and executed by the control unit 1110. Also, a server LLM application 8020 is executed in the large-scale language model server 20001. Specifically, the server LLM application 8020 is loaded into a memory provided in the large-scale language model server 20001 and executed by the control unit provided in the large-scale language model server 20001. The server LLM application 8020 may be an application that controls the entire large-scale language model provided in the large-scale language model server 20001. The server LLM application 8020 may be an application that controls the input and output of information to and from the large-scale language model provided in the large-scale language model server 20001. The client application 8010 and the server LLM application 8020 communicate with each other via, for example, the communication unit 1132 shown in FIG. 1B, the communication device 19011 shown in FIG. 1A, a network such as the Internet 19000, and a communication unit included in the server. In the example AI response output system of FIG. 8A, the client application 8010 and the server LLM application 8020 cooperate to output a more suitable AI response to the user. Here, the large-scale language model included in the large-scale language model server 20001 may be referred to as a server large-scale language model. The server LLM application 8020 may be referred to as a server large-scale language model application.

[0289] Next, in another example shown in FIG. 8B, a client application 8010 is executed in the AI ​​response output device 10010. As in FIG. 8A, the client application 8010 is loaded into the memory 1109 shown in FIG. 1B and executed by the control unit 1110. Also, a local LLM application 8015 is executed in the AI ​​response output device 10010. Specifically, the local LLM application 8015 is loaded into the memory 1109 shown in FIG. 1B in the AI ​​response output device 10010 and executed by the control unit 1110. The local LLM application 8015 may be an application that controls the entire large-scale language model possessed by the local LLM processing unit 10028 shown in FIG. 1B. The local LLM application 8015 may be an application that controls the input and output of information to and from the large-scale language model possessed by the local LLM processing unit 10028. The client application 8010 and the local LLM application 8015 communicate with each other via a communication path, such as a bus, within the AI ​​response output device 10010. In the AI ​​response output system shown in FIG. 8B, the client application 8010 and the local LLM application 8015 cooperate to output a more suitable AI response to the user. Here, the large-scale language model possessed by the local LLM processing unit 10028 may be referred to as a local large-scale language model. The local LLM application 8015 may be referred to as a local large-scale language model application.

[0290] Next, in another example shown in FIG. 8C , a client application 8010 and a local LLM application 8015 are executed in an AI response output device 10010. Details of the client application 8010 and the local LLM application 8015 are as described in FIGS. 8A and 8B , and therefore repeated description will be omitted. A server LLM application 8020 is executed in a large-scale language model server 20001. Details of the server LLM application 8020 are as described in FIG. 8A , and therefore repeated description will be omitted. The client application 8010 and the server LLM application 8020 communicate with each other via, for example, the communication unit 1132 shown in FIG. 1B , the communication device 19011 shown in FIG. 1A , a network such as the Internet 19000, and a communication unit provided in the server. The client application 8010 and the local LLM application 8015 communicate with each other via, for example, a communication path such as a bus within the AI ​​response output device 10010. In the artificial intelligence response output system of the example of Figure 8C, the client application 8010, the local LLM application 8015, and the server LLM application 8020 work together to output a more suitable artificial intelligence response to the user.

[0291] Next, an example of an AI response output process in the AI ​​response output system according to the seventh embodiment of the present invention will be described with reference to FIGS. 9A to 9G.

[0292] First, in Fig. 9A, an example of a user instruction sentence input by a user is shown on the right side. Details of the method for inputting a user instruction sentence are as explained in Examples 1 to 6, so repeated explanation will be omitted. On the left side of Fig. 9A, an example of output from an AI response output device 10010 of the AI ​​response output system is shown. Both the input of the user instruction sentence on the right side and the output from the AI ​​response output device 10010 on the left side are shown in chronological order from top to bottom.

[0293] Here, the example of output from the AI ​​response output device 10010 on the left side first shows fixed phrase response 1 (greeting). This is an example in which a fixed phrase greeting is output from the AI ​​response output device 10010 using the fixed phrase output control described in Figures 1C, 2L, 4B, 5A, or 6. The output control of this fixed phrase can be performed by the client application 8010. Next, when user instruction 1 is input to the AI ​​response output device 10010, the client application 8010 acquires it and transmits the user instruction 1 and control information to the large-scale language model. The large-scale language model that receives user instruction 1 performs inference at the timing indicated by the star mark 9001 and generates LLM response 1. The large-scale language model may be a large-scale language model of the large-scale language model server 20001. In this case, which corresponds to the state of Figure 8A, the server LLM application 8020 controls the input and output of information to the large-scale language model. The large-scale language model may also be a large-scale language model possessed by the local LLM processing unit 10028. In this case, which corresponds to the state of FIG. 8B, the local LLM application 8015 controls the input and output of information to the large-scale language model. The LLM response 1 generated by the large-scale language model is sent to the client application 8010 by these LLM applications and output as an AI response from the AI ​​response output device 10010. If the AI ​​response output device 10010 is a character conversation device, the AI ​​response output is recognized by the user as the response of the displayed character.

[0294] Next, an example is shown in which user instruction 2 is input to the AI ​​response output device 10010 in response to LLM response 1. Here, an example is shown in which the client application 8010, having acquired user instruction 2, outputs fixed-phrase response 2 (a backchannel response). This is an example in which a fixed-phrase backchannel response is output from the AI ​​response output device 10010 using the fixed-phrase output control described in Figures 1C, 2L, 4B, 5A, or 6. Multiple types of fixed-phrase backchannel responses can be prepared in advance, and controlled to be output randomly at intervals of at least a predetermined period to avoid unnatural high frequency. Meanwhile, as described in Figure 6, while outputting the fixed phrase, the client application 8010 transmits the user instruction 2 and control information to the above-mentioned LLM application (server LLM application 8020 or local LLM application 8015). The LLM application that acquires the user instruction sentence 2 controls the large-scale language model (the large-scale language model of the large-scale language model server 20001 or the large-scale language model possessed by the local LLM processing unit 10028) to perform inference at the timing of the star mark 9002. An LLM response 2 is generated by the inference of the large-scale language model. The LLM response 2 generated by the large-scale language model is sent to the client application 8010 by these LLM applications and output as an AI response from the AI ​​response output device 10010. If the AI ​​response output device 10010 is a character conversation device, the AI ​​response output is recognized by the user as the response of the displayed character.

[0295] A conversation between the AI ​​response output system and the user is carried out through the above series of processes. FIG. 9A shows an example of a series of conversations between a user and an AI regarding Haneda Airport in Japan, with greetings and interjections interspersed. The series of AI responses output from the AI ​​response output device 10010 appears to the user as a series of responses with a certain degree of consistency. However, these series of responses are composed of a combination of response sentences generated by a large-scale language model in which the output of standard phrases is controlled by a client application and the input and output of information is controlled by an LLM application. In other words, collaboration between the client application and the LLM application produces an AI response output that is more suitable for the user.

[0296] Next, using Figures 9B to 9G, we will explain an example of control of the character's conversation characteristics when a client application and an LLM application cooperate to output a series of AI responses in an AI response output system. Specifically, we will explain an example of control of fixed phrases for greetings and backchannels, and catchphrases and backchannels in responses output from a large-scale language model.

[0297] The example of FIG. 9B is an example in which the AI ​​response output system is a character conversation system or an AI assistant system. The example of FIG. 9B is also an example in which the AI ​​response output system can switch between and display multiple characters or AI assistants, as in FIG. 2H, FIG. 2I, FIG. 2L, FIG. 4A, or FIG. 4B. Here, the example of FIG. 9B will be described using an example of two characters. The explanation is given using only two characters for the sake of simplicity; three or more characters, including the characters described in the previous embodiments, may be switchably displayed. The two characters include a character with a character ID of 3 and a name of Necco. This character is also described in the previous embodiments. The two characters also include a new character with a character ID of 4 and a name of Airia.

[0298] The table in Figure 9B shows the character IDs, names, and character display examples for these two characters. Furthermore, the table in Figure 9B shows examples of standard greetings and standard backchannel responses that the client application inserts for each character. The process by which the client application inserts standard greetings and standard backchannel responses was explained in Figure 7, so a repeated explanation will be omitted.

[0299] In the table of FIG. 9B, distinctive expressions (keywords) are used in the standard greetings and responses to give these characters individuality. For example, Necco is a cat-like character, so her standard phrases end with "meow." Necco's first-person pronoun is "boku." For example, Airia's standard phrases end with "desuwa" or "masuwa." That is, her final words are set to end with "wa." Airia's first-person pronoun is "watakushi." These keywords may be referred to as keywords that indicate the character's individuality. These keywords can be used in common in each character's conversations to express the character's individuality. When the character switching process described in FIG. 2H is performed, the client application simply switches the keywords according to the setting information in the table of FIG. 9B.

[0300] Furthermore, the table in Fig. 9B has columns for setting preset instructions for the LLM application regarding catchphrases, interjections, etc. for these characters, but the example in Fig. 9B shows an example in which these settings are not made. The processing of preset instructions for the LLM application has been explained in Fig. 7, so a repeated explanation will be omitted.

[0301] The information shown in the table of Fig. 9B may be stored as table information in the storage unit 1170 or memory 1109 of the AI ​​response output device 10010 shown in Fig. 1B. This information may be used by a client application. The information shown in the table of Fig. 9B may also be referred to as setting information regarding the characteristics of the character's conversation.

[0302] Next, using Figures 9C and 9D, we will explain an example of a conversation when a client application performs the process of inserting the standard greetings and standard backchannel responses shown in the table of Figure 9B. To make the explanation easier to understand, the conversation examples in Figures 9C and 9D use the conversation example in Figure 9A as the original text, with only the parts that change due to the process of inserting the standard greetings and standard backchannel responses shown in the table of Figure 9B replaced.

[0303] First, FIG. 9C shows an example of a conversation in which the client application controls the AI ​​response output device 10010 to output an AI response as a response from the character Necco. The client application reads information about the character Necco's standard greetings and standard backchannel responses from the table information shown in FIG. 9B and uses the information to output standard AI responses. When the client application reads the setting information for each character from the table information, it can identify the information using a character ID or the like. As shown in the conversation example in FIG. 9C, standard greeting response 1 uses the character Necco's standard greetings, uses the first-person pronoun "boku" (I) and ends with "meow." Similarly, standard response 2 uses the character Necco's standard backchannel responses and ends with "meow." However, in the table of FIG. 9B, no catchphrases or backchannel responses have been set by the LLM application preset instructions. Therefore, LLM Response 1 and LLM Response 2 are the original text of Figure 9A, and do not end with "meow" or use the first person pronoun "boku." As a result, a user who is having a series of conversations with Necco will get the impression that the responses output by AI response output device 10010 from character Necco are inconsistent with the character's personality.

[0304] Similarly, FIG. 9D shows an example of a conversation in which the client application controls the AI ​​response output device 10010 to output an AI response as a response from the character Airia. In this case, the client application reads information about the character Airia's standard greeting and standard backchannel responses from the table information shown in FIG. 9B and uses these information to output the standard AI response responses. In the example of FIG. 9D, standard response 1 uses the character Airia's standard greeting, the first-person pronoun is "watashi" ("I am"), and the ending is "wa." Furthermore, standard response 2 uses the character Airia's standard backchannel responses and the ending is "wa." However, in the table of FIG. 9B, no catchphrases or backchannel responses have been set by the LLM application preset instructions. Therefore, LLM response 1 and LLM response 2 remain the original text of FIG. 9A, and the ending is not "watashi" and the first-person pronoun is not "watashi." As a result, even in the conversation example of FIG. 9D, a user who is engaged in a series of conversations will get the impression that the content output as the response of the character Airia from the AI ​​response output device 10010 is inconsistent with the character's personality.

[0305] In this way, when a client application and an LLM application work together to output a series of artificial intelligence responses, it is not easy to give the user a consistent impression of the character's personality even if the character's personality is set solely through the client application's template output control.

[0306] An example of improved control to solve this problem will be described with reference to FIGS. 9E to 9G.

[0307] FIG. 9E shows an example of an improved version of the table information in FIG. 9B. In the example of FIG. 9E, the character ID, character name, character image, and insertion of standard greetings and standard backchannel responses from the client application are the same as those in the table information in FIG. 9B, and therefore a repeated description will be omitted. In the example of FIG. 9E, settings for catchphrases and backchannel responses have been added to the preset instructions for the LLM application, which were not set in the table information in FIG. 9B. The preset instructions for the LLM application can be set by sending control information from the client application to the LLM application. This process is the same as that described in FIG. 7, and therefore a repeated description will be omitted. The information shown in FIG. 9E also relates to the settings related to the character's conversational characteristics. In the example of FIG. 9E, the settings for catchphrases and backchannel responses in the preset instructions for the LLM application include instructions for outputting characteristic expressions (keywords) common to the characteristic expressions (keywords) included in the standard greetings and standard backchannel responses inserted by the client application in the response of the large-scale language model. This allows the character's individuality to be reflected in the response output of the large-scale language model. Specifically, because Necco is a cat-like character, the settings for the preset greetings and responses inserted in the client application, as well as the settings for the catchphrases and responses in the preset instructions for the LLM application, are set to end with "meow." Also, the settings for the preset greetings and responses inserted in the client application, as well as the settings for the catchphrases and responses in the preset instructions for the LLM application, are set to end with "boku" (I). Similarly, for the character Airia, the settings for the preset greetings and responses inserted in the client application, as well as the settings for the catchphrases and responses in the preset instructions for the LLM application, are set to end with "wa" (Wa). Also, the settings for the preset greetings and responses inserted in the client application, as well as the settings for the catchphrases and responses in the preset instructions for the LLM application, are set to end with "wa." Also, the settings for the preset greetings and responses inserted in the client application, as well as the settings for the preset greetings and responses in the LLM application, are set to end with "watashi" (I).That is, in the example of table information shown in Fig. 9E, characteristic expressions (keywords) that indicate the individuality of each character and that are included in the settings of standard greetings and standard backchannels inserted by the client application and that overlap with the characteristic expressions (keywords) are also included in the settings of catchphrases and backchannels in the preset instructions of the LLM application. Note that when the character switching process described in Fig. 2H is performed, the client application simply switches the settings related to the conversation characteristics of the character shown in Fig. 9E, such as the settings of standard greetings, the settings of standard backchannels, and the settings of catchphrases and backchannels in the preset instructions of the LLM application, in accordance with the character switching process.

[0308] An example of a conversation to which the control of the standard phrases and preset instructions of the table information shown as a table in Fig. 9E, which has been improved in this way, is applied, will be described using Fig. 9F and Fig. 9G. To make the explanation easier to understand, the examples of conversation in Fig. 9F and Fig. 9G are shown as examples of a conversation to which the control of the standard phrases and preset instructions of the table information shown as a table in Fig. 9E is applied, using the example of conversation in Fig. 9A as the original text.

[0309] FIG. 9F is an example of a conversation in which the control of the fixed phrases and preset instructions in the table information shown in FIG. 9E is applied. FIG. 9F shows an example of a conversation in which a client application controls the output of an AI response from the AI ​​response output device 10010 as a response from the character Necco. The client application reads information on the character Necco's fixed greeting phrases and fixed backchannel responses from the table information shown in FIG. 9E and uses these to output fixed phrases for the AI ​​response. As shown in the conversation example in FIG. 9F, fixed phrase response 1 uses the character Necco's fixed greeting phrase, uses the first-person pronoun "boku" (I) and ends with "meow." Furthermore, fixed phrase response 2 uses the character Necco's fixed backchannel responses and ends with "meow." In the table of FIG. 9E, the catchphrases and backchannel settings based on the LLM application preset instructions include characteristic expressions (keywords) that overlap with the characteristic expressions (keywords) that express the personality of the character Necco, such as "meow" and the first-person pronoun "boku" (I), which are included in the endings and backchannel settings. As a result, in the conversation example of FIG. 9F, the ending in LLM Response 1 is "meow." Similarly, in LLM Response 2, the ending is "meow" and the first-person pronoun is "boku." Furthermore, LLM Response 2 includes the backchannel "hmm meow," which was generated by the large-scale language model in accordance with the LLM application preset instructions. This backchannel also ends with "meow." Thus, in the content output as responses by the character Necco from the AI ​​response output device 10010 in the series of conversations, the same characteristic expressions (keywords) are consistently used, whether the responses are standard phrases output by the client application or responses output by the large-scale language model controlled by the LLM application. This allows the user to get a more consistent impression of the character's personality from the content output as the character Necco's response from the AI ​​response output device 10010.

[0310] Similarly, FIG. 9G is another example of a conversation to which the control of the fixed phrases and preset instructions of the table information shown in FIG. 9E is applied. FIG. 9G shows an example of a conversation in which a client application controls the output of an AI response from the AI ​​response output device 10010 as a response from the character Airia. The client application reads information on the fixed greeting phrases and fixed backchannel responses of the character Airia from the table information shown in FIG. 9E and uses these to output the fixed phrases of the AI ​​response. As a result, as shown in the conversation example in FIG. 9G, fixed phrase response 1 uses the fixed greeting phrase of the character Airia, the first person pronoun is "watashi" (I) and the ending is "wa." Furthermore, fixed phrase response 2 uses the fixed backchannel responses of the character Airia and the ending is "wa." In the table of FIG. 9E, the catchphrases and backchannel settings based on the LLM application preset instructions include characteristic expressions (keywords) that overlap with the characteristic expressions (keywords) that characterize the character Airia's personality, such as "wa" and the first-person pronoun "watakushi" ("watashi") included in the ending and backchannel settings. As a result, in the example conversation of FIG. 9G, LLM Response 1 ends with "wa." Similarly, LLM Response 2 ends with "wa" and the first-person pronoun "watakushi" ("watashi"). Furthermore, LLM Response 2 includes the backchannel "uuuuu," generated by the large-scale language model in accordance with the LLM application preset instructions. This backchannel also ends with "wa." Thus, in the content output as the character Airia's responses from the AI ​​response output device 10010 in the series of conversations, the same characteristic expressions (keywords) are consistently used, whether the responses are standard phrases output by the client application or responses output by the large-scale language model controlled by the LLM application. This allows the user to get a more consistent impression of the character's personality from the content output as the character Airia's response from the AI ​​response output device 10010.

[0311] As described above, in the AI ​​response output system of this embodiment described with reference to FIGS. 9A to 9G, the client application and the LLM application cooperate to output a series of AI responses. Furthermore, in the AI ​​response output system of this embodiment described with reference to FIGS. 9A to 9G, control is performed so that common characteristic expressions (keywords) are consistently used for each character in the fixed phrase output by the client application and the response output of the large-scale language model controlled by the LLM application. This provides the effect of giving the user a more consistent impression of the character's personality in the content output from the AI ​​response output system. Furthermore, in this system, the AI ​​response output device 10010 controls the client application it executes to send and receive control information to the LLM application, thereby controlling so that common characteristic expressions (keywords) are consistently used for each character in the fixed phrase output by the client application and the response output of the large-scale language model controlled by the LLM application. This control of the client application executed by the AI ​​response output device 10010 provides the effect of giving the user a more consistent impression of the character's personality in the content output from the AI ​​response output system.

[0312] In the conversation examples shown in Figures 9F and 9G, output of fixed phrase response 2, which is a fixed phrase for a backchannel response, is initiated after receiving user instruction sentence 2 and before LLM response output 2, which is the response output of the large-scale language model controlled by the LLM application. Because inference of a large-scale language model takes some time, the client application controls to insert a fixed phrase for a backchannel response at this timing to prevent an unnatural gap in the response output from the AI ​​response output device 10010 to the user. This is the same method as in the control of Figure 6, where output of a fixed phrase response is initiated first and then the large-scale language model is caused to generate a response. In other words, one suitable timing for outputting a fixed phrase response is immediately before the start of inference of the large-scale language model.

[0313] 10A to 10D, an example will be described in which a client application, a local LLM application, and a server LLM application work together to output an artificial intelligence response to a user. This state corresponds to FIG. 8C.

[0314] The conversation example in Figure 10A illustrates an example in which a client application, a local LLM application, and a server LLM application work together to output an artificial intelligence response to a user. For simplicity, much of the content of the conversation example in Figure 10A is the same as that of the conversation example in Figure 9A. Therefore, the explanation of Figure 10A focuses on the differences from the conversation example in Figure 9A. The conversation example in Figure 10A does not include the various settings for expressing the character's personality using the table information described in Figure 9E. In the conversation example in Figure 10A, the standard response 1 and user instruction sentence 1 are identical to those in the conversation example in Figure 9A. Next, the server LLM response in Figure 10A is identical in content to LLM response 1 in Figure 9A, but this server LLM response is a response output based on the response of the server-side large-scale language model controlled by the server LLM application 8020. Therefore, the inference performed at the timing of the star mark 9003 is an inference based on the large-scale language model of the large-scale language model server 20001.

[0315] Here, the conversation example in FIG. 10A shows an example in which the network between the AI ​​response output device 10010 and the large-scale language model server 20001 becomes unavailable after the server LLM response. In the AI ​​response output system, the control for switching between the response of the server-side large-scale language model and the response of the local large-scale language model of the AI ​​response output device 10010 is as described in, for example, Example 5 ( FIG. 5A , etc.), and therefore a repeated description will be omitted. User instruction 2 in FIG. 10A is the same as user instruction 2 in FIG. 9A. However, because the network is unavailable, the client application 8010 sends user instruction 2 to the local LLM application 8015 rather than the server LLM application 8020. The client application 8010 also outputs a standardized response for network outages, i.e., a standardized response for network outages. Therefore, at the timing of the star mark 9004, inference is performed using the large-scale language model possessed by the local LLM processing unit 10028 controlled by the local LLM application 8015. 10A, a local LLM response is output based on inference using a large-scale language model possessed by the local LLM processing unit 10028. That is, in the conversation example of Fig. 10A, the output to the user from the AI ​​response output device 10010 of the AI ​​response output system will include a mixture of fixed phrase output by the client application 8010, output based on the response of a large-scale language model on the server side controlled by the server LLM application 8020, and output based on the response of a large-scale language model possessed by the local LLM processing unit 10028 controlled by the local LLM application 8015.

[0316] Next, an example of controlling various settings to express the individuality of the character according to this embodiment will be described with reference to the conversation series in FIG. 10A, using FIGS. 10B to 10D.

[0317] First, Fig. 10B shows table information obtained by improving the table information for various settings in Fig. 9E. Fig. 10B is an example of an improvement to the table information in Fig. 9E. In the example of Fig. 10B, the character ID, character name, and character display image are the same as the table information in Fig. 9E, so repeated explanations will be omitted.

[0318] In the example of FIG. 10B, the settings for the standard greetings and standard backchannel responses inserted by the client application are omitted in FIG. 10B for simplicity's sake, but are assumed to be the same as those in FIG. 9E. FIG. 10B also shows the settings for standard responses inserted by the client application when the network is down. The process by which the AI ​​response output device 10010 locally inserts standard responses when the network is down has been described in Example 2 of FIG. 5A, and therefore a repeated description will be omitted. Also, in the example of FIG. 10B, the preset instructions for the LLM application set in the table of FIG. 9E are stored as preset instructions in two locations: a server LLM application preset instruction and a local LLM application preset instruction. In the example of FIG. 10B, catchphrases and backchannel responses can be set for each of the two preset instructions, and are set as shown in FIG. 10B. The preset instructions for the server LLM application can be set by sending control information from the client application to the server LLM application. The preset instructions for the local LLM application can be set by sending control information from the client application to the local LLM application. These processes are the same as those explained in Fig. 7, so a repeated explanation will be omitted. Note that the information shown in the table of Fig. 10B is also setting information relating to the characteristics of the character's conversation.

[0319] In the table information of various settings in Figure 10B, the preset catchphrases and backchannels in the server LLM application and the local LLM application are set so that the individuality of the character indicated by the characteristic expressions (keywords) included in the standard phrases inserted by the client application when the network is down is reflected in the response output of the large-scale language model controlled by the server LLM application and the response output of the large-scale language model controlled by the local LLM application. In the example of Figure 10B, the settings of the preset catchphrases and backchannels in the server LLM application include instructions for outputting characteristic expressions (keywords) common to the characteristic expressions (keywords) included in the standard phrases inserted by the client application in the response of the large-scale language model controlled by the server LLM application. Similarly, the settings of the preset catchphrases and backchannels in the local LLM application include instructions for outputting characteristic expressions (keywords) common to the characteristic expressions (keywords) included in the standard phrases inserted by the client application in the response of the large-scale language model controlled by the local LLM application. In other words, characteristic expressions (keywords) common to those included in the template phrases inserted by the client application are also included in the catchphrases and backchannel settings of the server LLM application and in the catchphrases and backchannel settings of the preset instructions in the local LLM application. Specifically, because Necco is a cat-like character, the catchphrases and backchannel settings of the preset instructions in the client application, such as those for when the network is down, end with "meow." Furthermore, in these template phrases, Necco's first-person pronoun is set to "boku." Here, in the example of FIG. 10B, the catchphrases and backchannel settings of the preset instructions in the server LLM application are also set to end with "meow," and Necco's first-person pronoun is set to "boku." Similarly, the catchphrases and backchannel settings of the preset instructions in the local LLM application are also set to end with "meow," and Necco's first-person pronoun is set to "boku."The settings for catchphrases and backchannels for preset instructions in the server LLM application and the settings for catchphrases and backchannels for preset instructions in the local LLM application only need to share characteristic expressions (keywords), and the specific settings for catchphrases and backchannels do not need to be exactly the same. Similarly, in the example of FIG. 10B, the settings for the character Airia, such as for when the network is down, are configured to end with "wa." Furthermore, in these template messages, Airia's first-person pronoun is set to "watashi." Here, in the example of FIG. 10B, the settings for catchphrases and backchannels for preset instructions in the server LLM application are also configured to end with "wa" and to use Airia's first-person pronoun. Similarly, the settings for catchphrases and backchannels for preset instructions in the local LLM application are also configured to end with "wa" and to use Airia's first-person pronoun. In addition, when the character switching process described in Figure 2H is performed, the client application simply switches the settings related to the character's conversation characteristics shown in Figure 10B, such as the settings of various standard phrases, the settings of catchphrases and backchannels for preset instructions of the server LLM application, and the settings of catchphrases and backchannels for preset instructions of the local LLM application, in accordance with the character switching process.

[0320] An example of a conversation to which the control of the standard phrases and preset instructions of the table information shown in Fig. 10B is applied will be described using Fig. 10C and Fig. 10D. To make the explanation easier to understand, the conversation examples in Fig. 10C and Fig. 10D are shown using the conversation example in Fig. 10A as the original text, and Fig. 10B as an example of a conversation to which the control of the standard phrases and preset instructions of the table information is applied.

[0321] FIG. 10C is an example of a conversation to which the control of the fixed phrases and preset instructions of the table information shown in FIG. 10B is applied. FIG. 10C shows an example of a conversation in which a client application controls the output of an AI response from the AI ​​response output device 10010 as a response from the character Necco. In the conversation example of FIG. 10C, the control of the fixed greeting phrase by the client application shown in FIG. 9E is applied to fixed phrase response 1. The server LLM application preset instruction shown in FIG. 10B is applied to the server LLM response. The control of the fixed phrase for network outages shown in FIG. 10B is applied to the fixed phrase for network outages. Furthermore, the local LLM application preset instruction shown in FIG. 10B is applied to the local LLM response. In the example of FIG. 10C, the local LLM response includes the backchannel response "That's right, meow." generated by the large-scale language model of the local LLM processing unit 10028 controlled by the local LLM application. As a result, in a series of conversations, whether the content output as the character Necco's response from the AI ​​response output device 10010 is a standard phrase output by a client application, a response output from a large-scale language model controlled by a server LLM application, or a response output from a large-scale language model possessed by the local LLM processing unit 10028 controlled by a local LLM application, the content ends with "meow" and Necco uses the first-person pronoun "boku." In other words, a common characteristic expression (keyword) is consistently used. This allows the user to get a more consistent impression of the character's personality from the content output as the character Necco's response from the AI ​​response output device 10010.

[0322] Similarly, FIG. 10D is another example of a conversation to which the control of the standard phrases and preset instructions of the table information shown in FIG. 10B is applied. FIG. 10D shows an example of a conversation in which a client application controls the output of an AI response from the AI ​​response output device 10010 as a response from the character Airia. In the conversation example of FIG. 10D, the control of the standard greeting phrase by the client application shown in FIG. 9E is applied to standard phrase response 1. The server LLM application preset instruction shown in FIG. 10B is applied to the server LLM response. The control of the standard phrase for network outages shown in FIG. 10B is applied to the standard phrase for network outages. Furthermore, the local LLM application preset instruction shown in FIG. 10B is applied to the local LLM response. In the example of FIG. 10D, the local LLM response includes the backchannel response "Umm, that's it." generated by the large-scale language model of the local LLM processing unit 10028 controlled by the local LLM application. As a result, in a series of conversations, whether the content output as a response from the character Airia from the AI ​​response output device 10010 is a standard phrase output by a client application, a response output from a large-scale language model controlled by a server LLM application, or a response output from a large-scale language model possessed by the local LLM processing unit 10028 controlled by a local LLM application, the content ends with "wa" and Airia uses the first-person pronoun "watakushi." In other words, a common characteristic expression (keyword) is consistently used. This allows the user to get a more consistent impression of the character's personality from the content output as a response from the character Airia from the AI ​​response output device 10010.

[0323] As described above, in the AI ​​response output system of this embodiment described with reference to FIGS. 10A to 10D, the client application, the server LLM application, and the local LLM application cooperate to output a series of AI responses. Furthermore, in the AI ​​response output system of this embodiment described with reference to FIGS. 10A to 10D, control is performed so that common characteristic expressions (keywords) are used consistently for each character in the fixed phrase output by the client application, the response output of the large-scale language model controlled by the server LLM application, and the response output of the large-scale language model controlled by the local LLM application. This provides the effect of giving the user a more consistent impression of the character's personality in the content output from the AI ​​response output system. Furthermore, in this system, the AI ​​response output device 10010 controls the client application it executes to send and receive control information with the server LLM application and to send and receive control information with the local LLM application, thereby controlling so that common characteristic expressions (keywords) are used consistently for each character in the fixed phrase output by the client application, the response output of the large-scale language model controlled by the server LLM application, and the response output of the large-scale language model controlled by the local LLM application. This control of the client application executed by the AI ​​response output device 10010 has the effect of giving the user a more consistent impression of the character's personality in relation to the content output from the AI ​​response output system.

[0324] Example 8 Next, an eighth embodiment of the present invention is an improvement of the AI ​​response output device 10010 or AI response output system described in the drawings of the first to seventh embodiments. Specifically, this is an example in which output control of fixed phrases such as backchannels is more suitably performed in the response generation process of the AI ​​response output device 10010. In this embodiment, differences from these embodiments will be described, and repeated description of configurations similar to these embodiments will be omitted. Note that the LLM application in the description of this embodiment may be a server LLM application or a local LLM application.

[0325] In the example of FIG. 9E of Example 7, the table information of FIG. 9E is used to control the inclusion of characteristic expressions (keywords) for each character contained in template text inserted by the client application in preset instructions for the LLM application regarding the character's responses, such as backchannels. However, because the response output of the large-scale language model controlled by the LLM application is generated by inference, there may be cases where the desired response output for responses, such as backchannels, cannot be obtained probabilistically depending on the random numbers used in the inference and the performance of the large-scale language model. Therefore, in the AI ​​response output system according to Example 8 of the present invention, this problem is solved by the client application of the AI ​​response output device 10010 performing control using the table information shown in FIG. 11A.

[0326] The table information shown in FIG. 11A is an example in which a portion of the table information in FIG. 9E has been modified. The information stored in the table information in FIG. 11A also contains information about settings related to the character's conversational characteristics. When the character switching process described in FIG. 2H is performed, the client application simply switches the settings related to the character's conversational characteristics shown in FIG. 11A, such as various standard phrase settings, catchphrases, and backchannel settings in the preset instructions of the LLM application, in accordance with the character switching process. For simplicity's sake, only the differences between the table information shown in FIG. 11A and the table information in FIG. 9E will be described, and the same portions as those in FIG. 9E will not be described. Specifically, the example of the table information shown in FIG. 11A differs from the table information in FIG. 9E in that the preset instructions of the LLM application instruct the backchannel to "output backchannels using the symbols $%&." That is, in this embodiment, the client application sends a preset instruction to the LLM application to output backchannels using predetermined symbols rather than natural language. When the table information shown in FIG. 11A is used, the large-scale language model controlled by the LLM application that received the preset instruction outputs "$%&" as a backchannel response. The LLM application then transmits the backchannel response symbol "$%&" to the client application. Upon receiving the backchannel response symbol "$%&," the client application randomly replaces the backchannel response symbol "$%&" with one of the multiple backchannel response symbols for each character stored in the table information shown in FIG. 11A, and outputs the resulting response from the AI ​​response output device 10010 to the user. In this way, regardless of the large-scale language model controlled by the LLM application, both the backchannel response symbol inserted by the client application and the backchannel response symbol "$%&" output by the large-scale language model controlled by the LLM application that the client application replaces contain common expressions that correspond to the individuality of the characters.This allows the user to get a more consistent impression of the character's personality in response to the backchannels output by the AI ​​response output system. Note that the symbols "$%&" shown in FIG. 11A are just an example, and other symbols can be used. However, it is desirable to set the symbols using a combination of special symbols or characters so that they do not match words used in normal conversation. It is also desirable that the symbols do not include natural language characters that indicate vowels. Considering the possibility that the client application may fail to perform the above-mentioned replacement, if the symbols include natural language characters that indicate vowels, there is a possibility that the characters included in the symbols will be spoken when output as the character's voice. If the symbols do not include natural language characters that indicate vowels, the symbols cannot be spoken even in the case of a failed replacement. Therefore, the AI ​​response output device 10010 preferably ignores the symbols in the voice output.

[0327] Next, a conversation example using the table information in FIG. 11A will be described using FIGS. 11B and 11C. The conversation examples in FIGS. 11B and 11C are based on the conversation examples in FIGS. 9F and 9G, respectively. For simplicity, only differences from the conversation examples in FIGS. 9F and 9G will be described, and explanations of similarities to the conversation examples in FIGS. 9F and 9G will be omitted. Note that the LLM application in the conversation examples in FIGS. 11B and 11C may be a server LLM application or a local LLM application.

[0328] FIG. 11B shows an example of a conversation in which the client application controls the AI ​​response output device 10010 to output an AI response as a response from the character Necco. The conversation example in FIG. 11B differs from the conversation example in FIG. 9F in one respect: In the conversation in FIG. 9F, the backchannel response "hmm, meow" generated by the large-scale language model controlled by the LLM application has been replaced with the LLM backchannel response symbol "$%&." In the example in FIG. 11B, the client application replaces the LLM backchannel response symbol "$%&" output from the LLM application with "hmm, meow," selected by random numbers from the fixed backchannel response phrases of the character Necco in FIG. 11A, and outputs the resulting response to the user. As a result, in the content output as the character Necco's response from the AI ​​response output device 10010 in the series of conversations, the word "meow" is added to the end, and Necco uses "boku" (I) as her first-person pronoun. In other words, the same characteristic expressions (keywords) are used consistently throughout the conversation. As a result, even when the LLM interjection response symbol "$%&" in Figure 11B is used, the user can get a more consistent impression of the character's personality from the content output as the character Necco's response from the artificial intelligence response output device 10010.

[0329] Next, FIG. 11C illustrates an example of a conversation in which the client application controls the AI ​​response output device 10010 to output an AI response as a response from the character Airia. The conversation example in FIG. 11C differs from the conversation example in FIG. 9G in one respect: In the conversation in FIG. 9G, the backchannel response "Uh-huh" generated by the large-scale language model controlled by the LLM application is replaced with the LLM backchannel response symbol "$%&." In the example in FIG. 11C, the LLM backchannel response symbol "$%&" output from the LLM application is replaced with "Uh-huh," selected by the client application using a random number from the fixed backchannel response phrases of the character Airia in FIG. 11A, and output to the user. As a result, in the content output as the character Airia's response from the AI ​​response output device 10010 in the series of conversations, the ending "wa" is added, and Airia's first-person pronoun is "watashi." In other words, a common characteristic expression (keyword) is consistently used. As a result, even when the LLM interjection response symbol "$%&" in Figure 11B is used, the user can get a more consistent impression of the character's personality from the content output as the character Airia's response from the artificial intelligence response output device 10010.

[0330] 11A to 11C, in the AI ​​response output system of this embodiment, a preset instruction of the LLM application instructs that backchannels be output using predetermined symbols, and the client application replaces the predetermined symbols, i.e., the LLM backchannel response symbols generated by the large-scale language model, with fixed phrases for backchannels. This more reliably achieves the effect of giving the user a more consistent impression of the character's personality in relation to the content output from the AI ​​response output system.

[0331] In this system, the AI ​​response output device 10010 executes a client application that instructs the LLM application to output backchannels using predetermined symbols in a preset instruction sent to the LLM application, and receives the predetermined symbols, which are LLM backchannel response symbols generated by the large-scale language model, and replaces them with standard backchannel phrases. This more reliably achieves the effect of giving the user a more consistent impression of the character's personality in the content output from the AI ​​response output system.

[0332] Example 9 Next, a ninth embodiment of the present invention is an improvement of the AI ​​response output device 10010 or the AI ​​response output system described in the drawings of the first to eighth embodiments. Specifically, this is an example in which, in the response generation process of the AI ​​response output device 10010, the client application more suitably controls the insertion of backchannel responses by interrupting the natural language sentences in the response output of a large-scale language model controlled by an LLM application. In this embodiment, differences from these embodiments will be described, and repeated description of configurations similar to these embodiments will be omitted. Note that the LLM application in the description of this embodiment may be a server LLM application or a local LLM application.

[0333] In the conversation examples shown in Figures 9F and 9G, the client application starts outputting a canned response for a backchannel immediately before the start of inference from the large-scale language model. Also, in the conversation examples shown in Figures 10C and 10D, the client application outputs a canned response for a network outage when responding to a user during a network outage. Alternatively, the client application may output a canned response by interrupting a natural language sentence in the response output of the large-scale language model. For example, if the natural language sentence in the response output of the large-scale language model is longer than a predetermined number of characters, the response output may feel more natural to the user if a canned response for a backchannel can be inserted. In this case, the timing at which to insert the canned response for a backchannel in the natural language sentence in the response output of the large-scale language model becomes an issue. Therefore, the client application of the AI ​​response output device 10010 in the AI ​​response output system according to this embodiment determines the timing based on line breaks included in the natural language sentence in the response output of the large-scale language model. Specifically, the timing for inserting a backchannel response should be at least at the position of a line break symbol. Sentences in natural language are separated by commas and periods. Because a comma marks the middle of a sentence, it may not be natural to insert a backchannel response at that position. A period marks the end of a sentence, so it is preferable to insert a backchannel response at that position. However, there are cases where the sentence with the period and the sentence following the period are closely related in meaning in the context. In such cases, it may not be natural to insert a backchannel response at the position of a period between the previous and next sentences. In contrast, when multiple sentences are followed by a line break symbol, the position where the line break symbol is inserted signifies a temporary break in the continuity of the series of sentences. Therefore, the position where the line break symbol is inserted is a more natural and suitable timing for inserting a backchannel response. Furthermore, when two consecutive line breaks are inserted, the continuity of the text is temporarily interrupted and an empty line is inserted. This means that the continuity of the text is more strongly interrupted than when only one line break is inserted.Therefore, a position where there are two or more consecutive line breaks is a more natural and suitable timing to insert a backchannel.

[0334] A specific example of the response generation process of the AI ​​response output device 10010 according to the ninth embodiment of the present invention will be described with reference to FIGS. 12A to 12C.

[0335] First, FIG. 12A shows an example of a conversation in which a fixed phrase response is not inserted into the natural language sentence of the response output of the large-scale language model during the response generation process of the AI ​​response output device 10010. First, the client application outputs a fixed phrase response 1, a greeting. Next, the user inputs user instruction sentence 1, including a question about a Japanese ukiyo-e artist, into the AI ​​response output device 10010. The client application sends the input user instruction sentence 1 to the LLM application. Upon receiving user instruction sentence 1, the LLM application causes the large-scale language model to perform inference based on user instruction sentence 1 at the timing indicated by star mark 9005. The natural language sentence of the LLM response generated by the large-scale language model through this inference is sent from the LLM application to the client application. The client application outputs the natural language sentence of the LLM response to the user. Note that in this embodiment, for ease of explanation, line break symbols in the response sentence, which were omitted in the previous embodiments, are clearly shown in the figure. The line break symbols indicate line breaks that exist as data in the natural language sentence.

[0336] Next, using the conversation example in Figure 12A as the original text, we will explain, using Figures 12B and 12C, an example of a conversation in which the client application of the AI ​​response output device 10010 performs control using various setting information stored in the table information in Figure 9E, and further performs interrupt insertion processing of a standardized response into the middle of a natural language sentence in the response output of a large-scale language model.

[0337] First, FIG. 12B shows an example of a conversation in which a client application controls the AI ​​response output device 10010 to output an AI response as a response from the character Necco. First, in FIG. 12B, fixed phrase response 1 is a fixed phrase for the character Necco's greeting. User instruction sentence 1 is the same as in FIG. 12A. In the example of FIG. 12B, the LLM application preset instructions shown in the table information of FIG. 9E are applied to the LLM application. Similarly to FIG. 12A, the LLM application that receives user instruction sentence 1 causes the large-scale language model to perform inference based on user instruction sentence 1 at the timing of star mark 9005. Because the LLM application preset instructions shown in the table information of FIG. 9E are applied to this inference, the response generated by this inference uses characteristic expressions (keywords) that reflect the personality of the character Necco. For example, this can be achieved by adding "meow" to the end of a sentence. Here, it is assumed that the LLM response (division part 1) and LLM response (division part 2) shown in FIG. 12B are generated in one go without being divided in the inference of the large-scale language model executed at the timing of star mark 9005. The LLM application sends a response sentence including the LLM response (division part 1) and the LLM response (division part 2) to the client application. The client application that receives the response sentence determines the position within the response sentence at which to insert a standard backchannel response. The client application divides the response sentence into the LLM response (division part 1) and the LLM response (division part 2) at that position, as shown in the figure.

[0338] The client application first outputs the LLM response (divided part 1) to the user. Next, the client application inserts the character Necco's standard backchannel phrase "Um, meow" at the divided position and outputs it to the user. Next, the client application outputs the LLM response (divided part 2) to the user. This allows the user to recognize that backchannels are naturally included in the series of responses output from the client application of the AI ​​response output device 10010. Furthermore, the backchannels also use characteristic expressions (keywords) that show the personality of the character Necco, so the user can feel that the backchannels are natural.

[0339] Here, an example of a process in which a client application determines the position at which to split an LLM response and insert a standard backchannel response will be described. First, the client application acquires the full text of the LLM response from the LLM application. After determining the length of the full text of the LLM response, the client application determines not to split the LLM response and not to insert a standard backchannel response if the length is less than a predetermined length. This is to avoid the unnatural sound that would result from splitting a short LLM response and inserting a standard backchannel response. The predetermined length that serves as the threshold for this determination may be referred to as the period during which standard backchannel response responses are not inserted. Note that the unit of length used in this determination process may be the number of characters, the number of words, or the number of tokens. Next, if the length of the full text of the LLM response exceeds a predetermined length, the client application determines that it is acceptable to insert a standard backchannel response. Next, the client application analyzes the text from the beginning of the full text of the LLM response and determines the position at which to split the full text of the LLM response using a predetermined condition. An example of the predetermined condition is a position where there are two consecutive line breaks. After finding the candidate position, the client application may further randomly determine whether to split the entire text of the LLM response using a probability process using a random number. For example, a random number may be generated, and if the remainder when the generated random number is divided by a predetermined number, such as 3, is 1, the client application may determine to split the entire text of the LLM response. This allows the entire text of the LLM response to be randomly split with a one-in-three probability. The client application splits the entire text of the LLM response into two parts and outputs the next part of the LLM response. Then, it inserts and outputs a standard backchannel response at the split position. Next, the client application performs a process to determine the position to insert the standard backchannel response for the remaining part of the LLM response. This process can be performed by changing the target of the process to determine the position to insert the standard backchannel response for the entire text of the LLM response from the entire text of the LLM response to the remaining part of the LLM response. When the client application repeats this process, the process of splitting the LLM response and the output of the next part of the LLM response are repeated.After this, when the length of the remaining divided parts of the LLM response is equal to or less than the above-mentioned fixed phrase non-insertion period, the client application ends the process of determining whether or not division is necessary, and finally outputs the remaining LLM response. As a result, the entire LLM response originally obtained by the client application ...

Claims

1. A response output device, an input interface for accepting user input; A control unit; A memory unit; an output interface for outputting a response to a user; Equipped with the control unit is capable of executing a client application capable of transmitting and receiving information to and from a server external to the response output device or a large-scale language model application that controls a large-scale language model stored in the response output device; the client application is capable of generating instruction sentences for the large-scale language model based on user input received via the input interface, capable of sending control information different from the instruction sentences to the large-scale language model application, capable of sending the instruction sentences to the large-scale language model application, capable of receiving response sentences from the large-scale language model application which are results of inference performed by the large-scale language model, and capable of outputting a response based on the response sentences to a user via the output interface; The storage unit stores settings related to the characteristics of the character's conversation. Response output device.

2. The response output device according to claim 1, the client application transmits control information including a stationary instruction based on a setting related to the character's conversation characteristics to the large-scale language model application, and receives a response sentence resulting from the large-scale language model application controlling the large-scale language model to execute inference based on the stationary instruction based on the setting related to the character's conversation characteristics in the control information and an instruction sentence generated based on a user input received via the input interface; the client application controls outputting, to the user from the output interface, standard phrases associated with the character read from the storage unit based on settings related to the character's conversation characteristics and the response sentences received from the large-scale language model application as responses of the character in a series of conversations; Response output device.

3. The response output device according to claim 2, The settings related to the characteristics of the character's conversation include settings related to standard phrases for greetings or backchannel responses of the character, and settings related to constant instructions for causing the large-scale language model to output the character's catchphrases or backchannel responses. Response output device.

4. The response output device according to claim 3, The setting related to the constant instruction for causing the large-scale language model to output the character's catchphrases or backchannels includes an instruction for causing the large-scale language model to output keywords that are common to keywords that indicate the individuality of the character and are included in the setting related to the character's greetings or backchannels. Response output device.

5. 5. The response output device according to claim 4, The common keywords that indicate the individuality of the character include the ending expressions or first-person expressions among the characteristics of the character's conversation. Response output device.

6. 6. The response output device according to claim 5, the client application is capable of transmitting and receiving information to and from both a server large-scale language model application that controls a server large-scale language model stored in a server external to the response output device and a local large-scale language model application that controls a local large-scale language model stored inside the response output device; The settings related to the characteristics of the character's conversation stored in the storage unit include a setting related to a standard greeting or backchannel response of the character, a setting related to a first regular instruction for causing the server large-scale language model to output the character's catchphrase or backchannel response, and a setting related to a second regular instruction for causing the local large-scale language model to output the character's catchphrase or backchannel response. Response output device.

7. 7. The response output device according to claim 6, In the settings relating to the characteristics of the character's conversation stored in the storage unit, a setting relating to a first steady instruction for causing the server large-scale language model to output the character's catchphrases or backchannels, and a setting relating to a second steady instruction for causing the local large-scale language model to output the character's catchphrases or backchannels, both contain keywords that are common to keywords that indicate the personality of the character and are included in the settings relating to the character's standard greetings or backchannels. Response output device.

8. The response output device according to claim 7, The settings related to the conversation characteristics of the character stored in the storage unit include settings related to the conversation characteristics of each of a plurality of different characters, When the response output device performs a process of switching the character that outputs the response, the client application switches the setting related to the conversation characteristics of the character in accordance with the character switching process. Response output device.

9. The response output device according to claim 2, The setting regarding the characteristics of the character's conversation includes a setting regarding a fixed phrase of the character's response, the client application receives the user input via the input interface, and then outputs a backchannel response based on a set of standard backchannel responses for the character to the user via the output interface at a timing before the large-scale language model executes inference; Response output device.

10. The response output device according to claim 1, The settings related to the characteristics of the character's conversation stored in the storage unit include information on a constant instruction for outputting a symbol indicating the position of the character's backchannel in inference of the large-scale language model. Response output device.

11. The response output device according to claim 2, the client application performs a division process for dividing the response sentence received from the large-scale language model application and a process for inserting a fixed phrase associated with the character read from the storage unit between the divided phrases, and controls output of the character's response to the user from the output interface; Response output device.

12. The response output device according to claim 11, the client application determines the position at which to divide the response sentence based on a condition regarding a line break symbol included in the response sentence; Response output device.

13. The response output device according to claim 12, the client application determines a position where two consecutive line feed symbols are included in the response sentence as a candidate position for dividing the response sentence; Response output device.

14. The response output device according to claim 2, The settings related to the conversation characteristics of the character stored in the storage unit include a setting related to a reading speed of characters when the character reads out a standard phrase associated with the character, when outputting the response sentence received from the large-scale language model application to a user, the client application performs speed adjustment control to adjust the reading speed of the response sentence to approach the reading speed of characters when the character stored in the storage unit reads the standard phrase. Response output device.

15. The response output device according to claim 14, The settings related to the conversation characteristics of the character stored in the storage unit include settings related to the conversation characteristics of each of a plurality of different characters, When the response output device performs a process of switching the character that outputs the response, the client application switches the setting related to the conversation characteristics of the character in accordance with the character switching process. Response output device.

16. The response output device according to claim 15, when the setting related to the conversation characteristics of the character is switched in accordance with the character switching process, the client application changes the length of a silent period in the reading of characters at an insertion position of a space symbol or a line feed symbol included in the response sentence, in response to the character switching process, when outputting the response sentence received from the large-scale language model application to a user. Response output device.

17. The response output device according to claim 1, the settings related to the conversation characteristics of the character stored in the storage unit include settings related to keywords included in responses of the character and adjustment of intonation for the keywords; the client application performs control to adjust the intonation for the keyword when outputting the character's response to the user based on the setting related to the intonation adjustment; Response output device.

18. The response output device according to claim 17, The settings related to the conversation characteristics of the character stored in the storage unit include settings related to the conversation characteristics of each of a plurality of different characters, When the response output device performs a process of switching the character that outputs the response, the client application switches the setting related to the conversation characteristics of the character in accordance with the character switching process. Response output device.

19. The response output device according to claim 17, the client application uses a condition of a positional relationship between a keyword and a special character included in the response sentence received from the large-scale language model application as a condition for adjusting the intonation of the keyword included in the response sentence received from the large-scale language model application. Response output device.

20. The response output device according to claim 17, the client application transmits to the large-scale language model application a stationary instruction for inserting a predetermined keyword as a feature of the character's conversation and a stationary instruction for outputting an identification symbol indicating the position where the keyword has been inserted; the client application identifies the keyword, the intonation of which is to be adjusted, from among characters in the response sentence based on the identification symbol included in the response sentence received from the large-scale language model application; Response output device.

Citation Information

Patent Citations

  • Structural unit for tank construction

    JP1977008512A