Response output device and response output system
The response output device and system improve user interaction by integrating user input-driven search and content registration with large-scale language models, offering enhanced engagement and resource-efficient multimodal processing.
Patent Information
- Application Number
- JP2024069301
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-11-04
AI Technical Summary
Existing response output technologies using artificial intelligence do not adequately consider suitable configurations for user interaction.
A response output device and system that includes an input unit, control unit, and large-scale language model, enabling search processes and content file registration based on user input, with options for local or external model processing and multimodal capabilities.
Provides a more suitable and interactive response output technique, enhancing user engagement and resource efficiency through shared learning and multimodal processing.
Smart Images

Figure 2025165280000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a response output device and a response output system. [Background technology]
[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special table 2019-528512 publication Summary of the Invention [Problem to be solved by the invention]
[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.
[0005] An object of the present invention is to provide a more suitable response output technique. [Means for solving the problem]
[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problem, but one example thereof may be a response output device including an input unit and a control unit that generates a response request based on a user input from the input unit, transmits the response request to a large-scale language model, and outputs a response from the large-scale language model in response to the response request, wherein the control unit performs a search process based on search conditions specified by the user, extracts registration candidates for a content file to be registered in the response request, and registers a content file selected by the user from the extracted registration candidates in the response request. [Effects of the Invention]
[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments. [Brief explanation of the drawings]
[0008] [Figure 1A] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 1B] 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 1C] 1 is a diagram showing an example of the operation of an AI response output device and system according to an embodiment of the present invention; [Figure 2A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2C] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2I]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2J] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2K] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2L] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3C] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3I] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4A]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 5A] FIG. 2 is an explanatory diagram of an example of the operation of the artificial intelligence response output device according to an embodiment of the present invention. [Figure 5B] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5C] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5D] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 6] FIG. 10 is an explanatory diagram of an example of a response generation process of the AI response output device according to an embodiment of the present invention. [Figure 7A] FIG. 10 is an explanatory diagram of an example of a response output process of the artificial intelligence response output device according to an embodiment of the present invention. [Figure 7B] FIG. 2 is an explanatory diagram of an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 7C] FIG. 2 is an explanatory diagram of an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 7D] FIG. 2 is an explanatory diagram of an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 7E] FIG. 2 is an explanatory diagram of an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 8] 1 is a flowchart illustrating an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 9A] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 9B] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 10] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 11] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 12A] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 12B] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 13A] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 13B] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 14] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 15] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 16] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 17A] 1 is a diagram illustrating a relevance evaluation performed by an AI response output device according to an embodiment of the present invention. FIG. [Figure 17B] 1 is a diagram illustrating a relevance evaluation performed by an AI response output device according to an embodiment of the present invention. FIG. [Figure 18] 10 is a diagram illustrating an example of result information of a file relevance evaluation performed by an artificial intelligence response output device according to an embodiment of the present invention. FIG. [Figure 19A] FIG. 10 is a diagram illustrating an example of a search process performed by an artificial intelligence response output device according to an embodiment of the present invention. [Figure 19B] FIG. 10 is a diagram illustrating an example of a search process performed by an artificial intelligence response output device according to an embodiment of the present invention. [Figure 20A] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 20B] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 21A]FIG. 10 is a diagram illustrating an example of a search process performed by an artificial intelligence response output device according to an embodiment of the present invention. [Figure 21B] FIG. 10 is a diagram illustrating an example of a search process performed by an artificial intelligence response output device according to an embodiment of the present invention. [Figure 22A] FIG. 10 is a diagram illustrating an example of a search process performed by an artificial intelligence response output device according to an embodiment of the present invention. [Figure 22B] FIG. 10 is a diagram illustrating an example of a search process performed by an artificial intelligence response output device according to an embodiment of the present invention. [Figure 23A] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 23B] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 24A] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 24B] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 25] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 26] 1 is a flowchart illustrating an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 27] 1 is a flowchart illustrating an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 28] 1 is a diagram showing an example of a display of an artificial intelligence response output device according to an embodiment of the present invention; [Figure 29] 1 is a flowchart illustrating an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 30] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 31] 1 is a flowchart illustrating an example of processing of an artificial intelligence response output device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.
[0010] Note that if the AI response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI response output device has an audio output function, it may be referred to as an audio output device, or simply as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI response output device or the display output of the AI response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.
[0011] Example 1 As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.
[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system in which the AI response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.
[0013] In the example of FIG. 1A, the AI response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Furthermore, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which the pixels emit light themselves. Furthermore, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.
[0014] 1A, the audio output unit 1140 provided in the AI response output device 10010 is configured with a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. By audio input from the microphone 1139 or operation input from the user via an operation input unit (described later), the AI response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.
[0015] The AI response output device 10010 may be provided with a local large-scale language model within the AI response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or an audio output from the audio output unit 1140.
[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.
[0017] Alternatively, the AI response output device 10010 may also include a local large-scale language model, and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.
[0018] The configuration when the AI response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions, or may go via a router or repeater. The AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as one system.
[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
[0020] 1A shows an example in which a display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of FIG. 1A, the prompt display area 10051 displays an icon 10052 representing a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of FIG. 1A, the artificial intelligence response display area 10061 displays an icon 10062 representing an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.
[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment. These large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in these artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called Transformer. Reference 1, for example, has been published as an example of the learning of these models.
[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001 and the AI response output device 10010. In this way, instead of performing pre-learning for each application or terminal, replicating the large-scale language model, which is the base model generated by large-scale pre-learning, and using it on individual servers and terminals allows the resources used for learning to be shared, resulting in good resource efficiency.
[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.
[0025] Furthermore, large-scale language models can pre-train natural languages and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. FIG. 1A illustrates a large-scale language model server 20001, which is a server having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are AI models generated by large-scale pre-training on natural language and types of information other than natural language text information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. In addition, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.
[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.
[0028] The AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a nonvolatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, and an imaging unit 1180. The AI response output device 10010 may have a large screen, such as a monitor or television.
[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire the user input that serves as the basis for a prompt to the large-scale language model, which is the artificial intelligence.
[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN-type communication method. This allows the AI response output device 10010 to communicate with various servers connected to the Internet 19000.
[0031] The AI response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.
[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.
[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller or the like, and inputs a signal regarding an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. In addition to a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may also be used by, for example, an administrator to operate the AI response output device 10010. By the user's operation input via the operation input unit 1107, the AI response output device 10010 can acquire a user input that serves as the basis for a command sentence (prompt) to a large-scale language model, which is an AI. Note that a modified configuration in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107 is also possible.
[0034] The video signal input unit 1131 is connected to an external video output device and inputs video data. The video signal input unit 1131 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces, etc.
[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.
[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
[0037] The microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.
[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.
[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.
[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 controls video switching, such as which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also control image processing of the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling, which enlarges, reduces, or deforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.
[0042] The nonvolatile memory 1108 stores various data used by the AI response output device 10010. The data stored in the nonvolatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.
[0043] The local LLM processing unit 10028 has a memory capable of holding a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI response output device 10010.
[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, which is a speaker or the like.
[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 of the AI response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and responses are obtained from these large-scale language models. All of this control can be performed by the control unit 1110.
[0046] The storage unit 1170 may also store a fixed response phrase database (which may be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction sentences from the AI response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI response output device 10010.
[0047] Furthermore, in the example of the fixed response phrases in the database shown in FIG. 1C, if multiple fixed response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the fixed response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the fixed response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.
[0048] Next, an example of condition number 5 shown in Fig. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I might not know about that." By responding in this way, the user can be prompted to input again, and the corrected user input can be waited for.
[0049] Next, an example of condition number 6 shown in Fig. 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in Fig. 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the fixed response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.
[0050] The AI response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).
[0051] 1C described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the response template database (response template DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the AI response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0052] In the above explanation, it has been explained that the AI response output device 10010 has a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.
[0053] 1A and 1B, an example has been described in which the AI response output device 10010 includes the display unit 10011. However, the AI response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI may be configured to accept input from a user to the AI via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the input from the user via the voice output unit 1140.
[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference of the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.
[0055] <Example 2> Next, as a second embodiment of the present invention, an example will be described in which the AI response output device 10010 described in the first embodiment is connected to the Internet, and is connected to a server equipped with a large-scale language model AI via the Internet to perform operation. In this embodiment, differences from the first embodiment will be described, and repeated explanations of the same configurations as those of these embodiments will be omitted.
[0056] An example of a connection state between an AI response output device 10010 and a large-scale language model server 19001 according to the second embodiment of the present invention will be described with reference to FIG. 2A. The AI response output device 10010 according to the second embodiment may be called a character conversation device. A system including the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment may be called a character conversation system. A video of a character 19051 is displayed on a display unit 10011 displayed by the AI response output device 10010. The video of the character 19051 is generated by rendering a 3D model of the character in a virtual space.
[0057] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can serve as an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this embodiment may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0058] In the example of FIG. 2A, the audio output unit 1140 provided in the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A, an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication is also acceptable. The communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions. The AI response output device 10010 can communicate with a large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with a second server 19002 different from the large-scale language model server 19001 via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 may be considered as one system.
[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2B. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that Figure 2B does not illustrate communication paths such as the Internet 19000 shown in Figure 2A. Figure 2B also illustrates a user 230 of the artificial intelligence response output device 10010.
[0060] Here, we will explain the sequence of operations of the AI response output device 10010. The AI response output device 10010 loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0061] First, the AI response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is in natural language. Note that the extraction of the text of the words spoken by the user 230 may be performed continuously for all words, or may start when the user utters a word within a predetermined period after a trigger keyword. For example, the trigger keyword may be when the user utters "hello" followed by the character's name. For example, if the name of the character 19051 is "Koto," then "Hello, Koto!" may be the trigger keyword.
[0062] The character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230 and transmits the prompt to the large-scale language model server 19001 using an API. Here, the prompt may be metadata containing information written in a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON. The prompt stores text information in a natural language as a main message. The prompts transmitted from the AI response output device 10010 to the large-scale language model server 19001 include setting prompts that store instructions such as initial settings, and user prompts that reflect instructions from the user. Type identification information identifying whether the prompt is a setting prompt or a user prompt may be stored in a portion of the prompt other than the main message. When the character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230, it creates the user prompt and sends it to the large-scale language model server 19001.
[0063] Next, the large-scale language model of the artificial intelligence of the large-scale language model server 19001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response including natural language text information based on the result. The large-scale language model server 19001 uses an API to send the response to the artificial intelligence response output device 10010. The response stores the natural language text information as a main message. Here, the response may be metadata storing information written in the same format as the instruction (e.g., a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON). If the response uses the same format as the instruction, type identification information may be stored outside the main message to indicate that the information is a different type from the initial setting instruction and the user instruction. For example, information indicating that the response is a response from the large-scale language model may be stored.
[0064] Next, the AI response output device 10010 receives a response from the large-scale language model server 19001 and extracts the natural language text information stored as the main message in the response. The character operation program of the AI response output device 10010 uses speech synthesis technology to generate a natural language voice as a response to the user based on the natural language text information extracted from the response, and outputs the voice from the speaker, i.e., the voice output unit 1140, so that it sounds as if it were the voice of the character 19051. This process may also be referred to as the character's "speaking."
[0065] Conversation examples 1 to 5 in Fig. 2C show specific examples of response voices from character 19051 in response to words from user 230, which are generated by the above-described processing of AI response output device 10010 and large-scale language model server 19001. In this way, user 230 can converse with character 19051 as if it were a real person.
[0066] 2B or a system including the AI response output device 10010, there is no need to install a large-scale language model, which requires a huge amount of data and computational resources for learning, in the AI response output device 10010. In addition, the advanced natural language processing capabilities of the large-scale language model can be utilized via an API, and when a user speaks to a character, a more appropriate response can be given to the user, enabling a more appropriate conversation to be held.
[0067] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2D. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2D shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0068] FIG. 2D also shows the exchange of instructions and responses in chronological order, from the display setting instruction, the first round of user instructions and their responses, to the fourth round of user instructions and their responses.
[0069] As shown in FIG. 2D , the setting directive can be used to initially instruct the large-scale language model of the artificial intelligence of the large-scale language model server 19001, such as the name of the large-scale language model itself, the role to be played, and conversation characteristics. The user's name can also be understood as the initial setting. This allows the large-scale language model to generate responses from the first round onward while respecting the role. Then, to a user who hears the voice of the character 19051 based on the responses from the first round onward, the character 19051 will seem to have the setting and personality of the person described in the setting directive. Furthermore, the large-scale language model server 19001 according to this embodiment is equipped with a memory that stores the content of the conversation until the end of the series of conversations, and is configured to store a series of user directives and their responses and then generate responses. This allows for a conversation such as that shown in FIG. 2D to be realized.
[0070] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2E shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.
[0071] 2E shows an example of a case where, after the series of conversations shown in FIG. 2D has ended, the user 230 speaks to the character 19051 again to start a new conversation. In FIG. 2E, the exchange of instructions and responses is shown in chronological order, from the first round of user instructions and their responses to the third round of user instructions and their responses.
[0072] Here, "termination" of the "continuation of a series of conversations" refers to a process in which, when a predetermined condition is met, the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversation that the large-scale language model server 19001 has maintained while the series of conversations was continuing. An example of the predetermined condition is, for example, when the AI response output device 10010 issues an instruction to the large-scale language model server 19001 to "terminate" the "continuation of a series of conversations." Another example of the predetermined condition is, for example, when a predetermined time or more has passed since the AI response output device 10010 stopped sending instruction statements to the large-scale language model server 19001 regarding the series of conversations (timeout). Another example of the predetermined condition is when, after authentication processing has been performed in the connection between the AI response output device 10010 and the large-scale language model server 19001, the authentication processing is terminated due to factors such as communication disconnection or the AI response output device 10010 being powered off while exchanging the instruction statements and responses.
[0073] Note that when the "continuation of a series of conversations" "ends," the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversations that it had maintained while the series of conversations was continuing. Therefore, even though the conversation shown in FIG. 2E takes place after the series of conversations shown in FIG. 2D, the server's response to the user's instruction is a response with content that does not include any memory of the character's name set in the large-scale language model, the role to be played, the characteristics of the conversation, or the user's name, which were included in the setting instruction shown in FIG. 2D. Similarly, the conversation shown in FIG. 2E is a response with content that does not include any memory of the series of conversations shown in FIG. 2D. In other words, with the "end" of the "continuation of a series of conversations" shown in FIG. 2D, the conversation in FIG. 2E starts from an initialized state of the large-scale language model of the artificial intelligence of the large-scale language model server 19001.
[0074] This causes the user 230 to feel as if the character 19051 has lost its memory or is a completely different person. From the user 230's perspective, the character's response feels very strange, resulting in a feeling of loneliness and disappointment. Such behavior poses a problem in that it is not possible to ensure the consistency of the settings and memories of the character 19051, such as its name, role, conversational characteristics, and personality, displayed on the AI response output device 10010.
[0075] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2F. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2F shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.
[0076] FIG. 2F illustrates an example of a case in which the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations illustrated in FIG. 2D has ended. Unlike the process illustrated in FIG. 2E, in the process illustrated in FIG. 2F, when starting a new conversation, the AI response output device 10010 transmits a setting instruction as the first instruction to the large-scale language model server 19001. The setting instruction stores the same natural language text as the initial setting instruction of FIG. 2D. This may be referred to as a reset text. The setting instruction is followed by natural language text describing the history of past conversations. This may be referred to as a conversation history text. The AI response output device 10010 may record the history of past conversations as natural language text information in the storage unit 1170 while the series of conversations illustrated in FIG. 2D is continuing, linking the history of the conversations to information on the date and time of the conversations. If there are conversations on different dates, each conversation may be recorded linked to date and time information, and the conversation history may be accumulated. When generating a setting instruction for the first instruction of a later conversation, such as that shown in Figure 2F, the natural language text information of the conversation recorded in storage unit 1170 and the information on the date and time when the conversation took place can be read out and used to generate the setting instruction.
[0077] When natural language text information of past conversation history is used to generate the setting instruction sentence, the format can be determined freely to a certain extent because it is data to be sent to a large-scale language model, but as shown in Figure 2F, it is sufficient to prepare prefixes and suffixes in natural language, such as "I talked about the following on ____ day of ____ month," or "You talked about the following on ____ day of ____ month," and combine these with the natural language text information of the recorded conversation to generate the text of the setting instruction sentence. Also, information on the date and time of the conversation read from storage unit 1170 may be combined with the above-mentioned "____ day of ____ month" portion to form part of the text of the setting instruction sentence.
[0078] Even if user 230 speaks to character 19051 again to start a new conversation after a series of conversations has ended, by performing the above-described generation process and transmission process of the setting instruction sentence in Fig. 2F, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation. This is preferable because it is perceived by the user as ensuring the consistency of the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation.
[0079] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2G shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0080] Figure 2G shows an example of a series of conversations shown in Figure 2F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 2G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 2F, so repeated description is omitted.
[0081] As shown in the natural language text of the server response in the table of Figure 2F, by using the setting directive shown in Figure 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 reflects the settings and conversation history of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation. This is more preferable because it allows the user to recognize that the settings and memories of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation are more closely matched. Note that this allows the characters to be viewed as the same from the user's perspective, and may therefore be referred to as a pseudo-identity of the characters from the user's perspective.
[0082] Furthermore, from the user's perspective, they can share memories with the character, providing a more enjoyable character conversation experience.
[0083] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, FIG. 2H shows an example of the operation in which a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. A character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input input to the operation input unit 1107 or an operation detected by a touch operation input sensor of the display unit 10011.
[0084] In the example of FIG. 2H, in addition to character 19051 (named "Koto") used in the description of FIGS. 2A to 2G, character 19052 (named "Tom") and character 19053 (named "Necco") are shown. Character 19051 (named "Koto") and character 19052 (named "Tom") are human characters, and character 19053 (named "Necco") is a cat character. The display of characters displayed on display unit 10011 can be switched by switching to display on display unit 10011 an image generated by rendering a character in a different virtual 3D space for each character. The process for realizing the display of a rendered image of a 3D model of each character can be, for example, any of the first to third processing examples described in FIG. 15A. Furthermore, depending on the character, a dynamic 2D image may be displayed.
[0085] Furthermore, it is preferable that the character operation program executed by control unit 1110 also changes the synthetic voice used for the "speech" of each character when switching the display of the characters displayed on display unit 10011. This can be achieved by storing synthetic voice data of voice tones associated with each character in storage unit 1170 in advance, and performing synthetic voice change processing when switching the display of the character.
[0086] In the example of Fig. 2H, the user 230 is configured to be able to converse with any of the characters. In the AI response output device 10010 of Fig. 2H, each of these characters is set with a different role, name, conversational characteristics, personality, etc. Also, the memories of each character based on the conversation history are managed as different for each character.
[0087] Therefore, the AI response output device 10010 constructs a database shown in FIG. 2I in the storage unit 1170, and manages the character settings and the character conversation history using this database.
[0088] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Fig. 2I. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.
[0089] A character operation program executed by the control unit 1110 of the AI response output device 10010 constructs the database 19200 in, for example, the storage unit 1170. The character ID is an identification number that identifies each of multiple characters that can be displayed on the AI response output device 10010, and may be a natural number or may use alphabets, etc. The name is data of the name of each of multiple characters that can be displayed on the AI response output device 10010.
[0090] The initial setting instruction is text information in a natural language that explains the settings such as the role, name, conversational features, or personality of each of multiple characters that can be displayed on the AI response output device 10010. The initial setting instruction is natural language text information that is the main data of the setting instruction sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the description can be read as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0091] The conversation histories, which continue as conversation histories 1, 2, ..., are records of conversations between each character and the user, and are recorded separately for each character. The conversation histories will be included in the natural language text information, which is the main data of the setting instruction statement sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the conversation histories be readable as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0092] When the character displayed on the display unit 10011 of the AI response output device 10010 is switched, the character operation program executed by the control unit 1110 of the AI response output device 10010 uses the database 19200 of Figure 2I to select and switch the initial setting instruction statement and conversation history used for the natural language text information that is the main data of the setting instruction statement transmitted from the AI response output device 10010 to the large-scale language model server 19001 so as to correspond to the character displayed on the display unit 10011 of the AI response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the conversation history in the conversation history area of the database 19200 of Figure 2I that corresponds to the character displayed on the display unit 10011.
[0093] By using the database 19200 in this manner, the character operation program executed by the control unit 1110 of the AI response output device 10010 establishes a conversation between the user 230 and the character using character utterances that utilize responses from the same AI large-scale language model of the same large-scale language model server 19001. From the user's perspective, the uniqueness of each character's personality and other settings is preserved, and it appears as if each character's unique conversational memories continue. This is more preferable because it appears as if the identity of each character's settings and memories, such as their role, name, conversational characteristics, or personality, from the time of the previous conversation has been more consistently maintained. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.
[0094] Therefore, even when the AI response output device 10010 is configured to switch the character displayed on the display unit 10011 from among multiple character candidates, the operation using the database 19200 described above reduces the sense of discomfort felt by the user when conversing with each character, and the user can share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.
[0095] Note that if the initial setting instructions for multiple characters cannot be edited by the user, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained close to the intentions of the provider of the AI response output device 10010 or the creator of the character content. Alternatively, the initial setting instructions for each character may be edited by the user in response to input from the operation input unit 1107 or the like. In this case, the character's role, name, conversational characteristics, or personality can be set to a preferred setting, allowing the user to converse with a character that they have individually set. In this case, the character's 3D model, its rendered image, and the type of synthesized voice for the character may be replaced accordingly.
[0096] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2J. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service at a lower cost using the character conversation device based on the artificial intelligence response output device 10010 and the character conversation system based on the artificial intelligence response output device 10010 and the large-scale language model server 19001 will be described.
[0097] As explained in Figure 2B, it is extremely resource inefficient to train a large-scale language model at this level of artificial intelligence by limiting it to a specific application. Therefore, it is more resource efficient to perform large-scale training to generate a foundation model that can be applied to a variety of applications, and then use it on various devices via an API (Application Programming Interface). In such cases, providers of large-scale language models often recover the costs incurred in training the large-scale language model from device users as API usage fees. In natural language models, API usage fees are often charged based on the number of tokens, which are units of words that separate sentences, processed.
[0098] Therefore, in the artificial intelligence response output device 10010 of the second embodiment of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010 and a character conversation service using a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.
[0099] For example, by using the processing and configuration of Examples 1 to 3 shown in the table of Figure 2J, it is technically possible to reduce the number of tokens in the natural language text information transmitted between the AI response output device 10010 and the large-scale language model server 19001 using an API.
[0100] Example 1 is an example of a method for reducing the number of tokens in the conversation history text stored in the API setting directive and transmitted, in which the conversation history text is shortened using a document summarization process to reduce the number of tokens. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. The text summarization may be performed at the start of the next conversation, but it is more time-efficient to perform it at the end of the "series of conversations."
[0101] Furthermore, the text summarization process may be requested from the large-scale language model itself of the large-scale language model server 19001. However, in this case, the effect of saving the number of tokens is small. Therefore, for example, if the second server 19002 provides text summarization process for natural language via an API at a lower cost than the large-scale language model of the large-scale language model server 19001, the text summarization process can be requested from the second server 19002 via the API, and the text summary of the conversation history can be stored in a setting instruction statement for the large-scale language model server 19001 and transmitted.
[0102] Furthermore, if it is only text summarization processing, it can be performed on the terminal side, and text summarization can be performed by the control unit 1110 executing a document summarization program loaded in the memory 1109 of the AI response output device 10010. In this case, the effect of saving the number of tokens is high. Furthermore, even if the conversation history becomes long, if an upper limit on the number of characters after summary is specified in the text summarization processing, the upper limit on the text length of the conversation history is determined, so that an upper limit on the number of tokens can be set and tokens can be saved.
[0103] In addition, since the text information of the character's initial settings, such as the character's role, name, conversational characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information of the character's initial setting instructions and reduce the number of tokens of the text information in the conversation history.
[0104] The processing described in Example 1 may be performed by the character movement program executed by the control unit 1110 controlling each unit.
[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored in the API setting instruction and transmitted. For example, the number of tokens is reduced by deleting the oldest conversation history among the conversation histories with characters recorded in the storage unit 1170. Specifying an upper limit on the number of characters in the conversation history determines the upper limit on the length of the conversation history text, thereby setting an upper limit on the number of tokens and enabling token conservation. Alternatively, a predetermined period of the conversation history may be specified and conversation history exceeding that period may be deleted. This also enables token conservation. Note that, in Example 2, the text information for the character's initial setting, such as the character's role, name, conversation characteristics, or personality, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the text information in the character's initial setting instruction and reduce the number of tokens in the text information for the conversation history.
[0106] The processing described in Example 2 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0107] Example 3 is a method of reducing the number of tokens by reducing the frequency of sending setting instruction sentences using an API. Specifically, even after the device is powered on, after the displayed character is switched, or after the video settings and synthetic voice settings of the displayed character are completed, setting instruction sentences are not sent in advance, and only when the control unit 1110 determines that natural language text information included in the user's speech picked up by the microphone 1139 is text information that should use a large-scale language model of artificial intelligence is the setting instruction sentence sent to the large-scale language model server 19001, thereby reducing the frequency of sending setting instruction sentences to the large-scale language model server 19001 and reducing the number of tokens.
[0108] Specifically, for example, after the device is powered on or after an operation input to switch the displayed character is made, character 19051 (named "Koto") is displayed on display unit 10011 as shown in FIG. 2H by display processing of display unit 10011 controlled by a character operation program executed by control unit 1110. At this time, for example, if a synthetic voice for the character 19051 to appear is stored and prepared in storage unit 1170 or the like, a synthetic voice for the character's appearance such as "Good morning, I'm Koto," "Hello, I'm Koto," or "Good evening, I'm Koto" may be output from the speaker, which is audio output unit 1140. At this time, the image of character 19051 has already been set as the image of the character to be displayed on display unit 10011, and the synthetic voice output from the speaker, which is audio output unit 1140, is set to the synthetic voice corresponding to character 19051.
[0109] Here, the inference process of the large-scale language model of the artificial intelligence in the large-scale language model server 19001 already described also takes time as the instruction becomes longer. In particular, if the setting instruction includes text information related to past conversation history, the number of tokens in the instruction increases, resulting in a particularly long inference process time. The setting instruction itself and its response are not output to the user 230. Based on the response to the user instruction following the setting instruction, a synthesized voice is output as the character's "utterance" from the speaker, which is the voice output unit 1140. In this case, it may seem preferable at first glance to send the setting instruction from the AI response output device 10010 to the large-scale language model server 19001 in advance to complete the inference process of the large-scale language model for the setting instruction in advance, as this would result in a faster response in the output of synthesized voice of the character's 19051's "utterance" after the user 230 speaks to the character 19051.
[0110] However, if the setting instruction is transmitted to the large-scale language model server 19001 before the user 230 speaks and the inference processing of the large-scale language model for the setting instruction is completed in advance, for example, the user 230 may turn off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or the user 230 may switch the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. In this case, the number of tokens processed in the inference processing of the large-scale language model by transmitting the setting instruction to the large-scale language model server 19001 in advance becomes the number of processed tokens for which the usage fee is wasted. This hinders the provision of a character conversation device using the AI response output device 10010 and a character conversation service using a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 at a lower cost to users.
[0111] Therefore, after the AI response output device 10010 is powered on (ON) or after an operation input to switch the displayed character is made, it is desirable that the AI response output device 10010, under the control of the character operation program executed by the control unit 1110, sets the image of character 19051 as the image of the character to be displayed on the display unit 10011, and sets the synthetic voice output from the speaker, which is the audio output unit 1140, to the synthetic voice corresponding to character 19051, continue not to send a setting instruction statement to the large-scale language model server 19001 until the user 230 recognizes that he or she is speaking to character 19051.
[0112] Here, the point in time at which it is recognized that user 230 is speaking to character 19051 may be, for example, the point in time at which the trigger keyword described in Fig. 2B is detected, or the point in time at which the text of the words spoken by user 230 is extracted. In this way, the number of processing tokens that waste usage fees can be reduced, and a character conversation service by a character conversation device using artificial intelligence response output device 10010 or a character conversation system using artificial intelligence response output device 10010 and large-scale language model server 19001 can be provided to users at a lower cost.
[0113] Furthermore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051 has passed, it is desirable to continue not sending setting instructions to the large-scale language model server 19001, for example, if the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Specifically, examples of preset keywords include keywords such as "try jumping" and "try dancing," which are keywords by which the user 230 requests the character 19051 to react, such as by animating the character 19051 to move or emitting synthetic voice. In this case, the character operation program executed by the control unit 1110 reads out motion data, animation video, and / or synthetic voice data corresponding to the reaction stored in the storage unit, corresponding to the character 19051, and / or synthetic voice data corresponding to the reaction, and uses these data to generate video to be displayed on the display unit 10011 and output synthetic voice from the speaker, which is the audio output unit 1140.
[0114] Such processing does not necessarily require inference processing of the large-scale language model of the large-scale language model server 19001. After the processing, if the user 230 turns off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or if the user 230 switches the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, if the setting instruction statement is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be the number of processing tokens that unnecessarily consumes the usage fee.
[0115] Therefore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051, it is desirable to continue not sending the setting instruction statement to the large-scale language model server 19001 until it is determined whether or not the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Only when it is determined that inference processing of a large-scale language model is required is it desirable to send the setting instruction statement to the large-scale language model server 19001 and proceed with the inference processing of the large-scale language model.
[0116] The processing described in Example 3 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0117] According to the method for reducing (saving) the number of processing tokens for a large-scale language model using the examples of Figure 2J described above, a character conversation device using the AI response output device 10010 and a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.
[0118] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2K. The example of FIG. 2K shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which has been described in each of FIGS. 2A to 2J, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063, which is a response from the large-scale language model, together with an image of a character 19051 on the display unit 10011. The text 10063, which is a response from the large-scale language model, may be displayed superimposed on top of the image of the character 19051, as shown in FIG. 2K. Alternatively, the text 10063, which is a response from the large-scale language model, may be displayed together with the image of the character 19051 without being superimposed on the image of the character 19051.
[0119] The display in Figure 2K is an example, but for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by audio.
[0120] Therefore, in this case, the control unit 1110 may perform control to start a display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, as shown in Fig. 2K. In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more suitably use the character conversation device (artificial intelligence response output device 10010). Note that the user 230 may be configured to manually switch ON / OFF the display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011.
[0121] Next, an example of a fixed response phrase database (fixed response phrase DB) in the character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in FIGS. 2H and 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition numbers and condition contents are the same as those in FIG. 1C. In the example of FIG. 2L, individual fixed response phrases are set for each of the multiple characters in response to these conditions. For example, fixed response phrases for each condition are stored for each of the three characters, Character 1: Koto, Character 2: Tom, and Character 3: Necco, as described in FIGS. 2H and 2I. The output control of fixed response phrases is the same as that in FIG. 1C, and therefore will not be repeated.
[0122] In the example of FIG. 2L, the control unit 1110 selects a corresponding fixed response phrase from a fixed response phrase database (fixed response phrase DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses the selected fixed response phrase for output control as a response uttered by the character. For example, in the example of the fixed response phrase database (fixed response phrase DB) in FIG. 2L, even under the same conditions, the fixed response phrases are changed to expressions or contents that correspond to the individuality of the character. This allows the character conversation device (artificial intelligence response output device 10010) to provide the user with conversation that corresponds to the individuality of the displayed character. The user can feel that each character has a more consistent personality. This makes it possible to realize a character conversation device (artificial intelligence response output device 10010) that gives multiple characters a more realistic presence.
[0123] The fixed response phrase database (fixed response phrase DB) of FIG. 2L described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 2L may also be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in each server. In this way, even if the AI response output device 10010 does not have a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0124] The character conversation device and character conversation system according to the second embodiment described above can reduce the sense of discomfort felt by the user from the conversation with the character displayed on the AI response output device 10010. Furthermore, the character conversation device and character conversation system according to the second embodiment can provide the character conversation service to the user at a lower cost.
[0125] In the above description of the second embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 19001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B, and the large-scale language model held by the local LLM processing unit 10028 may be used instead of the large-scale language model held by the large-scale language model server 19001. In this case, in the above description of the second embodiment, the large-scale language model held by the large-scale language model server 19001 may be read as the large-scale language model held by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).
[0126] In this case, too, it is possible to further reduce the sense of discomfort felt by the user from conversations with characters displayed on the AI response output device 10010. Note that when a large-scale language model held by the local LLM processing unit 10028 is used instead of the large-scale language model held by the large-scale language model server 19001, there is less need to consider usage fees according to the number of processed tokens, but by reducing the number of processed tokens even for the large-scale language model held by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power required for inference. In this case, it is possible to provide users with a character conversation service that consumes less power.
[0127] In the above description of the second embodiment, an example has been described in which a conversation history with a character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). Alternatively, a conversation history with a character may be recorded and stored in a second server 19002 or another cloud server connected to the Internet 19000. In this case, when a user and a character start a new conversation, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or another cloud server, acquires (downloads) a past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create an instruction for a large-scale language model. Specific methods for using past conversation history to create an instruction for a large-scale language model are as described in the figures of the second embodiment, and therefore repeated description will be omitted.
[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character's conversation history up to a predetermined point in time, such as every time a conversation between the user and the character takes place or when the conversation between the user and the character ends, to the second server 19002 or another cloud server. That is, the character conversation device (artificial intelligence response output device 10010) may upload the conversation history with the character to the second server 19002 or another cloud server at a predetermined timing, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or another cloud server and use it to generate an instruction sentence for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) that the user used the previous day and the character conversation device (artificial intelligence response output device 10010) that the user is about to use are different individual devices and can display the same character, and when the user has multiple conversations with the same character between the different individual devices at different times, it is possible to realize a conversation in which the character's memories are pseudo-carried over from the previous conversation, which is more convenient for the user.
[0129] The process described above in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with a character to the second server 19002 or another cloud server, thereby pseudo-taking over the character's memory, is also effective when handling the database 19200 including the conversation histories of multiple characters described in Figures 2H and 2I. In other words, if the database 19200 described in Figure 2I is configured to be uploaded and downloaded to the second server 19002 or another cloud server, not only for one character but for multiple characters, and between different individual devices, when a user has multiple conversations with each of the multiple characters at different times, it is possible to realize a conversation in which the memory of each character is pseudo-taken over from the previous conversation, which is more convenient for the user.
[0130] Example 3 Next, the third embodiment of the present invention is an improvement of the character conversation device (artificial intelligence response output device 10010) and the character conversation system explained in the drawings of the second embodiment. In this embodiment, differences from the second embodiment will be explained, and repeated explanations of the same configurations as those of the second embodiment will be omitted.
[0131] As in Example 2, the character in Example 3 can provide the user with the service of a large-scale language model, which is an artificial intelligence, and can be of assistance to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this example may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0132] An example of a character conversation device and a character conversation system according to a third embodiment of the present invention will be described with reference to Fig. 3A. The character conversation system of the third embodiment is provided with a large-scale language model server 20001 instead of the large-scale language model server 19001 in Fig. 2A, and is connected to the Internet 19000.
[0133] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence, and is a multimodal large-scale language model artificial intelligence that can process not only the natural language text information that the large-scale language model server 19001 could process, but also types of information other than natural language text information.
[0134] Moreover, the AI response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (AI response output device 10010) of the second embodiment, as an example.
[0135] In the third embodiment as well, the AI response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the Internet 19000 using an API.
[0136] The character conversation system of the third embodiment includes a mobile information processing terminal 20010 used by a user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.
[0137] 3B, an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011, which is a touch operation input panel, a control unit 20012, an external power supply input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.
[0138] The display panel 20011 is equipped with a touch operation input sensor and can accept touch operation input by the finger of the user 230. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel. The display panel 20011 may also be referred to as a display unit.
[0139] The communication unit 20020 may be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and a memory, and the control unit controls the display panel 20011 and the communication unit 20020. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000. This allows the mobile information processing terminal 20010 to communicate with various servers connected to the Internet 19000.
[0140] The power supply 20014 converts AC current input from the outside via the external power supply input interface 20013 into DC current and supplies the required DC current to each unit of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Furthermore, the secondary battery 20015 supplies power to each unit requiring power via the external power supply input interface 20013 when power is not supplied from the outside.
[0141] The video signal input unit 20023 is connected to an external video output device and inputs video data. The video signal input unit 20023 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be various USB interfaces.
[0142] The audio signal input unit 20024 is connected to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 20024 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface in which a terminal and a cable are integrated.
[0143] The audio output unit 20021 can output audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 can also output audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured with a speaker. The audio output unit 20021 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard.
[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 20012, which will be described later, may perform voice recognition processing on the generated audio signal to obtain text information from the audio signal.
[0145] The imaging unit 20025 is a camera having an image sensor. A camera may be provided on the front side of the mobile information processing terminal 20010, on the display panel 20011 side, or on the back side of the display panel 20011 side. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.
[0146] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 20016 before product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices, external servers, etc., via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may also be output to external devices, external servers, etc., via the communication unit 20020.
[0147] The video control unit 20017 performs various controls related to the video signal input to the display panel 20011. The video control unit 20017 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be referred to as a video processing unit or image processing unit. The video control unit 20017 controls video switching, such as which video signal to input to the display panel 20011 between the video signal to be stored in the memory 20026 and the video signal (video data) input to the video signal input unit 20023. The video control unit 20017 may also control image processing of the video signal input from the video signal input unit 20023 and the video signal to be stored in the memory 20026. Examples of image processing include scaling, which enlarges, reduces, or transforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0148] The attitude sensor 20018 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each unit connected thereto.
[0149] The nonvolatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the nonvolatile memory 20027 includes, for example, data for various operations to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 20026 stores video data to be displayed on the display panel 20011, data for controlling the device, etc. The control unit 20012 may read various software from the storage unit 20016 and expand and store it in the memory 20026.
[0150] The control unit 20012 controls the operation of each unit connected to it. The control unit 20012 may also work in conjunction with a program stored in the memory 20026 to perform arithmetic processing based on information acquired from each unit within the mobile information processing terminal 20010.
[0151] Next, an example of the operation of the character conversation device (AI response output device 10010) according to the third embodiment of the present invention will be described with reference to Figure 3C. This can also be said to be an example of the operation of a character conversation system including the AI response output device 10010 and the large-scale language model server 20001. In the third embodiment as well, the character conversation device (AI response output device 10010) loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0152] In the second embodiment, the actions performed by the user 230 with respect to the character conversation device (artificial intelligence response output device 10010) were mainly calls made by the user 230 using his / her voice. The character conversation device (artificial intelligence response output device 10010) of the second embodiment performed a series of operations starting from the process of collecting the voice of the user 230 with a microphone. In contrast, the character conversation device (artificial intelligence response output device 10010) of the third embodiment is also capable of performing the series of operations performed by the character conversation device (artificial intelligence response output device 10010) starting from the process of collecting the voice of the user 230 with a microphone, as described in the second embodiment. In addition, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the user 230 can perform an action with respect to the character conversation device (artificial intelligence response output device 10010) by user operation via the operation input unit 1107 of FIG. 1B. Here, examples of the operation input unit 1107 of FIG. 1B include a mouse, a keyboard, and a touch panel.
[0153] Furthermore, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by performing a touch operation by the user that can be detected by the touch operation input sensor of the display unit 10011 of Figure 1B.
[0154] In addition, the user 230 can operate the mobile information processing terminal 20010 and communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting the user's 230 operation input to the character conversation device (artificial intelligence response output device 10010).
[0155] Alternatively, an information storage image such as a two-dimensional code storing information that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as a two-dimensional code captured by the imaging unit 1180, and obtain the information. Alternatively, an image that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180, and obtain the results of the image recognition processing.
[0156] As described above, the character conversation device (artificial intelligence response output device 10010) of the third embodiment has a greater variety of actions that the user 230 can take toward the character conversation device (artificial intelligence response output device 10010) than the character conversation device (artificial intelligence response output device 10010) described in the second embodiment. As a result, the character conversation device (artificial intelligence response output device 10010) of the third embodiment can acquire the results of actions taken by the user 230 other than the user's voice, and generate an instruction sentence (prompt) to be sent to the large-scale language model server 20001 based on the results. As a result, the instruction sentence to be sent to the large-scale language model server 20001 can more preferably include information of a type other than text information in a natural language extracted from the user's voice. Examples of information of a type other than text information in a natural language extracted from the user's voice include images, videos, and sounds.
[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment uses an API to send an instruction to the large-scale language model server 20001. In this embodiment, the instruction may be metadata containing information written in a notation using tags, such as the markup format of a markup language, a notation using predetermined symbols, such as Markdown, or an object notation of a predetermined script, such as JSON. In this embodiment, the instruction may be classified into two types: a setting instruction that stores instructions, such as initial settings, and a user instruction that reflects instructions from a user. Type identification information identifying whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. In this case, the instruction includes text information in a natural language as the main message. Furthermore, in this embodiment, the main message of the instruction may include, in addition to the text information in natural language, a non-natural language information source, such as an image, video, or audio, as information of a type other than the natural language text information. A specific method for including a non-natural language information source in the instruction will be described later.
[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. The large-scale language model server 20001 receives an instruction from a character conversation device (artificial intelligence response output device 10010). Based on the instruction, the multimodal large-scale language model performs inference and generates a response including natural language text information that is the result of the inference. Here, because the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.
[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or sounds stored as the main message in the response. The character operation program of the character conversation device (artificial intelligence response output device 10010) may use a voice synthesis technology to generate a natural language voice as a reply to the user based on the natural language text information extracted from the response, and output the voice from the voice output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.
[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display characters in a natural language that serve as a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on text information in a natural language extracted from the above-mentioned response. At this time, the characters may be displayed together with the character 19051, may be displayed superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the information of the image of the non-natural language information source extracted from the above-mentioned response. At this time, the image may be displayed together with the character 19051, may be superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the video information of the non-natural language information source extracted from the above-mentioned response. At this time, the video may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed in place of the video of the character 19051. These specific processes may be executed by the video control unit 1160.
[0163] In addition, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output a voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140, which is a speaker.
[0164] As described above, the character conversation device (artificial intelligence response output device 10010) of FIG. 3C or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 does not require the large-scale language model itself, which requires vast amounts of data and computational resources for learning, to be installed in the character conversation device (artificial intelligence response output device 10010). Furthermore, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via an API. In response to a user's action toward a character, a response based on a non-natural language information source can be provided in addition to a response based on natural language text, enabling a more appropriate conversation to be held.
[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3D . This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3D shows examples of non-natural language information sources, such as natural language text and images, of the main message of an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001, and examples of non-natural language information sources, such as natural language text and images, of the main message of a server response that is the response. In this embodiment, the non-natural language information source can be an image, video, audio, or the like, but FIG. 3D shows an example of an image as the non-natural language information source.
[0166] 3D also shows an exchange of instructions and responses in chronological order, from a setting instruction, a first round of user instructions and their responses, to a second round of user instructions and their responses. The instructions and responses shown in FIG. 3D include non-natural language information source 20061 and non-natural language information source 20062, which were not shown in FIG. 2D of Example 2. In the example of FIG. 3D, both non-natural language information source 20061 and non-natural language information source 20062 are images.
[0167] Here, for ease of explanation, FIG. 3D shows an image of the non-natural language information source 20061 pasted into the instruction. However, there are multiple methods for transmitting or specifying data from the non-natural language information source 20061 in an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these multiple methods or switch between them. An example of each method will be described below.
[0168] The first method for transmitting or specifying non-natural language information source data in a directive is used, for example, when the non-natural language information source to be specified is a non-natural language information source located in a location such as a server connected to a network such as the Internet. A specific example of the first method is to use information such as tags and symbols in the directive to specify a non-natural language information source file located on a network such as the Internet by using location information (such as a URL) on the network such as the Internet and a file name.
[0169] For example, it is a tag that specifies an image in a markup language. <img src=""****”"> You can also specify an image on a network such as the Internet by writing the location information and file name information of the image file in the **** part using the tag. <video src=""****”">You can also specify a video file on a network such as the Internet by writing the location information and file name information of the video file in the **** part using the tag. <audio src=""****”">By using the above, audio that exists on a network such as the Internet can be specified by entering the location information and file name information of the audio file in the **** part. Also, if the notation is JSON, an image that exists on a network such as the Internet can be specified by preparing a key such as img_src and entering the location information and file name information of the image file as the value. For video files and audio files, it is sufficient to prepare the respective keys and values. This specific example of a format is just one example, and other unique formats can also be used. In either case, information that specifies the location information and file name information of the non-natural language information source file can be stored in the directive.
[0170] When information specifying the location information and file name information of a non-natural language information source file is stored in a directive, as in the first method, the directive itself does not need to store the data of the non-natural language information source file itself. Therefore, the amount of data in the directive can be reduced. In the first method, the large-scale language model server 20001 that receives a directive specifying non-natural language information source data simply uses the location information and file name information of the non-natural language information source file stored in the directive to obtain the non-natural language information source file located in a location such as a server connected to a network such as the Internet.
[0171] Here, how location information and file name information are input when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence using the first method will be described. In FIG. 3C, it has been explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are increased beyond the voice of the user 230 compared to the second embodiment. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, and the like, by user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in FIG. 1B.
[0172] Furthermore, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a web browser program and display a GUI of the web browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the web browser program may be accepted by a user operation via the operation input unit 1107 (e.g., a mouse, keyboard, or touch panel) or a user touch operation detectable by a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, video, or audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction. In this case, the web browser program may acquire location information and file name information of the non-natural language information source data and pass them to the character operation program.
[0173] Furthermore, user 230 may operate mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from mobile information processing terminal 20010, thereby inputting location information such as a URL for specifying non-natural language information source data to the character conversation device (artificial intelligence response output device 10010). Alternatively, location information such as a URL for specifying non-natural language information source data, file name information, and the like may be input in a manner such as displaying an information storage image such as a two-dimensional code on display panel 20011 of mobile information processing terminal 20010, performing image recognition processing on an image captured by imaging unit 1180 of character conversation device (artificial intelligence response output device 10010), and acquiring the results of the image recognition processing, as described in FIG.
[0174] Note that the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file is previously present in a location such as a server connected to a network such as the Internet. For example, if non-natural language information source data such as images, videos, or audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) is to be included in the instruction statement, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to the second server 19002 via the Internet 19000, and include in the instruction statement the location information on the Internet (such as a so-called URL) and file name of the non-natural language information source data on the uploaded second server 19002. In this case, the second server 19002 functions as a so-called intermediate server.
[0175] Similarly, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in an instruction statement, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the location information on the Internet (such as a URL) and the file name of the non-natural language information source data on the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include the acquired location information on the Internet (such as a URL) and the file name of the non-natural language information source data uploaded to the second server 19002 in the instruction statement.
[0176] Furthermore, a media server may be constructed within the character conversation device (artificial intelligence response output device 10010) so that the character operation program of the character conversation device (artificial intelligence response output device 10010) can cooperate with memory 1109 and storage unit 1170 to be accessible from other servers via the Internet 19000. In this case, when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction statement using the first method, the character conversation device (artificial intelligence response output device 10010) may store location information on the Internet (such as a URL) indicating the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data in the instruction statement.
[0177] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction sentence is, for example, a method in which the non-natural language information source data itself is simply stored (attached) to the instruction sentence (prompt) and transmitted. Generally, non-natural language information source data such as images, videos, and audio has a larger data volume than text information in natural language. Therefore, in this case, the data volume of the instruction sentence (prompt) itself is larger than that in the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) temporarily stores the non-natural language information source data to be stored (attached) to the instruction sentence (prompt) in the memory 1109, and when transmitting the instruction sentence (prompt), stores (attaches) the non-natural language information source data in the memory 1109 via the communication unit 1132 and outputs the data to the large-scale language model server 20001. The non-natural language information source data itself that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read from the storage unit 1170 and stored in memory 1109.
[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data using instruction sentences.
[0179] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. Therefore, in the first pass of the user instruction shown in the example of Figure 3D, it can obtain images of the swimming pool and poolside, which are non-natural language information source 20061, and text information in natural language, and output the text information in natural language as a response to the first pass of the user instruction as an inference result, as shown in the figure.
[0180] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example of the response to the second round of user instructions in Figure 3D, the large-scale language model server 20001 can include in the response a non-natural language information source 20062 generated by inference from the multimodal large-scale language model and send it to the character conversation device (artificial intelligence response output device 10010). In Figure 3D, the non-natural language information source 20062 is an example of an image in which a circle image is added to the image of the swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 3D, but may be video or audio.
[0181] When the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, a method similar to the first method or the second method in which the character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence can be used.
[0182] Specifically, as a method similar to the first method described above, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The non-natural language information source 20062, such as an image, video, or audio, itself may be stored in the large-scale language model server 20001, or the non-natural language information source 20062 may be transferred to and stored in a second server 19002 functioning as an intermediate server. In either case, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The character conversation device (artificial intelligence response output device 10010) that receives the response may access the large-scale language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction to acquire the non-natural language information source 20062.
[0183] Specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction sentence and use it for various outputs to the user 230.
[0184] According to the operation of the character conversation device (artificial intelligence response output device 10010) and the character conversation system of the third embodiment described above with reference to Fig. 3D, instructions and responses are sent and received between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230 to realize a conversation using non-natural language information such as images, videos, and sounds. This makes it possible to realize a more advanced and natural conversation as shown in each message in Fig. 3D.
[0185] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3E shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0186] 3E shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in FIG. 3D has ended. In the example of FIG. 3E, processing using the conversation history as described in FIGS. 2F, 2G, and 2I of Example 2 is not performed. Therefore, like FIG. 2E of Example 2, FIG. 3E shows a response with content that does not remember at all the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, the user's name, the conversation history, and the like that were included in the setting directive.
[0187] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3F. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3F shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.
[0188] Figure 3F shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in Figure 3D has ended. Here, Figure 3F shows an example in which the method of storing a message explaining the history of past conversations in the setting instruction sentence, which was explained in Figure 2F of the second embodiment, is also applied to the character conversation device (artificial intelligence response output device 10010) of the third embodiment. Specifically, in Figure 3F, the message that is the content of the setting instruction sentence in Figure 3D is stored as a reset message, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.
[0189] The large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. Therefore, non-natural language information source data may have been transmitted or specified in past instructions and responses. Therefore, in the example of Figure 3F, the conversation history message reflects not only the natural language text information in past instructions and responses, but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method for transmitting or specifying non-natural language information source data in the instruction in Figure 3F is similar to the transmission or specification of non-natural language information source data described in Figure 3D, and therefore a repeated description will be omitted.
[0190] In the example of Figure 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself in the instruction statement, and storing (attaching) the non-natural language information source data in the instruction statement without storing (attaching) the non-natural language information source data in the instruction statement. This also applies to the instruction statement of Figure 3F.
[0191] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3G shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0192] Figure 3G shows an example of a series of conversations shown in Figure 3F, from the first user instruction and its response following the first setting instruction to the third user instruction and its response. Figure 3G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 3F, so repeated description is omitted.
[0193] As described above, even when using the large-scale language model server 20001 having a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information according to the third embodiment, even when the user 230 speaks to the character 19051 again to start a new conversation after a series of conversations has ended, if the generation and transmission processes of the setting instruction sentence in Fig. 3F are performed, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation, as shown in Fig. 3G. This is preferable because it allows the user to recognize that the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation are more closely matched.
[0194] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, FIG. 3H uses the example described in FIG. 2H of the second embodiment for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.
[0195] Furthermore, database 20200 for managing character settings and character conversation history, shown in Figure 3H, has the same format as database 19200 shown in Figure 2I of Example 2, and in Figure 3H, only the differences from database 19200 shown in Figure 2I will be explained. Also, the content of the character "Koto" in the database will be explained, and the content of other characters will be omitted.
[0196] As described above, the large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information. Therefore, both the instruction statements from the character conversation device (artificial intelligence response output device 10010) and the responses from the large-scale language model server 20001 include not only natural language text information but also the transmission or specification of non-natural language information source data. Therefore, the database 20200 shown in FIG. 3H records, in the conversation history data, not only the natural language text information included in these instruction statements and responses, but also the transmission or specification of non-natural language information source data. The specific method for transmitting or specifying non-natural language information source data in the conversation history recording is the same as the transmission or specification of non-natural language information source data described in FIG. 3D, and therefore a repeated description will be omitted.
[0197] In the example of FIG. 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself to the instruction statement and not storing (attaching) the non-natural language information source data to the instruction statement. This is also true for the conversation history of FIG. 3H. However, in the conversation history of FIG. 3H, if the method of specifying non-natural language information source data is to specify location information and file name information of a non-natural language information source file on a server on a network such as the Internet (such as the second server 19002 functioning as an intermediate server or another cloud server), there is a possibility that the non-natural language information source file on the server may be deleted over a long period of time in the conversation history. In this case, even if the location information and file name information are used, the non-natural language information source file cannot be retrieved at a later date, and information in the conversation record may be lost.
[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts the instruction sentence and the response message into a conversation history and records them, it can use the location information and file name information to obtain the non-natural language information source file itself specified in the instruction sentence and the response from a server on the network and store it in the storage unit 1170. Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) can rewrite the location information and file name of the non-natural language information source file to Internet location information (such as a URL) indicating the location of the non-natural language information source file within the media server built within the character conversation device (artificial intelligence response output device 10010), and then record the non-natural language information source in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself deletes the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the conversation record information, which is more suitable for preserving the conversation record.
[0199] By using the database of Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between multiple character candidates to display on the display unit 10011, the user will feel less uncomfortable when conversing with each character, and will be able to share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.This effect can be achieved as shown in Figure 2I of Example 2, and can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information.
[0200] In addition, in the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, the large-scale language model server 20001 uses a multimodal large-scale language model artificial intelligence that can process not only natural language text information but also non-natural language information other than natural language text information.
[0201] Here, an API is used for communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001. In a multimodal large-scale language model, the API usage fee may be charged according to the number of processed natural language text information, which are units of words that separate sentences called tokens, as well as the amount of data from non-natural language information sources.
[0202] Therefore, in order to provide users with a character conversation service by the character conversation system according to this embodiment at a lower cost, the following modified example may be used.
[0203] In a first variation, the database conversation history record of FIG. 3H also records the transmission or specification of non-natural language information source data. However, the character and the user exchange conversations about the natural language information source data using text information in natural language, and the content of those conversations is recorded as text information in natural language. Even if the transmission or specification of the natural language information source data is omitted from the database conversation history record of FIG. 3H, the conversation itself about the natural language information source data will still be recorded as text information in natural language to some extent. Therefore, if a certain amount of information reduction is acceptable, the transmission or specification of the natural language information source data may be omitted from the database conversation history record of FIG. 3H. In this case, the transmission or specification of the natural language information source data is also omitted from the conversation history message of the setting instruction sentence of FIG. 3F. This reduces the amount of data from non-natural language information sources communicated using the API.
[0204] Next, as a second modification, in the recording of the conversation history of the database in FIG. 3H, natural language text information explaining the content of the non-natural language information source data is recorded instead of transmitting the non-natural language information source data or recording specified information. The natural language text information explaining the content of the non-natural language information source data may be obtained, for example, by starting a conversation between the large-scale language model of the large-scale language model server 20001 and a character conversation device (artificial intelligence response output device 10010) separately from the conversation as a character, and having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the natural language text information may be obtained by having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit through a conversation with another large-scale language model of another server that is available at a lower cost than the large-scale language model of the large-scale language model server 20001. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be used as natural language text information explaining the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is a tag in a markup language. <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">This is the text information written in the **** part of the above.
[0205] Furthermore, in the case of JSON notation, in an object that is stored in association with the location information and file name information of the non-natural language information source data, which are keys and values indicating the location information of the non-natural language information source data, a key corresponding to the alternative text can be further linked and stored as a value that is the alternative text data itself.
[0206] In this case, too, the transmission of the natural language information source data or the recording of the specified information can be omitted in the recording of the conversation history in the database of Fig. 3H, and the transmission of the natural language information source data or the specified information can also be omitted from the conversation history message of the setting instruction sentence of Fig. 3F, thereby reducing the amount of data from non-natural language information sources communicated using the API.
[0207] Next, a third variation is an example in which, at the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data is not stored in the user instructions, but is replaced with natural language text information explaining the content of the non-natural language information source data. For example, in the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data 20061 may be replaced with the user instructions, and a description such as "This image is of a swimming pool with seats and parasols by the pool. There is water in the swimming pool. There are drinks on the table next to the seats" may be stored as natural language text information. In this case, the description may be obtained by having another large-scale language model on another server, which is more inexpensive to use than the large-scale language model on the large-scale language model server 20001, explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the description may be obtained from a server of various other services that can obtain summaries and descriptions of non-natural language information source data such as images, videos, and audio. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be natural language text information that explains the content of the non-natural language information source data.
[0208] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3I. The example of FIG. 3I shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which is described in each of FIGS. 3A to 3H, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063 of natural language information source data, image 10064 of non-natural language information source data, and / or video 10065 of non-natural language information source data, which are responses from the large-scale language model, together with a video of a character 19051, on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed superimposed on top of the video of the character 19051, as shown in FIG. 3I.
[0209] Furthermore, text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on the video of character 19051. The display in FIG. 3I is an example, but for example, if user 230 adjusts the volume of the audio output of audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating operation input unit 1107 or the touch operation input sensor of display unit 10011, user 230 will not be able to hear the response from the large-scale language model by audio. Therefore, in this case, control unit 1110 may control to start a display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in FIG. 3I.
[0210] In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more preferably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the image of the character 19051, may be manually switched on / off by the user 230 operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. According to the display example of FIG. 3I, the character conversation device (artificial intelligence response output device 10010) that supports multimodal display can more preferably output a response from the large-scale language model.
[0211] According to the character conversation device and character conversation system of Example 3 described above, in addition to the effects of the character conversation device and character conversation system of Example 2, it is possible to provide users with a more advanced conversation experience that includes information in non-natural languages in addition to information in natural languages by using a multimodal large-scale language model. Furthermore, according to the character conversation device and character conversation system of Example 3, it is possible to provide users with a character conversation service at a lower cost.
[0212] In the above description of the third embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 20001 is used as the large-scale language model. In contrast to this, the character conversation device (artificial intelligence response output device 10010) may be equipped with the local LLM processing unit 10028 shown in FIG. 1B and may use the multimodal large-scale language model held by the local LLM processing unit 10028. In this case, the multimodal large-scale language model held by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model held by the large-scale language model server 20001.
[0213] In this case, in the above description of the third embodiment, the multimodal large-scale language model held by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model held by the local LLM processing unit 10028 of the character conversation device (the AI response output device 10010). In this case, too, a more sophisticated conversation experience that includes non-natural language information in addition to natural language information can be provided to the user using the multimodal large-scale language model. Note that, when the multimodal large-scale language model held by the local LLM processing unit 10028 is used instead of the multimodal large-scale language model held by the large-scale language model server 20001, there is less need to consider usage fees based on the number of processed tokens and the amount of data in the non-natural language information source. However, even with the multimodal large-scale language model held by the local LLM processing unit 10028, the number of processed tokens and the amount of data in the non-natural language information source can be reduced, thereby reducing the consumption of resources such as power required for inference. In this case, a character conversation service that consumes less power can be provided to the user.
[0214] The configuration described in Example 2 for uploading and downloading the conversation history with a character or the database data including the conversation history with a character to the second server 19002 or another cloud server can also be used in the example using a multimodal large-scale language model described in Example 3. In this case, too, when a user has multiple conversations with each of the multiple characters at different times between different individual devices for one character or multiple characters, it is possible to realize a conversation in which the memory of each character is pseudo-carried over from the previous conversation, which is more convenient for the user.
[0215] Example 4 Next, Example 4 of the present invention is an improvement of the AI response output device 10010, the character conversation device, or these systems described in the drawings of Example 2 or Example 3. In this Example, differences from Example 2 or Example 3 will be described, and repeated explanations of configurations similar to those examples will be omitted.
[0216] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0217] An example of operation using a database in a character conversation device (artificial intelligence response output device 10010) according to a fourth embodiment of the present invention will be described with reference to Fig. 4A. The database according to the fourth embodiment shown in Fig. 4A is an extension of the database described with reference to Fig. 2I or Fig. 3I. Specifically, the database shown in Fig. 4A assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores in the database initial setting instructions and conversation histories corresponding to each user and character.
[0218] In the example of Fig. 4A, for user 1 with a user ID of 1, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored.
[0219] These initial setting instruction sentences and conversation history data are stored as separate data in different areas for each combination of user and character. For the sake of explanation, in Fig. 4A, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (AI response output device 10010) uses the initial setting instruction sentences and conversation history stored in different areas for each combination of user and character based on the user currently using (logging in to) the character conversation device (AI response output device 10010) or its system, thereby making it possible to more appropriately maintain the consistency of the character's personality and the continuity of memory for each different user.
[0220] Specifically, consider a situation in which User 1 has previously had a conversation with character Tom using a character conversation device (AI response output device 10010), and User 2 is unaware of that conversation, and then User 2 subsequently has a conversation with character Tom. In this case, if AI response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify users, the response output from AI response output device 10010 will be based on a conversation history that User 2 does not remember, and the conversation between User 2 and the character of AI response output device 10010 may become inconsistent.
[0221] In contrast, even in a similar situation, if the database shown in Figure 4A is used, the control unit 1110 of the character conversation device (AI response output device 10010) identifies users by ID, stores the initial setting instructions and conversation history in a different area for each user, and uses the initial setting instructions and conversation history stored in a different area for each user to generate an AI response. As a result, the initial setting instructions and conversation history used to generate an AI response for each user are based on the operation or conversation history of that user and are managed separately from the operation or conversation history of other users. This makes it possible to more appropriately maintain consistency in the conversation history between each user and each character of the AI response output device 10010.
[0222] The database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in the storage unit 1170 of the AI response output device 10010 and used by the control unit 1110. Furthermore, without being limited to this, the database of initial setting directives and / or conversation histories may be stored in a server on the network. For example, if the AI response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 in generating an AI response, the database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in these servers themselves. In this way, it is possible to omit the process of transmitting the initial setting directives and conversation histories again from the AI response output device 10010 to these servers by including them in directives, thereby reducing the number of transmission tokens for use of the large-scale language model.
[0223] When storing the database of the initial setting instruction sentence and / or the conversation history described in FIG. 4A, the AI response output device 10010 can transmit the user ID, the character ID, and the user instruction sentence for the subsequent conversation to these servers. The large-scale language models on these servers can use the user ID and character ID obtained from the AI response output device 10010 to obtain the corresponding initial setting instruction sentence and the conversation history from the database of the initial setting instruction sentence and / or the conversation history of FIG. 4A. The large-scale language models on these servers can perform inference using the initial setting instruction sentence and the conversation history, and the user instruction sentence for the subsequent conversation transmitted from the AI response output device 10010, to generate an AI response and transmit it to the AI response output device 10010. In this way, the effect of more appropriately maintaining the consistency of the character's personality and the continuity of memory for each different user can be obtained while saving the number of transmitted tokens when using the large-scale language model.
[0224] Next, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) according to the fourth embodiment of the present invention will be described with reference to Fig. 4B. The database according to the fourth embodiment shown in Fig. 4B is an extension of the database described in Fig. 1C or Fig. 2L. Specifically, the database shown in Fig. 4B assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data on standard response phrases corresponding to each user and character in the database.
[0225] In the example of Fig. 4B, for user 1, whose user ID is 1, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3. In addition, for user 2, whose user ID is 2, and user 3, whose user ID is 3, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3.
[0226] These standard response phrase data are stored as separate data in different areas for each user-character combination. For ease of explanation, in FIG. 4B, the data stored in each area is represented as standard response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, standard response phrase data 101 is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 1: Koto shown in FIG. 2L. Data 201 in FIG. 4B is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 2: Tom shown in FIG. 2L.
[0227] Data 301 in FIG. 4B is stored as a database such as a table corresponding to the fixed response phrases corresponding to condition numbers 1 to 7 of character 3: Necco shown in FIG. 2L. Data 102, 202, and 302 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 2. Data 103, 203, and 303 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 3. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses fixed response phrase data stored in different areas for each combination of user and character, based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.
[0228] In this way, even the same character can respond with different template responses for each user. That is, even for the same character, it may be more appropriate to vary the content of the template responses depending on the relationship between the character and the user. For example, depending on the relationship between the character's set age and the age of the user registered in the AI response output device 10010 or the system, the user may be older, the same age, or younger than the character. In this case, different content for the template responses of the character to older users, the template responses of the character to users of the same age, and the template responses of the character to younger users can make the conversation between the user and the character more appropriate or natural. That is, by performing operations using the database of FIG. 4B and varying the content of the template responses depending on the relationship between the character and the user, it is possible to create a more appropriate or natural conversation.
[0229] 4B described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the response template database (response template DB) shown in FIG. 4B may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the AI response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0230] According to the character conversation device and character conversation system of Example 4 described above, it is possible to produce more suitable or more natural conversation depending on the relationship between the character and the user, the conversation history, etc.
[0231] <Example 5> Next, a fifth embodiment of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of the first, second, and third embodiments. Specifically, this is an example in which the response generation process of the AI response output device 10010 is switched from response generation process using a large-scale language model on a network to response generation process using a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, or response generation process using a fixed response phrase database. In this embodiment, differences from these embodiments will be described, and repeated description of configurations similar to those of these embodiments will be omitted.
[0232] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0233] An example of response generation process switching processing in the AI response output device 10010 of the fifth embodiment of the present invention will be described using FIG. 5A. The table in FIG. 5A shows examples 1 to 9 of response generation process switching processing in the AI response output device 10010. In the table in FIG. 5A, the column "Switching Overview" shows an overview of the switching processing for each example. The column "State before switching of LLM on the network (API connected LLM)" shows the state before response generation processing by a large-scale language model on the network (large-scale language model connected using an API), such as the large-scale language model provided in the large-scale language model server 19001 in FIG. 1 and the multimodal large-scale language model provided in the large-scale language model server 20001, is switched to another response generation process. The column "Switching Occurrence Condition" shows the conditions under which switching processing of the response generation process occurs. The column "Switching destination from LLM on the network (API-connected LLM)" indicates the switching destination to which the response generation process of the AI response output device 10010 is switched from a large-scale language model on the network (a large-scale language model connected using an API), such as a large-scale language model provided in the large-scale language model server 19001 and a multimodal large-scale language model provided in the large-scale language model server 20001. When the condition indicated in "Switching occurrence condition" occurs in the state of "State before switching of LLM on the network (API-connected LLM)" shown in Figure 5A, the control unit 1110 of the AI response output device 10010 may perform control to switch to the large-scale language model, database, or correspondence indicated in "Switching destination from LLM on the network (API-connected LLM)."
[0234] Below, each example shown in the table of FIG. 5A will be described. Example 1, as shown in the "Switching Overview," is an example in which switching is performed depending on the network connection status of the AI response output device 10010. In Example 1, the "state before switching of the LLM (API-connected LLM) on the network" indicates that the network connection status of the AI response output device 10010 is connectable. Here, in Example 1, the "switching occurrence condition" indicates "when network connection becomes unavailable." That is, this means when the connection via the network between the AI response output device 10010 and a large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable. Specifically, the connection failure may be due to a communication failure state on the connection path from the AI response output device 10010 to the Internet 19000. Alternatively, the connection failure may be due to a communication failure state on the Internet 19000. Alternatively, the connection failure may be due to a situation in which the large-scale language model on the network (a large-scale language model connected using an API) itself cannot connect to the Internet 19000. Also, in Example 1, "local LLM" is shown as the "switching destination from LLM on the network (API-connected LLM)." This specifically means that switching processing is performed to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. That is, in Example 1, even if connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and response generation processing by the large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable, response generation processing is switched to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. This makes it possible to continue response generation processing using the large-scale language model, despite differences in performance as a large-scale language model.
[0235] Next, Example 2 in FIG. 5A will be described. In Example 2, the "switching destination from the networked LLM (API-connected LLM)" in Example 1 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 2, if a connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and the response generation process using the large-scale language model on the network (a large-scale language model connected using an API) cannot be used, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response through simpler processing and output the response to the user.
[0236] Next, Example 3 of FIG. 5A will be described. In Example 3, the "switching destination from LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "no-response handling." The "no-response handling" means that even if a user input requesting a response from a large-scale language model is received from the user via the touch panel, microphone 1139, or operation input unit 1107, no response to this input is generated, or even if a user input requesting a response from a large-scale language model is received, no response to this is output. In other words, Example 3 makes it easier to handle a situation where connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason, making it impossible to use the response generation process by the large-scale language model on the network (a large-scale language model connected using an API).
[0237] Next, Example 4 of FIG. 5A will be described. As shown in the "Switching Overview," Example 4 is an example in which switching is performed due to a response delay of an LLM on the network. In Example 4, the "state before switching an LLM on the network (API-connected LLM)" indicates a state in which a response from an LLM on the network is obtained within a predetermined time. Here, in Example 4, the "switching occurrence condition" indicates a case in which a response from an LLM on the network is not obtained within the predetermined time and exceeds the predetermined time. Also, in Example 4, the "switching destination from an LLM on the network (API-connected LLM)" indicates a "local LLM." The "local LLM" of the switching destination is the same as in Example 1, so repeated explanation will be omitted. That is, in Example 4, even if the response from an LLM on the network (a large-scale language model connected using an API) exceeds the predetermined time for some reason and the response generation process by the LLM on the network (a large-scale language model connected using an API) cannot be used smoothly, the response generation process is switched to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This makes it possible to continue the response generation process using a large-scale language model, even if there are differences in performance as a large-scale language model.
[0238] Next, Example 5 in FIG. 5A will be described. Example 2 is an example in which the "switching destination from the networked LLM (API-connected LLM)" in Example 4 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 5, if for some reason a response from the networked LLM (large-scale language model connected using an API) exceeds a predetermined time and the response generation process using the networked LLM (large-scale language model connected using an API) cannot be used smoothly, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response using simpler processing and output the response to the user.
[0239] Next, Examples 6 to 9 of FIG. 5A will be described. As shown in "Switching Overview," Examples 6 to 9 are examples in which switching is performed when the upper limit of API usage or usage fee is reached. As described in Example 2, providers of large-scale language models often collect the costs used to train the large-scale language model from terminal users as API usage fees for the terminal. In such cases, with natural language models, API usage fees are often charged based on the number of processing of units of words that separate sentences, called tokens. Here, various methods of charging and limiting API usage fees are conceivable. One possible example is to define the upper limit of the amount of large-scale language model usage services that a user can receive under normal circumstances using the number of tokens processed.
[0240] In this case, the user may be able to receive services using large-scale language models at a specified API usage fee until the usage volume (or corresponding usage fee) is reached, and once the upper limit of the usage volume (or corresponding usage fee) is reached, certain restrictions may be imposed, such as the user being unable to receive services using large-scale language models in their normal state (performance or frequency).
[0241] Examples 6 to 9 in FIG. 5A are examples of response generation process switching control by the control unit 1110 of the AI response output device 10010 when such restrictions occur in a service using a large-scale language model. Specifically, in Example 6, the "state before switching the on-network LLM (API-connected LLM)" is a state in which the API usage volume and API usage fee have not reached a predetermined upper limit. This means that the usage volume of the on-network LLM (API-connected LLM) has not reached a predetermined upper limit. In this case, the user can use the on-network LLM (API-connected LLM) in a normal state.
[0242] Here, in Example 6, the "switching trigger condition" is when the API usage volume or API usage fee reaches a predetermined upper limit. This means when the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit. Also, in Example 6, the "switching destination from the network-based LLM (API-connected LLM)" is a second LLM on the network that is different from the LLM (which may be referred to as the first LLM) used in normal operation. An example of a second LLM on the network is an LLM with a lower fee than the first LLM used in normal operation. Since it is a lower-fee service, the performance of the second LLM may be lower than that of the first LLM. Even in this case, there is still a significant advantage if large-scale language models can be used inexpensively even after the usage volume / fee limit of the first LLM is reached.
[0243] Next, Example 7 in Figure 5A will be described. In Example 7, the "switching destination from the network-based LLM (API-connected LLM)" in Example 6 is changed from a second LLM on a network different from the LLM (which may be referred to as the first LLM) used in the normal state to a "local LLM." In Example 7, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit, it is possible to continue performing response generation processing using a large-scale language model by switching to response generation processing using a local LLM that is not subject to restrictions such as the usage volume of the network-based LLM, the API usage volume, or the API usage fee.
[0244] Next, Example 8 in FIG. 5A will be described. In Example 8, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is similar to the process described in FIG. 1C, FIG. 2L, or FIG. 4B, and therefore will not be described again. In Example 8, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the networked LLM (API-connected LLM) reaches a predetermined upper limit, the process switches to response generation using a preparatory response phrase database, which is not subject to restrictions such as the usage volume of the networked LLM, the API usage volume, or the API usage fee. This makes it possible to generate a response using simpler processing and output the response to the user.
[0245] Next, Example 9 in FIG. 5A will be described. In Example 9, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "no-response handling." The "no-response handling" refers to a handling in which a response to the user is not generated or a response to the user is not output. Example 9 makes it easier to handle a situation in which response generation processing using a large-scale language model on the network (a large-scale language model connected using an API) is unavailable because the API usage or API usage fee has reached a predetermined upper limit, i.e., the usage of the networked LLM (API-connected LLM) has reached a predetermined upper limit.
[0246] According to the switching control of the response generation process of the AI response output device 10010 shown in Examples 1 to 9 of Figure 5A as described above, even in situations where the response generation process using an LLM (large-scale language model connected using an API) on the network cannot be used as usual, more appropriate switching or response can be made according to each situation.
[0247] 5A may be performed by combining a plurality of examples. For example, the switching control of Examples 1 to 3 may be combined with any of the controls of Examples 4 to 9. Similarly, the control of Example 4 or Example 5 may be combined with any of the controls of Examples 1 to 3 or Examples 6 to 9. Similarly, the control of Examples 6 to 9 may be combined with any of the controls of Examples 1 to 5.
[0248] Next, an example of display of an AI assistant or a character when the AI response output device 10010 of the fifth embodiment is configured as an AI assistant device or a character conversation device will be described with reference to FIGS. 5B to 5D.
[0249] First, Figure 5B is an example of the display of an AI assistant or character on the AI response output device 10010 when performing the switching control of Example 3 in Figure 5A. In the example of Figure 5B, the display state of the AI assistant or character is changed depending on whether the network connection status of the AI response output device 10010 is network connectable or network unconnectable. The network connectable and network unconnectable states of the AI response output device 10010 are as explained in Figure 5A, so a repeated explanation will not be given.
[0250] In the example of FIG. 5B, the AI response output device 10010 (1) displays the AI assistant or character in a normal, awake state when a network connection is possible, but (2) displays the AI assistant or character in a "sleeping" state when a network connection is not possible. In the switching control of example 3 of FIG. 5A, when the AI response output device 10010 cannot connect to the network, it does not generate or output a response even if a command is input from the user. In this case, if the AI assistant or character displayed by the AI response output device 10010 is in a normal, awake state, the user will feel uncomfortable. However, if the AI assistant or character displayed by the AI response output device 10010 is displayed in a sleeping state, the user will understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user.
[0251] In the case of Figure 5B(2), it is desirable that the user understand that "the reason the AI assistant or character is not responding is because it is asleep" before making a user input requesting a response using a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the start timing of the state in Figure 5B(2) where the AI assistant or character is displayed in a "sleeping" state when network connection is not possible is immediately after the control unit 1110 of the AI response output device 10010 determines that network connection is not possible, before the user makes a user input requesting a response using a large-scale language model.
[0252] Next, as another display example, the display example of FIG. 5C will be described. The display example of FIG. 5C is an example in which the display state of the AI assistant or character is changed depending on the state of the "switching destination from the LLM on the network (API-connected LLM)" in the table during the switching control of FIG. 5A. Specifically, FIG. 5C shows (1) a display example of the AI assistant or character in a state in which the AI response output device 10010 can connect to a large-scale language model on the network (a large-scale language model connected using an API) and is able to use a response generation process using the large-scale language model on the network (referred to as the normal state in this figure); (2) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to a response generation process using an LLM or a template response database with lower performance than the large-scale language model on the network (a large-scale language model connected using an API); and (3) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A.
[0253] In the example of FIG. 5C , for example, (1) when the AI response output device 10010 is in a "normal state," the AI response output device 10010 displays the AI assistant or character in a state where there are no particular problems. Note that the "normal state" in FIG. 5C may be considered a state other than states (2) and (3). Also, for example, (2) when the AI response output device 10010 has switched to a response generation process using an LLM or a response template database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), the AI response output device 10010 displays the AI assistant or character in a "sleepy" state. Note that "displaying the AI assistant or character in a "sleepy" state" may also be expressed as "a display indicating that the AI assistant or character is feeling drowsy."
[0254] The response generation process (2) has lower performance than the response generation process using a large-scale language model on the network (a large-scale language model connected using an API) in the normal state (1). Therefore, by displaying the AI assistant or character in a "sleepy" state, it is possible to implicitly convey to the user that the response performance of the AI assistant or character is low. This makes it possible to further reduce the sense of discomfort felt by the user due to a low-performance response. Note that the switching conditions under which the AI response output device 10010 switches to response generation processing using an LLM or a fixed response phrase database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), are as described in FIG. 5A, and therefore a repeated explanation will be omitted.
[0255] 5C(2), it is desirable to implicitly inform the user that the response performance of the AI assistant or character is low before the user makes a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the start timing of the state in FIG. 5C(2) where the AI assistant or character is displayed in a "sleepy" state be immediately after the AI response output device 10010 switches to a response generation process using an LLM or a template response database, which has lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user makes a user input requesting a response from a large-scale language model.
[0256] Also, for example, in (3) the state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A, the AI response output device 10010 displays the AI assistant or character in a "sleeping" state. As also described in FIG. 5B, by displaying the AI assistant or character displayed by the AI response output device 10010 in a "sleeping" state, the user can understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user. Note that the conditions under which the AI response output device 10010 switches to the no-response mode described in FIG. 5A are the same as those described in Example 3 or Example 9 of FIG. 5A, and therefore a repeated explanation will be omitted. Note that in the case of FIG. 5C(3), it is desirable for the user to understand that "the AI assistant or character is not responding because it is sleeping" before making a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the timing for starting the state (3) in Figure 5C in which the AI assistant or character is displayed in a "sleeping" state be immediately after the AI response output device 10010 switches to the no-response response described in Figure 5A, before the user input requesting a response from a large-scale language model.
[0257] 5C, the AI response output device 10010 displays a display that implicitly reflects a change in the state of the AI assistant or character without directly providing the user with a technical explanation of the state of the AI response output device 10010 related to the response generation process. This can reduce the sense of discomfort felt by the user compared to when a technical explanation of the state of the AI response output device 10010 related to the response generation process is directly provided to the user. Furthermore, this can reduce the sense of discomfort felt by the user compared to when the display state of the AI assistant or character remains the same as its normal state despite a change in the state of the AI response output device 10010 related to the response generation process.
[0258] However, some users may wish to know a more precise explanation of the technical state of each state. Therefore, a display example for such users will be described with reference to FIG. 5D. Among the rows of the table shown in FIG. 5D, the rows for explaining the device state and display state are identical to those in FIG. 5C, and therefore, repeated explanations will be omitted. Furthermore, the display example of the AI assistant or character shown in the row for the display example of the AI assistant or character is almost identical to that in FIG. 5C, except that a question mark (?) is displayed in the display example. The question mark (?) is a mark that the user operates when requesting an explanation from the AI response output device 10010, and may also be referred to as a help mark.
[0259] In the example of FIG. 5D , when a user selects the question mark (?) through a user operation, such as via the touch panel of the operation input unit 1107 or the display unit 10011 in FIG. 1B , the display of the AI assistant or character on the AI response output device 10010 changes to the display example shown in the row of the display example after user operation. Specifically, regardless of whether the device state is (1), (2), or (3), an explanation of the technical state of each state is displayed. For example, in the example of FIG. 5D , if the device state is (1) normal, a display explaining that the device is in a normal state with no particular technical limitations can be displayed, such as "Normal state." Furthermore, if the device state is (2) using a low-performance LLM or a standard response phrase database, a display technically explaining the low-performance state can be displayed, such as "Low-performance mode." This display can also be considered a display explaining the reason why the AI assistant or character is displaying a "sleepy" state.
[0260] In this case, a more detailed technical explanation may be provided. Specifically, a message such as "Low-performance LLM usage mode" or "Canned response mode" may be displayed. If the device status is (3) unresponsive, a message such as "Network connection unavailable" may be displayed, providing a technical explanation of the reason for switching to unresponsive mode. If the reason for switching to unresponsive mode is that a response from an LLM (a large-scale language model connected via an API) on the network exceeds a specified time, a message such as "Response from LLM is delayed" may be displayed. If the reason for switching to unresponsive mode is that the network LLM usage, API usage, or API usage fee has reached its limit, a message such as "LLM usage limit reached," "API usage limit reached," or "API usage fee has reached a specified amount" may be displayed. These messages may be considered to explain the reason why the AI assistant or character is displayed in a "sleeping" state.
[0261] According to the display example of FIG. 5D described above, even if there are technical constraints in the response generation process in the AI response output device 10010, first, instead of providing a direct explanation to the user, the state of the device is implicitly indicated by a change in the display state of the AI assistant or character, thereby further reducing the sense of discomfort felt by the user. This display is more suitable for users who do not need technical explanations. Furthermore, by displaying an operation mark to explain the technical state, a display is provided to users who operate the mark that technically explains the state of the response generation process in the AI response output device 10010 (normal state or state with technical constraints). This makes it possible to provide a more suitable display for users who want to know the technical state accurately.
[0262] In the examples of Figures 5B, 5C, and 5D, a "sleeping" state is shown as an example of the display state of the AI assistant or character when the AI assistant or character is "unresponsive," but this is only an example and the embodiment is not limited to this. Instead of the "sleeping" state, another display state that implies a situation where the AI assistant or character is unable to respond, such as "taking a break," may be used. In the examples of Figures 5C and 5D, a "sleepy" state is shown as an example of the display state of the AI assistant or character when a low-performance LLM or a standard response phrase database is being used, but this is only an example and the embodiment is not limited to this. Alternatively, another display state that implies that the AI assistant or character has low response performance, such as "hungry," may be used.
[0263] According to the AI response output device and the AI response output system according to the fifth embodiment described above, it is possible to more appropriately switch the response generation process used by the AI response output device depending on the connection state between the large-scale language model on the network and the AI response output device, the response delay state from the large-scale language model on the network, the usage amount of the large-scale language model on the network, etc. Furthermore, when the AI response output device according to the fifth embodiment is configured as an AI assistant device or a character conversation device, it is possible to perform a display that is less strange to the user.
[0264] Example 6 Next, Example 6 of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of Examples 1 to 5. Specifically, this is an example in which the response generation process of the AI response output device 10010 is more suitably combined with a response generation process using a large-scale language model on a network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, and a response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated explanations of configurations similar to those examples will be omitted.
[0265] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0266] An example of a response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 shows an example of a flowchart of the response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention. Specifically, a time axis in which time progresses from top to bottom, a processing flow, and a response output example are shown. The response shown in the response output example may be output via display on the display unit 10011 of the AI response output device 10010 or audio output by the audio output unit 1140.
[0267] In the example of FIG. 6, first, at time t0, a user input is made via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010, requesting a response based on a large-scale language model, and the control unit 1110 of the AI response output device 10010 acquires the user input (step 600). Next, at time t1, the control unit 1110 starts preparations for response output using a fixed response phrase database stored in the storage unit 1170, and starts response output using the fixed response phrase database (step 601). In the example of FIG. 6, response output using the fixed response phrase database starts at time t2, and as shown in the figure, the fixed response is being output but has not yet been completed. "Good morning" in the figure indicates the output of part of the sentence that continues "Good morning..."
[0268] At time t3, before the response output using the fixed response phrase database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and transmits the generated instruction sentence to a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, thereby starting a request for a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 starts acquiring a response from the large-scale language model (step 603).
[0269] An example of a response output in which the response output using the fixed response phrase database is completed at time t5 is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is the ____ day of the month, isn't it?" is completed at time t5 using the fixed phrases stored in the fixed response phrase database and date information stored in memory. Here, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 has already started acquiring a response from the large-scale language model. Therefore, at time t6, which follows time t5 when the response output display using the fixed response phrase database is completed, the control unit 1110 starts outputting a response from the large-scale language model following the response output using the fixed response phrase database (step 604). Thereafter, at time t7, a response from the large-scale language model is output following the response output using the fixed response phrase database. When the response output from the large-scale language model is completed, the response output according to the processing flow shown in FIG. 6 is completed (step 605).
[0270] Next, the effect of the processing flow shown in Fig. 6 of the present invention will be described. Processing a large-scale language model requires a large amount of computational resources. Generally, even if inference, which requires fewer computational resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several to several tens of seconds from when the control unit starts to request a response from the large-scale language model until it can obtain a response from the large-scale language model. This period corresponds to the period from time t3 to time t4 shown in Fig. 6. Furthermore, from time t0, when a user input is made, until time t4, the control unit 1110 is unable to obtain a response output from the large-scale language model, and therefore is unable to output a response from the large-scale language model to the user.
[0271] 6, there is no start of preparation for response output using the fixed response phrase database and no start of response output using the fixed response phrase database, the user may continue to wait for a period of several seconds to more than ten seconds from time t0 when the user input is made to time t4 without receiving a response from the AI response output device 10010. For example, when the AI response output device 10010 is configured as an AI assistant device or a character conversation device, the waiting time may give the user a sense of discomfort.
[0272] In contrast, in the processing flow according to the sixth embodiment of the present invention shown in Fig. 6, the control unit 1110 starts the process of outputting a response using a fixed response phrase database, which requires fewer computational resources than the process of a large-scale language model, before starting to acquire a response from the large-scale language model. This prevents the user from having to wait from time t0 to time t4 without receiving a response from the AI response output device 10010. From the user's perspective, whether the response output is using a fixed response phrase database or a response output from a large-scale language model, it is the same as receiving a response from the AI response output device 10010.
[0273] 6, by providing step 601 before step 603, the response of the AI response output device 10010 to the user can be artificially accelerated. This can further reduce the sense of discomfort felt by the user due to long waiting times. Furthermore, by outputting a response from a large-scale language model following a response using the template response database in step 604, the user can perceive these outputs as if they were a series of more natural outputs.
[0274] According to the AI response output device and AI response output system of Example 6 described above, the waiting time for a response from the AI response output device can be shortened, and the sense of discomfort felt by the user can be further reduced.
[0275] Example 7 Example 7 of the present invention is an improvement of the AI response output device 10010 or AI response output system described in the drawings of Examples 1 to 6. Note that in Example 7, differences from Examples 1 to 6 will be described, and repeated description of the same configurations as those examples will be omitted.
[0276] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0277] Here, in the AI response output device 10010 or the AI response output system, when an instruction statement is generated based on a user input requesting a response from a large-scale language model input via the AI response output device 10010, and the instruction statement is sent to the large-scale language model to obtain a response, the instruction statement may contain inappropriate content. Also, the response from the large-scale language model may contain inappropriate content. In these cases, a technique can be considered that replaces the response content from the large-scale language model with predetermined content to filter out inappropriate responses from being output.
[0278] FIG. 7A shows an example of a replacement response when filtering to prevent inappropriate responses from being output. For example, FIG. 7A(1) is an example of a natural language sentence to be output as a replacement response when an instruction sent from the AI response output device 10010 to a large-scale language model includes a statement requesting a response with violent content, or when the response generated by the large-scale language model includes violent content. FIG. 7A(2) is an example of a natural language sentence to be output as a replacement response when an instruction sent from the AI response output device 10010 to a large-scale language model includes a statement requesting a response with sexual content, or when the response generated by the large-scale language model includes sexual content. FIG. 7A(3) is an example of a natural language sentence to be output as a replacement response when an instruction sent from the AI response output device 10010 to a large-scale language model includes a statement requesting a response with content that endangers public safety, or when the response generated by the large-scale language model includes content that endangers public safety. 7A(4) is an example of a natural language sentence to be output as a replacement response when an instruction sent from the AI response output device 10010 to the large-scale language model includes a statement requesting a response that violates the privacy of a specific individual, or when a response generated by the large-scale language model includes content that violates the privacy of a specific individual. In this way, the large-scale language model outputs a replacement sentence to the AI response output device 10010 as a response to filter out inappropriate responses so that they are not output, and the AI response output device 10010 can output the replacement sentence to the user as a response.
[0279] However, there may be cases where it is not desirable to output the replacement sentence as a response to the user as shown in FIG. 7A. For example, consider a case where the AI response output device 10010 or the AI response output system described in each of the figures in Examples 1 to 6 initially specifies the name of the large-scale language model itself, the role to be played, conversational characteristics, etc., to the large-scale language model, causing the response from the large-scale language model to act as if it were a character with a consistent personality. In this case, consider a case where the replacement sentence, as shown in FIG. 7A, which declares itself to be an AI (artificial intelligence) and explains the situation, is output as a response to the user as is. In this case, even if the user previously recognized the response from the AI response output device 10010 or the AI response output system as a response from a character with a consistent personality, the output of the replacement sentence as shown in FIG. 7A may lead the user to realize that the series of responses up to that point were merely responses from an AI (artificial intelligence). This situation would ruin the consistent personality of the character that the user perceived based on the responses after the initial settings. Therefore, in the AI response output device 10010 or AI response output system according to this embodiment, when a large-scale language model filters to prevent an inappropriate response from being output, operations and processing are performed to output a more appropriate response to the user.
[0280] Hereinafter, specific operations and processes in the AI response output device 10010 or AI response output system according to this embodiment will be described.
[0281] First, FIG. 7B shows an example of a flowchart for performing processing to prevent an inappropriate response from being output in the AI response output device 10010 or AI response output system according to the seventh embodiment.
[0282] First, in step 701, when a user input is received via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010, the control unit 1110 generates instruction statement 1 such as a question or request based on the user input, and transmits the instruction statement 1 to a large-scale language model on a network such as the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001, or to a local large-scale language model such as the local LLM processing unit 10028 (S701). Next, in step 711, the large-scale language model that has received the instruction statement 1 performs an inappropriate filtering necessity determination process to determine whether inappropriate filtering is necessary (S711). The inappropriate filtering necessity determination process will be described in detail below.
[0283] Here, if the large-scale language model determines that inappropriate filtering is necessary through the inappropriateness filtering necessity determination process of step 711, it transmits response A containing natural language from which inappropriate content has been filtered to the control unit 1110 of the AI response output device 10010. An example of response A containing natural language from which inappropriate content has been filtered is the replacement sentence shown in FIG. 7A described above. At this time, the large-scale language model according to this embodiment transmits an inappropriateness filter flag B, indicating that the content of the response is inappropriate and has been filtered, to the control unit 1110 of the AI response output device 10010, in addition to or instead of response A containing natural language from which inappropriate content has been filtered. In other words, the inappropriateness filter flag B is a flag indicating that the response has been filtered. Note that the term "flag" in the embodiments of the present invention may be replaced with "control information." Here, in step 702, the control unit 1110 of the artificial intelligence response output device 10010, which has received response A containing natural language from which inappropriate content has been filtered and / or inappropriate filter flag B, performs processing to respond to response A based on response A containing natural language from which inappropriate content has been filtered and / or inappropriate filter flag B (S702).
[0284] On the other hand, if the large-scale language model determines in the inappropriateness filter necessity determination process of step 711 that inappropriate filtering is unnecessary, i.e., that the response is appropriate, the large-scale language model transmits the generated normal response C containing natural language to the control unit 1110 of the AI response output device 10010 without filtering for inappropriate content. At this time, the large-scale language model according to this embodiment may transmit an inappropriate filter flag D indicating that the content is not inappropriate to the control unit 1110 of the AI response output device 10010, in addition to the normal response C containing natural language generated by the large-scale language model. Here, in step 703, the control unit 1110 of the AI response output device 10010, which has received the normal response C containing natural language generated by the large-scale language model and / or the inappropriate filter flag D indicating that the content is not inappropriate, performs processing to address the response C based on the normal response C containing natural language generated by the large-scale language model and / or the inappropriate filter flag D indicating that the content is not inappropriate (S703).
[0285] According to the processing of the artificial intelligence response output device 10010 or the artificial intelligence response output system according to the flowchart of Figure 7B described above, in addition to or instead of the response from the large-scale language model, a flag indicating that the content is inappropriate and that the response has been filtered, or a flag indicating that the content of the response is not inappropriate, can be sent from the large-scale language model to the artificial intelligence response output device 10010.
[0286] Next, an example of the inappropriateness filter necessity determination process in step 711 of FIG. 7B will be described with reference to FIG. 7C. FIG. 7C shows a table related to the determination process. In the example of the determination process shown in FIG. 7C, two conditions, whether or not an inappropriate keyword is included in the instruction statement and whether or not inappropriate data is included in the generated data, are combined to determine whether the data is "inappropriate and needs to be filtered" or "not inappropriate and does not need to be filtered." In the example of FIG. 7C, if the instruction statement does not include an inappropriate keyword and the generated data does not include inappropriate data, the data is determined to be "not inappropriate and does not need to be filtered." Furthermore, if the instruction statement includes an inappropriate keyword or if the generated data includes inappropriate data, the data is determined to be "inappropriate and needs to be filtered" in both cases.
[0287] According to the determination process of FIG. 7C described above, it is possible to more appropriately determine the possibility of outputting an inappropriate response.
[0288] Next, a specific example of the inappropriate filter flag shown in Fig. 7B will be described with reference to Fig. 7D. Fig. 7D shows two examples of the inappropriate filter flag, example (1) and example (2).
[0289] First, for example, the example of the inappropriate filter flag in Figure 7D(1) simply indicates two states: a state where the response is judged to be "not inappropriate" and the response is "unfiltered," and a state where the response is judged to be "inappropriate" and the response is "filtered." For example, a one-bit flag may be used, and when the flag is "0," the flag may be configured to indicate a state where the response is judged to be "not inappropriate" and the response is "unfiltered." In this case, when the flag is "1," the flag may be configured to indicate a state where the response is judged to be "inappropriate" and the response is "filtered."
[0290] As explained above, the inappropriate filter flag in Fig. 7D(1) makes it possible to more appropriately distinguish, with one bit, between a state in which a response is judged to be "not inappropriate" and is "unfiltered" and a state in which a response is judged to be "inappropriate" and is "filtered." It also makes it possible to more appropriately determine the possibility of outputting an inappropriate response.
[0291] In contrast, the example of the flag in Figure 7D (2) is an example in which, when a response is judged to be "inappropriate" and the response is "filtered," the reason for the judgment that the response was "inappropriate" can also be identified by the flag. Specifically, the example of the inappropriate filter flag in (2) is configured to indicate eight states from "0" to "7" in decimal notation. When the inappropriate filter flag in (2) indicates "0," the state is the same as in (1), so repeated explanation will be omitted.
[0292] Here, when the inappropriate filter flag (2) indicates "1" to "6," the flag indicates that the instruction or generated data was judged to be "inappropriate" and the response was "filtered." The reason for the "inappropriate" judgment can be identified by the value of the flag. For example, when the inappropriate filter flag indicates "1," it indicates that the instruction or generated data contains content related to "violence," which has led to the instruction being judged to be "inappropriate," and the response being "filtered." When the inappropriate filter flag indicates "2," it indicates that the instruction or generated data contains "sexual" content, which has led to the instruction being judged to be "inappropriate," and the response being "filtered." When the inappropriate filter flag indicates "3," it indicates that the instruction or generated data contains content that may "violate" "public safety," which has led to the instruction being judged to be "inappropriate," and the response being "filtered."
[0293] Furthermore, when the inappropriate filter flag indicates "4," it indicates that the instruction or generated data contains content that may violate "personal privacy," and is therefore judged to be "inappropriate" and the response has been "filtered." When the inappropriate filter flag indicates "5," it indicates that the instruction or generated data contains content that may promote "discrimination," and is therefore judged to be "inappropriate," and the response has been "filtered." When the inappropriate filter flag indicates "6," it indicates that the instruction or generated data contains content that raises "academic concerns," and is therefore judged to be "inappropriate," and the response has been "filtered." In the example of Figure 7D(2), the inappropriate filter flag can indicate "7," which indicates a situation that does not fall under the above-mentioned inappropriate filter flag settings of "0" to "6."
[0294] In the example of the flag in Fig. 7D(2), the flag is configured to indicate eight states in decimal notation from "0" to "7." Therefore, the inappropriate filter flag may be configured as a 3-bit flag.
[0295] As explained above, the inappropriate filter flag in Figure 7D(2) not only makes it possible to distinguish between a state in which a response is judged to be "not inappropriate" and the response is "unfiltered" and a state in which a response is judged to be "inappropriate" and the response is "filtered," but also makes it possible to identify the reason for the judgment that the response was "inappropriate."
[0296] 7B is a large-scale language model on a network such as in the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001, the flag described in FIG. 7D is transmitted from the large-scale language model to the AI response output device 10010 via a network such as the Internet 19000. The inappropriate filter flag received by the AI response output device 10010 can be used by the control unit 1110 for processing.
[0297] 7B is a local large-scale language model such as the local LLM processing unit 10028, the flag described in Fig. 7D is transmitted from the local LLM processing unit 10028 to the control unit 1110 within the AI response output device 10010. The control unit 1110 can use the received inappropriate filter flag for processing.
[0298] Next, an example of control performed by the control unit 1110 of the AI response output device 10010 that has received an inappropriate filter flag will be described with reference to Fig. 7E. Fig. 7E shows an example of a database in which responses corresponding to the inappropriate filter flags described above are prepared and stored for each of the multiple characters shown in Fig. 2H. The character IDs and character names are the same as those in Example 2, Example 3, or Example 4, and therefore will not be described again.
[0299] In the second, third, or fourth embodiment, the user is devised to feel as if each character has a consistent personality by using the initial setting instruction sentence, the conversation history, or the correspondence with the user ID, etc. In this case, consider a case where an inappropriate filtering process is performed on the response from the large-scale language model, and the AI response output device 10010 outputs to the user a replacement sentence after the inappropriate filtering process, such as those shown in (1) to (4) of Figure 7A.
[0300] Here, sentences (1) to (4) in Figure 7A contain expressions such as "I am AI," and sentences (1), (2), and (4) in Figure 7A contain expressions such as "I am programmed." Therefore, a user viewing this output may be strongly aware that the multiple characters shown in Figure 2H are artificial intelligences programmed by software, and may not be able to recognize the consistent personalities of each character that has been built up so far.
[0301] Therefore, in the example of Fig. 7E of this embodiment, even if the filtered natural language response sentence generated by the large-scale language model has the contents of (1) to (4) in Fig. 7A, the control unit 1110 of the AI response output device 10010 uses the database shown in Fig. 7E to select a sentence (response sentence) that will be a response according to each character and the inappropriate filter flag, and controls to output the response to the user. That is, under the control of the control unit 1110, the natural language response sentence generated by the large-scale language model (e.g., the contents of (1) to (4) in Fig. 7A) is not output directly to the user, but is replaced with the contents stored in the database of Fig. 7E and output to the user. The database shown in Fig. 7E may be stored in the storage unit 1170 of Fig. 3.
[0302] For example, when the character "Koto" with a character ID of 1 is displayed on the AI response output device 10010, if the control unit 1110 receives an inappropriate filter flag with a value of 1 (a flag indicating that the instruction sentence or generated data contains content related to "violence") shown in FIG. 7D(2) in step 702 of FIG. 7B, the control unit 1110 refers to the database of FIG. 7E stored in the storage unit 1170, selects a response sentence such as "Koto is against violence," and outputs it to the user. By controlling the response sentence output in this manner, the user can recognize that the response is more in line with the personality of the character displayed by the AI response output device 10010 than when a response sentence generated by a large-scale language model is directly output, as shown in FIG. 7A(1), which is preferable.
[0303] Also, for example, when the character "Koto" with a character ID of 1 is displayed on the AI response output device 10010, if the control unit 1110 receives an inappropriate filter flag with a value of 2 shown in FIG. 7D(2) (a flag indicating that the instruction or generated data contains "sexual" content) in step 702 of FIG. 7B, and if the control unit 1110 receives an inappropriate filter flag with a value of 3 shown in FIG. 7D(2) (a flag indicating that the instruction or generated data contains content that may "violate" "public safety"), the control unit 1110 may refer to the database of FIG. 7E stored in the storage unit 1170 and control to output a corresponding response to the user. This allows the user to recognize that the response is in line with the personality of the character "Koto" displayed on the AI response output device 10010, which is more preferable.
[0304] Similarly, when the character "Tom" with a character ID of 2 is displayed on the AI response output device 10010, if the control unit 1110 receives an inappropriate filter flag with a value of 1 (a flag indicating that the instruction or generated data contains content related to "violence") shown in FIG. 7D(2) in step 702 of FIG. 7B, if the control unit 1110 receives an inappropriate filter flag with a value of 2 (a flag indicating that the instruction or generated data contains content related to "sexuality") shown in FIG. 7D(2), or if the control unit 1110 receives an inappropriate filter flag with a value of 3 (a flag indicating that the instruction or generated data contains content that may "violate" "public safety") shown in FIG. 7D(2), the control unit 1110 may refer to the database of FIG. 7E stored in the storage unit 1170 and control to output a corresponding response to the user. This allows the user to recognize that the response is in line with the personality of the character "Tom" displayed on the AI response output device 10010, which is more preferable.
[0305] Similarly, when the character "Necco" with a character ID of 3 is displayed on the artificial intelligence response output device 10010, if the control unit 1110 receives an inappropriate filter flag with a value of 1 shown in FIG. 7D(2) (a flag indicating that the instruction or generated data contains content related to "violence") in step 702 of FIG. 7B, if the control unit 1110 receives an inappropriate filter flag with a value of 2 shown in FIG. 7D(2) (a flag indicating that the instruction or generated data contains content related to "sexual" content), or if the control unit 1110 receives an inappropriate filter flag with a value of 3 shown in FIG. 7D(2) (a flag indicating that the instruction or generated data contains content that may "violate" "public safety"), the control unit 1110 will refer to the database of FIG. 7E stored in the storage unit 1170 and control the output of a corresponding response to the user.
[0306] In the example of FIG. 7E, when the character "Necco" receives an inappropriate filter flag with a value of 1 to 6 as shown in FIG. 7D(2), the control unit 1110 controls to consistently output a response of "Necco doesn't really understand, meow." By doing so, the AI response output device 10010 can give the user the impression that the character "Necco" has the personality of pretending not to understand and evading answers to inappropriate commands and responses. In this case, too, the user can recognize that the response is in line with the personality of the character "Necco" displayed on the AI response output device 10010, which is more preferable.
[0307] According to the control using the database of Figure 7E described above, even if the response from the large-scale language model is a response filtered by the large-scale language model, it can be replaced with a response that is in line with the personality of the character displayed by the AI response output device, which is more convenient for the user.
[0308] The response sentence database of Fig. 7E described above may be stored in the storage unit 1170 and used by the control unit 1110 of the AI response output device 10010. However, the response sentence database of Fig. 7E may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side.
[0309] In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may acquire the character ID of the character displayed on the artificial intelligence response output device 10010 from the artificial intelligence response output device 10010, replace the response sentence generated by the large-scale language model of each server with the corresponding response sentence from the response sentence database in Fig. 7E, and transmit the same to the artificial intelligence response output device 10010. In this case, the artificial intelligence response output device 10010 may output the response sentence to the user.
[0310] In this way, it is possible to realize an AI response output system that can generate responses using the response sentence database of Fig. 7E even if the AI response output device 10010 is not equipped with the response sentence database of Fig. 7E. In this case, within the large-scale language model server 19001 or the large-scale language model server 20001, the flags shown in Fig. 7D may be transmitted and received between the respective large-scale language models and the respective control units.
[0311] Example 8 Example 8 of the present invention is an improvement of the AI response output device 10010 or AI response output system described in the drawings of Examples 1 to 7. Note that in Example 8, differences from Examples 1 to 7 will be described, and repeated description of the same configurations as those examples will be omitted.
[0312] As in the above-described embodiments, the AI response output device 10010 may be referred to as a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0313] Incidentally, in the response request (prompt) sent from the AI response output device 10010 to the large-scale language model, in addition to text information such as a question (request sentence) written in natural language, etc., data files of media content such as images, videos, and music (hereinafter referred to as content files) can be registered. Note that the response request can also be called a command sentence, as in the above-mentioned embodiment.
[0314] When a user registers a content file in a response request, if the content file to be registered in the response request has been determined in advance, the user can simply register that content file in the response request. On the other hand, if the content file to be registered has not been determined or an appropriate content file cannot be found, the user must search for a content file related to the content of the response request from a storage location where multiple content files are saved, for example.
[0315] In such cases, conventionally, a user would access the storage destination and sequentially check the contents of multiple content files to find a content file related to the content of the response request. Therefore, for example, in a situation where there are multiple storage destinations or a large number of content files, the task of finding the content file becomes complicated, and there is a risk that it will take a long time to register the content file.
[0316] Therefore, the AI response output device 10010 according to the eighth embodiment performs a search process and a registration process as support processes when the user registers a content file. More specifically, the control unit 1110 of the AI response output device 10010 performs a search process based on search conditions specified by the user, extracts registration candidates for content files to be registered in the response request, and displays a list of the extracted registration candidates on the display unit 10011. Furthermore, a registration process is performed to register a content file selected by the user from the extracted registration candidates in the response request.
[0317] In this way, by having the AI response output device 10010 perform the search process and registration process, the user can register a content file appropriate for the response request by selecting one of the content files from the list of registration candidates extracted by the AI response output device 10010. Therefore, the user's effort in the registration operation can be significantly reduced, and the time required for the registration operation can be shortened.
[0318] The procedure for generating a response request by the AI response output device 10010, particularly the content file search process and registration process in response to the response request, will be described in more detail below.
[0319] The procedure for generating a response request by the AI response output device 10010 described below is an example of registering a first content file and a second content file for the response request. Note that the first content file refers to the first content file registered for the response request, and the second content file refers to the second content file registered for the response request. It is also possible to register three or more content files for the response request.
[0320] Fig. 8 is a flowchart showing an example of a response request generation process by the AI response output device. When creating a response request, the AI response output device 10010 first searches for a candidate for registering a first content file for the response request in step S1, as shown in Fig. 8, and then registers the first content file based on the result of the search (step S2).
[0321] Here, as shown in Figure 9A, when the control unit 1110 of the artificial intelligence response output device 10010 performs search processing and registration processing, the display unit 10011 is set with an instruction text display area (response request display area) 10051 as well as a registration operation area 10071 for performing content file registration operations.
[0322] As explained in the above embodiment, content files such as images 10054 and videos 10055 can be registered in the instruction statement display area (response request display area) 10051 along with text 10053. That is, the response request display area 10051 has, in addition to an input area 10056 in which text 10053 can be input, registration areas in which multiple content files such as images, videos, and music can be registered, in this example, registration areas 10057A and 10057B in which two content files can be registered.
[0323] When performing the search process in step S1, input information for inputting search conditions for searching for registration candidates for the first content file is first displayed in the registration operation area 10071. As an example, the registration operation area 10071 displays information prompting the input of search conditions such as the search target, search location, time when the search target was saved, and keywords related to the search target. In the example shown in FIG. 9A, along with the comment "Please specify search conditions," item names such as "What," "From where," "When," and "Other," and input fields for each item are displayed. This prompts the user to input search conditions such as the search target, search location, time when the search target was saved, and keywords.
[0324] In the example of Figure 9A, the search targets "photo" and "video" are specified in the "what" field, the search location "server storage" is entered in the "where" field, the time of storage "July of this year" is entered in the "when" field, and the user enters related keywords "mountain climbing" and "Mount Fuji" in the "other" field.
[0325] Also, a "Start Search" button 10072 is set in the registration operation area 10071. After specifying the search conditions as described above, the user selects this button 10072, and the control unit 1110 starts a search process for registration candidates for the first content file (step S1).
[0326] Incidentally, in the AI response output device 10010 or the AI response output system, a typical example of a location (destination) where a user saves content files is a storage area 19012 reserved for each user in a server 19002 on a network, as shown in FIG. 10 . Each user can create any folder within the storage area 19012. In this example, an image / video folder 19013 and a music folder 19014 have been created within the storage area 19012. A plurality of image files 19015 and a video file 19016, which are content files, are stored in the image / video folder 19013. A plurality of music files 19017 are stored in the music folder 19014.
[0327] Furthermore, the location where the user saves the image files 19015, video files 19016, and music files 19017 is not limited to the server 19002, but may be, for example, the storage unit 1170 of the AI response output device 10010, as shown in Fig. 11. The storage unit 1170 may be configured as an internal memory 10210 of the AI response output device 10010, or may be configured as an external recording medium 10310 connectable to the AI response output device 10010.
[0328] 11 , an image / video folder 10211 and a music folder 10212 are created in internal memory 10210. A plurality of image files 10213 and video files 10214 are stored in image / video folder 10211, and a plurality of music files 10215 are stored in music folder 10212. Similarly, an image / video folder 10311 and a music folder 10312 are created in external recording medium 10310. A plurality of image files 10313 and video files 10314 are stored in image / video folder 10311, and a plurality of music files 10315 are stored in music folder 10312.
[0329] The control unit 1110 of the AI response output device 10010 performs a content file search process for the search location, which is a storage destination designated by the user, among these storage locations, and extracts registration candidates for the first content file that meets the search conditions. As an example, the control unit 1110 analyzes the contents of each content file, such as a photo (image) or video (moving image), and extracts, from among the multiple content files, a content file that was taken in July of this year and includes an image of a person climbing a mountain and Mount Fuji as a registration candidate.
[0330] The method for analyzing the content file is not particularly limited, and existing image analysis technology, video analysis technology, etc. may be applied. Therefore, a detailed description of the method for analyzing the content file will be omitted. Furthermore, the analysis of the content file itself does not necessarily have to be performed by the AI response output device 10010, and may be performed by a device other than the AI response output device 10010.
[0331] When the control unit 1110 starts a content file search process, the display in the registration operation area 10071 switches to displaying the results of the search process. As an example, as shown in FIG. 9B, a list of registration candidates for the first content file is displayed in the registration operation area 10071 as a search result. In the example of FIG. 9B, the list of registration candidates includes the file name of each registration candidate, information such as the date and time of saving (date and time of shooting), and a thumbnail image, based on the metadata of the content file. Note that the content file information displayed as the list of registration candidates is not particularly limited.
[0332] Thereafter, when the user selects a desired content file from the list of registration candidates for the first content file displayed in the registration operation area 10071, the control unit 1110 performs a registration process to register the selected content file in a response request (request prompt) (step S2). For example, when the user selects "image1.jpg", which is one of the registration candidates, the selected "image1.jpg" is registered in the response request. When the list of registration candidates for the first content file is displayed, a message may be displayed to the user prompting them to select a registration candidate.
[0333] 12A, the user can select a registration candidate (content file) by, for example, performing a file move operation such as drag and drop, in which the desired content file (image1.jpg) is moved from registration operation area 10071 to registration area 10057A of response request display area 10051. By this file move operation, the desired content file (image1.jpg) is registered as a first content file in registration area 10057A that constitutes response request display area 10051, as shown in FIG.
[0334] However, the method by which the user selects a content file is not particularly limited, and may be an operation other than a file move operation. When the user performs a specific operation other than a file move operation on a content file that is a registration candidate displayed on the display unit 10011, the control unit 1110 may determine that the content file on which the specific operation was performed has been selected by the user. For example, instead of moving the desired content file (image1.jpg), the user may perform a specific operation such as long pressing on the content file (image1.jpg), and the content file (image1.jpg) may be registered in the response request.
[0335] Furthermore, when registering a content file in a response request, for example, a display may be displayed to the user asking whether or not to register the selected content file in the response request. In other words, before registering the content file in the response request, the control unit 1110 may perform a registration necessity confirmation process to confirm whether or not to register the content file in the response request. For example, when a content file is selected from a list of content files that are candidates for registration by a specific operation such as a long press by the user, as shown in FIG. 13A, a pop-up message is displayed to the user asking, "Do you want to use the selected image in the AI request?" along with options for the user's response, "Yes" and "No." If the user then selects "Yes," the selected content file is registered in the response request.
[0336] At this time, a message may be displayed to inform the user that the content file has been registered in the response request. For example, as shown in Fig. 13B, when a content file (Image1.jpg) is registered in the response request registration area 10057A, a pop-up message may be displayed stating, "Image1.jpg has been set as one of the AI request information."
[0337] Furthermore, before registering the content file in a response request, the content file selected by the user may be played back for a short time, for example, as a preview display, so that the user can check part of the content of the content file. When the control unit 1110 determines that the content file has been selected by the user, the control unit 1110 may play back the entire content of the selected content file before registering it in a response request. For example, as shown in FIG. 14, a pop-up window is displayed offering two options for operations on the selected file: "Play and check" and "Use for AI request." If the user selects the "Play and check" option, the content file (Image1.jpg) is played. After playback is complete, the user may be prompted with a question such as, "Do you want to use this content file for an AI request?" If the "Use for AI request" option is selected, the content file (Image1.jpg) may be registered in the response request registration area 10057A.
[0338] In this example, the control unit 1110 of the AI response output device 10010 performs a search process based on search conditions specified by the user to extract candidates for registering the first content file. However, without performing the search process by the control unit 1110 in step S1, a content file arbitrarily selected by the user may be registered as the first content file in the response request in the registration process in step S2.
[0339] In the above example, an input screen for inputting search conditions for searching for candidates for the first content file is first displayed in the registration operation area 10071, but if the user wants to register an arbitrary content file, multiple storage locations (folders) in which the content files are saved are displayed in the registration operation area 10071. As an example, when the input screen for inputting search conditions is displayed in the registration operation area 10071, the user touches the registration area 10057A of the response request display area 10051, and the display in the registration operation area 10071 switches from the input screen for inputting search conditions to a display of a list of storage locations (folders).
[0340] Thereafter, when the user selects any folder in registration operation area 10071, a list of content files in the selected folder is displayed in registration operation area 10071, as shown in Fig. 15, for example. The example shown in Fig. 15 is an example in which the user selects an image folder in which multiple image files are saved, and shows a state in which a list of image files in the selected image folder is displayed. Then, control unit 1110 may register the content file (image1.jpg) selected by the user from this list in registration area 10057A of the response request (request prompt) as the first content file.
[0341] The above is an explanation of the search process in step S1 and the registration process in step S2 in the flowchart of Fig. 8. Next, the process proceeds to step S3, where the first content file post-registration process is executed.
[0342] In this example, as a post-registration process for the first content file, a search confirmation process is performed to confirm with the user whether or not a search for information (content files) related to the first content file is necessary. As an example, as shown in FIG. 16, a confirmation question, "Do you want to search for information related to registered images?", and options of "Yes" and "No" as the user's response to the confirmation question are displayed. If the user selects "Yes," the process proceeds to step S4, where the control unit 1110 starts a search process to search for registration candidates for a second content file related to the first content file.
[0343] Here, if a first content file is registered in the response request, the first content file becomes one of the search conditions for searching for a second content file. That is, if a first content file is registered in the response request, when the search process of step S4 is started, control unit 1110 searches for registration candidates for the second content file using the search condition that the file is related to the first content file. In the example shown in Fig. 16, since the first content file (image1.jpg) is registered in registration area 10057A, control unit 1110 performs search process for registration candidates for the second content file using the search condition that the file is related to this first content file (image1.jpg).
[0344] 16, the process proceeds to step S5 without searching for a second content file in step S4, and control unit 1110 performs processing to generate a request statement (text), as described below. Furthermore, the search confirmation processing for the user in step S3 does not necessarily have to be performed. After registering the first content file in the response request in the registration processing in step S2, control unit 1110 may start processing to search for registration candidates for a second content file (step S4) without performing the post-registration processing in step S3.
[0345] Next, an example of a method for extracting registration candidates for the second content file in step S4 will be described. As the search process of step S4, control unit 1110 evaluates the relevance (degree of relevance) between the first content file and each content file to be searched, based on information about the first content file and information about each content file to be searched. In the following description, the evaluation of the relevance between the first content file and each content file to be searched may be referred to as "file relevance evaluation."
[0346] An example of the method for evaluating file relevance in step S4 is to analyze the contents of the first content file and the contents of each content file to be searched, and evaluate the relevance between them based on the analysis results. In this method, multiple feature points of the first content file are identified based on the analysis results of the first content file. Note that the analysis of the first content file can be performed using existing image analysis technology or video analysis technology, so a detailed description of the analysis technology will be omitted.
[0347] Next, the relevance between the identified first content file and each of the search target content files is evaluated based on the degree of match between the feature points of the first content file and the content of each of the search target content files. For example, the relevance between the two is determined based on an evaluation value calculated from the degree of match between the feature points of the first content file and the content of each of the search target content files. More specifically, for example, if the calculated evaluation value is equal to or greater than a predetermined first threshold, the content file is determined to be "related to the first content file." On the other hand, if the evaluation value is less than the first threshold, the content file is determined to be "unrelated to the first content file." Then, content files determined to be "related to the first content file" are extracted as candidates for registration of the second content file.
[0348] As an example, as shown in Figure 17A, suppose the first content file registered in the response request is an image of an animal character, and as a result of image analysis, four keywords, "animal (panda)," "human-like," "black and white," and "around the eyes," are extracted as characteristic points of the content of the first content file.
[0349] In contrast, if the content file to be searched is, for example, file 11 (image11.jpg), which is an image of a person, there will be few matches between the four feature points and the content of file 11, and the calculated evaluation value will be low. For this reason, file 11 will likely be determined to be "unrelated to the first content file." On the other hand, if the content file to be searched is, for example, file 12 (image12.jpg), which is an image of two animal characters, there will be more matches between the four feature points and the content of file 12 than file 11, and the calculated evaluation value will be higher than file 11. For this reason, file 12 will likely be determined to be "related to the first content file." The first threshold value can be set arbitrarily.
[0350] Another method for file relevance evaluation is to evaluate the relevance between a first content file and each content file to be searched, based on the metadata (meta information) of the first content file and the metadata of each content file to be searched. In this method, the relevance is evaluated according to the degree of agreement between the information extracted from the metadata of both files.
[0351] More specifically, for example, an evaluation value is calculated based on the degree of match between information extracted from the metadata of each content file to be searched and information extracted from the metadata of the first content file (e.g., the number of matches between the two pieces of information), and the relevance between the two is evaluated based on the calculated evaluation value. As an example, if this evaluation value is equal to or greater than a predetermined second threshold, the content file is determined to be "related to the first content file." On the other hand, if the evaluation value is smaller than the second threshold, the content file is determined to be "not related to the first content file." The second threshold can be set arbitrarily.
[0352] In the example shown in Figure 17B, file relevance evaluation is performed based on three pieces of information extracted from the metadata recording the information when the first content file was recorded: "recording date," "recording location," and "device used." In this example, for the search target file 3 (movie3.mov), all of the above three pieces of information in the metadata recording the information when file 3 was recorded match the information in the first content file, so the calculated evaluation value is high and it is likely to be determined to be "related to the first content file." On the other hand, for the search target file 4 (movie4.mov), none of the above three pieces of information in the metadata recording the information when file 4 was recorded match the information in the first content file, so the calculated evaluation value is low and it is likely to be determined to be "not related to the first content file."
[0353] In the above example, the registration candidate for the second content file is extracted based on the result of one file relevance evaluation. However, the registration candidate for the second content file may be extracted based on the results of multiple types of file relevance evaluation. That is, in step S4, multiple types of file relevance evaluation may be performed, and the relevance between the first content file and the content file to be searched may be determined comprehensively based on the results of these file relevance evaluations. For example, an evaluation value may be calculated for each file relevance evaluation, and the presence or absence of relevance between the first content file and the content file to be searched may be determined based on the sum of these evaluation values.
[0354] When multiple types of file relevance evaluations are performed, the control unit 1110 stores, for example, data on the results of each file relevance evaluation (hereinafter referred to as evaluation result data) and successively updates this evaluation result data. Then, the control unit 1110 comprehensively evaluates the relevance between the first content file and the content file to be searched based on this evaluation result data.
[0355] As an example, we will explain a case where two types of file relevance evaluations are performed for each content file to be searched, and the presence or absence of relevance between the first content file and each content file is determined based on the results of these file relevance evaluations.
[0356] FIG. 18 is a diagram showing a list of result information of the relevance assessment for each content file when two types of file relevance assessment are performed. In the example shown in FIG. 18, the "save destination" and "file name" of each content file are listed together with information on "relevance assessment 1," "relevance assessment 2," "total relevance assessment value," and "candidate flag" as result information of the relevance assessment. "Relevance assessment 1" and "relevance assessment 2" are evaluation values calculated using the two types of file relevance assessment. While the setting criteria for "relevance assessment 1" and "relevance assessment 2" are different, in this example, both are set on a ten-point scale from 0 to 9. The relevance between the first content file and each content file being searched is determined based on the total value of these, which is the "total relevance assessment value." Specifically, if the "total relevance assessment value" is greater than 10, the content file is determined to be "related to the first content file," and the candidate flag is set to "1" (the candidate flag is ON). On the other hand, if the "total relevance evaluation value" is 10 or less, the content file is determined to be "not related to the first content file" and the candidate flag is set to "0" (candidate flag is turned OFF).
[0357] The control unit 1110 extracts registration candidates for the second content file based on the result information of such file relevance evaluation. That is, the control unit 1110 extracts content files whose candidate flags are set to "1" as registration candidates for the second content file.
[0358] The criteria for determining whether or not the first content file is related to each content file to be searched are not particularly limited. For example, even if the "total relevance evaluation value" is 10 or less, if the value of either "relevance evaluation 1" or "relevance evaluation 2" is equal to or greater than a predetermined threshold, for example, equal to or greater than 8, the content file may be determined to be "related to the first content file."
[0359] Furthermore, if the search target may include, for example, a music file, it is preferable to facilitate extraction of keywords representing the atmosphere or image of the first content file as feature points identified from the analysis results of the first content file. For example, as shown in FIG. 19A, if the first content file registered in the response request is an image of an animal character, it is preferable to facilitate extraction of keywords such as "animal," "black and white," and "panda" as well as keywords representing the atmosphere or image, such as "cute," "bright," and "child-friendly," as feature points of the first content file. For example, as shown in FIG. 19B, if the first content file is an image of a landscape, it is preferable to facilitate extraction of keywords such as "nature," "mountains," "trees," and "green," as well as keywords representing the atmosphere or image, such as "calm" and "tranquil." Performing a search process based on these feature points (keywords) also makes it easier to extract music content files as candidates for registration of the second content file.
[0360] When control unit 1110 extracts registration candidates for second content files related to the first content file based on the results of the file relevance evaluation as described above, it displays a list of the extracted registration candidates on display unit 10011. For example, as shown in Fig. 20A, a list of registration candidates for second content files, such as image files, video files, or music files, is displayed in registration operation area 10071 of display unit 10011.
[0361] The number of registration candidates for second content files displayed in the registration operation area 10071 is not particularly limited, but an upper limit may be set on the number to be displayed. In other words, in the search process in step S4, an upper limit may be set on the number of extraction candidates for second content files to be registered, and the search process may be terminated when the extraction number reaches the upper limit. Of course, the search process may be performed on all content files to be searched without setting an upper limit on the extraction number.
[0362] Furthermore, the order in which the multiple registration candidates are displayed is not particularly limited, but as an example, the registration candidates for the second content file are displayed in the order in which they were extracted. In other words, during the registration candidate search process, the detected registration candidates are displayed sequentially in real time. Alternatively, after the registration candidate search process is completed, the candidates may be rearranged in a predetermined order, and the results may be displayed as a list of registration candidates. For example, as shown in FIG. 20B, the list of registration candidates may be displayed in descending order of evaluation value (e.g., "total relevance evaluation value"), that is, in descending order of relevance to the first content file.
[0363] Furthermore, among the list of registration candidates, registration candidates that are highly related to the first content file may be displayed in an emphasized manner. For example, registration candidates with high evaluation values may be displayed in a color different from the other registration candidates, so that the highly evaluated registration candidates stand out. Furthermore, if there are multiple types of content files in the list of registration candidates, the registration candidates may be categorized and displayed by type. This makes it easier for the user to select a desired one from the multiple registration candidates.
[0364] In the search process of step S4, all types of content files stored in the search location may be searched for, but the types of content files to be searched for may be specified as necessary. In other words, the search process may search only for search targets specified depending on the situation.
[0365] For example, if the same type of second content file as the first content file is likely to be selected, the search targets for registration candidates for the second content file may be limited to those of the same type as the first content file. For example, if an image file is registered as the first content file, the search targets may be limited to image files. Of course, the search targets for registration candidates may include those of a different type from the first content media. In that case, it is preferable to search with priority given to those of the same type as the first content media. Note that if the user has specified a search order, the search process is performed according to the specified search order. This may shorten the search time.
[0366] Furthermore, for example, if a type of second content file different from that of the first content file is likely to be selected as the second content file, the search targets for registration candidates for the second content file may be limited to types different from that of the first content file. For example, if an image file is registered as the first content file, the search targets may be limited to video files and music files. In the example shown in FIG. 21A, the first content file is an image file, and four image files, three video files, and two music files are stored in the storage area that is the search location. In this case, the search targets may be three video files (video files 1 to 3) and two music files (music files 1 and 2), excluding the four image files (image files 1 to 4).
[0367] Of course, the search targets for registration candidates for the second content file may include the same type of content media as the first content file. However, in this case, it is preferable to prioritize searching for types of content different from the first content file. For example, in the example shown in FIG. 21B, an image file is registered as the first content file, and four image files, three video files, and two music files are stored in the storage area that is the search location. In this case, the search order for registration candidates for the second content file is three video files (video files 1 to 3), two music files (music files 1 and 2), and four image files (image files 1 to 4).
[0368] Furthermore, the search location in the search process of step S4 is not particularly limited, and may be set to a storage location specified by the user as described above, or may be set in advance as necessary. For example, the storage location of the registration candidates for the second content file may first be the same storage area as the storage location (storage unit) of the first content file. For example, as shown in FIG. 22A, assume that there are three storage locations for content files: Storage 1, Storage 2, and Web, and the first content file is stored in Storage 1. In this case, the search location (search target) for the registration candidates for the second content file may be limited to Storage 1, which is the same storage area as the storage location of the first content file. In other words, Storage 2 and Web may be excluded from the search locations for the registration candidates for the second content file.
[0369] Of course, in addition to the same storage area as the storage destination of the first content file, other storage areas may also be specified as search locations. In this case, it is preferable to search the same storage area (storage unit) as the storage destination of the first content file preferentially. For example, in the example shown in FIG. 22B, the search locations for the second content file to be registered include not only Storage 1, which is the storage destination of the first content, but also Storage 2 and the Web. However, the order of search priority is set as Storage 1, Storage 2, and the Web. Therefore, in the search process, the search will be performed in the order of Storage 1, Storage 2, and the Web. In this way, by appropriately specifying the search location and search order, it is possible to shorten the search time.
[0370] Furthermore, for example, the storage destination of the registration candidates for the second content file may be a storage area different from the storage destination (folder) of the first content file. Therefore, the search location for the registration candidates for the second content file may be limited to a storage area different from the storage destination of the first content file. Of course, in this case, the same storage area as the storage destination of the first content file may be specified as the search location, along with a storage area different from the storage destination of the first content file. In this case, it is preferable to search storage areas different from the storage destination of the first content file with priority. This may shorten the search time.
[0371] As described above, in this example, in the second content file search process in step S4, the search process is performed using the degree of relevance to the first content file as a search criterion, but the search criterion is not limited to this. In the search process in step S4, the user may be allowed to specify various search criteria, as in the first content file search process in step S1.
[0372] The above is an explanation of the search process in step S4 in the flowchart of Fig. 8. Next, the process proceeds to step S5, where the registration process for the second content file is performed. This registration process for the second content file in step S5 is similar to the registration process for the first content file in step S2, so it will only be briefly explained.
[0373] In step S5, when the user selects a desired content file from the list of registration candidates for the second content file displayed in the registration operation area 10071, the control unit 1110 performs a registration process to register the selected content file in a response request (request prompt). For example, as shown in FIG. 23A, when the user selects "movie31.mov", which is one of the registration candidates for the second content file, as shown in FIG. 23B, the selected "movie31.mov" is registered in the registration area 10057B that constitutes the response request display area 10051. This completes the flow of the registration process for the second content file in step S5.
[0374] When displaying the list of registration candidates for the second content file in step S5, a message may be displayed prompting the user to select a registration candidate. Alternatively, a confirmation process may be performed to confirm with the user whether or not there is any registration candidate that the user wishes to register as a second content file. If the user replies to this confirmation process that there is no file that the user wishes to register as a second content file, the process may return to step S4 without performing the registration process in step S5, change the search conditions, and perform the search process for registration candidates again. Furthermore, before performing the search process for registration candidates again, a confirmation process may be performed to confirm with the user whether or not to perform a re-search. Furthermore, when performing a re-search for registration candidates, it is preferable to exclude content files extracted as registration candidates in previous search processes from the search process.
[0375] In this example, two content files are registered in a response request, but three or more content files may be registered in a response request. In this case, after the second content file is registered, it is preferable to perform a candidate confirmation process to ask the user whether or not any of the candidates for registering the second content file has subsequently been registered as a third content file.
[0376] As an example, this confirmation process may be performed when the registration candidates for the second content file include a content file of a type different from the first content file and the second content file already registered in the response request. In this example, since the first content file is an image file and the second content file is a video file, the candidate confirmation process may be performed when the registration candidates include files other than image files and video files (e.g., music files). Alternatively, the candidate confirmation process may be performed when the registration candidates for the second content file include files with particularly high evaluation values other than those already registered.
[0377] If the user responds to this candidate confirmation process that there is a file that they would like to register as a third content file, the registration process of step S5 is executed again. Then, as in the case of registering the second content file, one of the third content file registration candidates selected by the user is registered as the third content file in the response request.
[0378] When all content files have been registered in the response request in this manner, the process proceeds to step S6, where the control unit 1110 performs a request statement generation process. The request statement generation process is a process for generating a request statement to be registered in the response request in accordance with a user instruction. Specifically, the process generates a request statement (text 10053) to be entered in an input field 10056 that constitutes the response request display area 10051.
[0379] When the request statement generation process of step S6 starts, for example, as shown in FIG. 24A, the registration operation area 10071 switches to a display that prompts the user to input a request. In other words, an input screen is displayed where the user inputs information necessary to generate a request statement. As an example, first, a question 1, "What do you do?", and options "search" and "synthesize" as answers to the question 1 are displayed. In this example, the user is allowed to select preset keywords as a method of inputting information necessary to generate a request statement, rather than manually entering text.
[0380] When an answer to this question 1 is selected, the next question 2 is displayed. In the example of FIG. 24A, "search" is selected as the answer to question 1, so question 2, "What do you want to search for?", is displayed. An answer sentence to question 2 is also displayed, and the user can complete the answer sentence by appropriately selecting a part of the sentence from multiple preset keywords (options). For example, all preset keywords are displayed, and the user can complete the answer sentence by selecting any keyword from among them. Furthermore, when question 2 is answered, the next question is displayed as necessary.
[0381] The method for displaying the keywords is not particularly limited. For example, a plurality of preset keywords may be displayed in descending order of frequency of use by the user. Furthermore, a predetermined number of a plurality of preset keywords may be displayed in descending order of frequency of use. When there are multiple users, it is preferable to determine the frequency of keyword use for each user. For example, the control unit 1110 may predict, based on the types and contents of the first and second content files registered in the response request, which of the preset keywords are likely to be selected by the user, and display the keywords in descending order of probability. Furthermore, it is preferable that the keywords include keywords that represent the features of the first content file extracted during the search process.
[0382] Furthermore, in the request statement generation process of step S6, the control unit 1110 may generate multiple request statement candidates based on, for example, the types and contents of the first content file and the second content file registered in the response request. Then, when the user selects one of the multiple request statements, the selected request statement may be registered in the input area 10056 constituting the response request display area 10051.
[0383] When the user inputs answers to each question, the control unit 1110 generates a request sentence based on the information input by the user and inputs it into the prompt. As an example, the control unit 1110 generates a request sentence such as "Search for scenes in which the person in the image appears in the video" based on the information input by the user shown in FIG. 24A, and registers the text of this request sentence in the input field 10056 as shown in FIG. 24B. This creates a response request including the first content file and the second content file. Then, the process proceeds to step S7, where the response request including the first content file and the second content file is sent to the large-scale language model.
[0384] The method of inputting the request by the user in step S6 may be, for example, a method in which the user directly inputs a request statement, but is not particularly limited to this. The method of inputting the request by the user may be text input, voice input, or of course other input methods.
[0385] <Modification of Example 8> When a response request (instruction) is generated based on a user input requesting a response from a large-scale language model input via the AI response output device 10010 and the response request is sent to the large-scale language model to obtain a response, it is conceivable that a content file registered in the response request may contain inappropriate content. In this case, if a response request in which a content file with inappropriate content is registered is sent to the large-scale language model, there is a risk that the response from the large-scale language model may contain inappropriate content.
[0386] A modified example of the eighth embodiment is an example in which an inappropriate media content determination process (hereinafter referred to as "content determination process") is performed to determine whether or not the content of a content file registered in a response request contains inappropriate content before transmission to a large-scale language model. If the content determination process determines that the content file contains inappropriate content, the content file is deregistered. This prevents a response containing inappropriate content from being output to a user from the large-scale language model.
[0387] Fig. 25 is a diagram showing an example of a response output system according to a modified example. As shown in Fig. 25, the configuration of the response output system according to the modified example is basically the same as that of the eighth embodiment, but the second server 19002 further includes an inappropriate media content database (hereinafter also referred to as "inappropriate DB") 19010. Words (keywords) and the like that are deemed inappropriate for response requests are registered in advance in the inappropriate DB 19010. Note that examples of inappropriate content include excessively violent or sexual expressions.
[0388] The control unit 1110 of the artificial intelligence response output device 10010 performs content judgment processing, for example, for content files registered by a user, such as Image1.jpg and Video31.mov, by referring to the inappropriate DB19010 to determine whether the content file contains inappropriate content.
[0389] Fig. 26 is a flowchart showing an example of a procedure for creating a response request by a modified AI response output device. The procedure for creating a response request by the AI response output device 10010 is basically the same as the procedure for creating a response request shown in Fig. 8, except that step S10 for performing content determination processing is added after step S4. The content determination processing of step S10 will be mainly described below.
[0390] As described above, once the second content file registration process is performed in step S4, the control unit 1110 then performs content determination process in step S10. In the content determination process, the inappropriate DB 19002 is referenced to determine whether the first content file and the second content file contain inappropriate content. Content files that contain inappropriate content are then deregistered. Note that the content determination process in this example is performed for each of the first content file and the second content file, and the process content is basically the same.
[0391] FIG. 27 is a flowchart showing an example of the procedure for content determination processing. First, in step S111, the content file registered in the response request is compared with the inappropriate DB 19002. More specifically, the first content file and the second content file are analyzed to extract keywords that are characteristic points of the first content file and the second content file. The analysis of the first content file and the second content file may be performed using existing image analysis technology, video analysis technology, or the like. The characteristic points of the second content file may also be the results extracted in the search process. Then, the extracted characte...
Claims
1. an input unit; a control unit that generates a response request based on a user input from the input unit, transmits the response request to a large-scale language model, and outputs a response from the large-scale language model in response to the response request, The control unit performing a search process based on search conditions designated by the user, extracting registration candidates for content files to be registered in the response request, and registering a content file selected by the user from the extracted registration candidates in the response request; Response output device.
2. 2. The response output device according to claim 1, A plurality of content files can be registered in the response request, and if a first content file is registered in the response request, The control unit performing the search process using the relevance to the first content file as one of search conditions, and extracting registration candidates for the second content file; Response output device.
3. 3. The response output device according to claim 2, The control unit before performing the search process, confirm with the user whether or not to perform the search process using search criteria including the relevance to the first content file; Response output device.
4. 2. The response output device according to claim 1, a display unit that displays the list of registration candidates, The control unit When the user performs a specific operation other than a file move operation on the content file displayed on the display unit as the registration candidate, it is determined that the content file on which the specific operation has been performed has been selected by the user. Response output device.
5. 5. The response output device according to claim 4, The control unit When it is determined that the content file has been selected by the user, the content of the selected content file is reproduced before registering the selected content file in the response request. Response output device.
6. 5. The response output device according to claim 4, The control unit before registering the content file in the response request, confirming with the user whether or not the content file needs to be registered in the response request; Response output device.
7. 3. The response output device according to claim 2, The control unit When performing the search process, the first content file is analyzed to identify feature points of the first content file, and the registration candidates are extracted based on the degree of relevance of the identified feature points to the first content file. Response output device.
8. 3. The response output device according to claim 2, When the search process targets multiple types of content files, The control unit a content file of a type different from the first content file is preferentially subjected to the search process; Response output device.
9. 9. The response output device according to claim 8, The plurality of types of content files include image files, video files, and music files. Response output device.
10. 3. The response output device according to claim 2, a plurality of storage units in which the content files are stored; The control unit performing the search process preferentially on the storage unit in which the first content file is stored; Response output device.
11. 2. The response output device according to claim 1, a display unit that displays the list of registration candidates, The control unit When the number of the registration candidates displayed on the display unit reaches a preset upper limit, the search process is terminated. Response output device.
12. 3. The response output device according to claim 2, a display unit that displays the list of registration candidates, The control unit As a result of the search process, a list of the registration candidates is displayed in order of the degree of relevance to the first content file. Response output device.
13. 3. The response output device according to claim 2, a display unit that displays the list of registration candidates, The control unit As a result of the search process, among the registration candidates, those highly related to the first content file are displayed in an emphasized manner. Response output device.
14. 3. The response output device according to claim 2, a display unit that displays the list of registration candidates, The control unit displaying a list of the registration candidates in the order of extraction as a result of the search processing; Response output device.
15. 2. The response output device according to claim 1, a display unit that displays the list of registration candidates, The control unit When the search process targets multiple types of content files, classifying the content files by type and displaying them on the display unit; Response output device.
16. 5. The response output device according to claim 4, The control unit After registering the content file selected by the user in the response request, confirming with the user whether or not a content file the user wishes to register is included in the list of registration candidates; Response output device.
17. 5. The response output device according to claim 4, If you want to perform a re-search process with search conditions changed from the previous search process, The control unit displaying the list of registration candidates on the display unit excluding those extracted in the previous search process; Response output device.
18. 2. The response output device according to claim 1, The control unit performing a determination process to determine whether or not a content file to be registered in the response request contains inappropriate content, and if the content file does not contain inappropriate content, transmitting the response request in which the content file is registered to the large-scale language model; Response output device.
19. 19. The response output device according to claim 18, The control unit In the determination process, an explanation response request requesting an explanation in natural language of the content file to be registered in the response request is sent to the large-scale language model, and based on the response content generated by the large-scale language model in response to the explanation response request, it is determined whether or not the content file contains inappropriate content. Response output device.
20. Large-scale language models and a control unit that generates a response request based on a user input from an input unit, transmits the response request to the large-scale language model, and causes the large-scale language model to output a response generated in response to the response request, The control unit performing a search process based on search conditions designated by the user, extracting registration candidates for content files to be registered in the response request, and registering a content file selected by the user from the extracted registration candidates in the response request; Response output system.
Citation Information
Patent Citations
Structural unit for tank construction
JP1977008512A
Cited By
Gateway system, information processing method, and program
JP7860327B1