Response output system
The response output system enhances interaction with users by utilizing a large-scale language model to generate and output suitable responses through a clarification process, addressing the need for improved user interaction in existing technologies.
Patent Information
- Application Number
- JP2024096040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-12-25
AI Technical Summary
Existing technologies using artificial intelligence for response generation do not adequately address the need for a more suitable configuration for providing responses to users.
A response output system comprising an input unit, a control unit, and an output unit, which includes a clarification process to generate a response instruction sentence based on a question sentence, utilizing a large-scale language model to provide more suitable responses.
The system provides more suitable response outputs by enhancing the interaction with users through improved response generation and clarification processes.
Smart Images

Figure 2025187339000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a response output system. [Background technology]
[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special table 2019-528512 publication Summary of the Invention [Problem to be solved by the invention]
[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.
[0005] An object of the present invention is to provide a more suitable response output technique. [Means for solving the problem]
[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problem, and one example thereof may be a response output system including: an input unit into which a question sentence is input by a user; a control unit that generates a response instruction sentence for a large-scale language model based on the question sentence and acquires the response sentence generated by the large-scale language model in response to the response instruction sentence; and an output unit that performs output based on the response sentence acquired by the control unit, wherein the control unit executes a clarification process for the question sentence and generates the response instruction sentence based on the question sentence after the clarification process. [Effects of the Invention]
[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments. [Brief explanation of the drawings]
[0008] [Figure 1A] 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. [Figure 1B] 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 1C] 1 is a diagram showing an example of the operation of an AI response output device and system according to an embodiment of the present invention; [Figure 2A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2C] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 2E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2I]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2J] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2K] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 2L] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3A] 1 is an explanatory diagram illustrating an example of a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3C] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3D] 1 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; [Figure 3E] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3F] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3G] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3H] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 3I] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4A]1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 4B] 1 is an explanatory diagram illustrating an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. [Figure 5A] FIG. 2 is an explanatory diagram of an example of the operation of the artificial intelligence response output device according to an embodiment of the present invention. [Figure 5B] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5C] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 5D] FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 6] FIG. 10 is an explanatory diagram of an example of a response generation process of the AI response output device according to an embodiment of the present invention. [Figure 7] 1 is a diagram illustrating an example of a response output system according to an embodiment of the present invention. [Figure 8A] FIG. 1 is a diagram illustrating an example of a processing flow of a response output system according to an embodiment of the present invention. [Figure 8B] FIG. 2 is a diagram showing an example of a processing flow of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 9] FIG. 10 is a diagram showing a detailed example of a disambiguation processing step in the processing flow of the response output system according to one embodiment of the present invention. [Figure 10A] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 10B] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 10C] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 10D] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 11]FIG. 10 is a diagram showing a modified example of the disambiguation processing step in the processing flow of the response output system according to an embodiment of the present invention. [Figure 12] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 13A] FIG. 1 is a diagram illustrating an example of a processing flow of a response output system according to an embodiment of the present invention. [Figure 13B] FIG. 2 is a diagram showing an example of a processing flow of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 14] 1 is a diagram illustrating an overview of a response output system according to an embodiment of the present invention. [Figure 15] FIG. 10 is a diagram showing a detailed example of a relevance confirmation process and processing steps when there is a relevance in the processing flow of the response output system according to an embodiment of the present invention. [Figure 16] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 17] FIG. 2 is an explanatory diagram of an example of the operation of the response output system according to an embodiment of the present invention. [Figure 18] FIG. 1 is a diagram illustrating an example of a processing flow of a response output system according to an embodiment of the present invention. [Figure 19A] FIG. 1 is a diagram illustrating an example of a processing flow of a response output system according to an embodiment of the present invention. [Figure 19B] FIG. 2 is a diagram showing an example of a processing flow of an artificial intelligence response output device according to an embodiment of the present invention. [Figure 20] FIG. 10 is a diagram illustrating a detailed example of a log information saving operation of the response output system according to an embodiment of the present invention. [Figure 21] 10A and 10B are diagrams illustrating detailed examples of log information selection and readout operations of a response output system according to an embodiment of the present invention. [Figure 22] FIG. 10 is a diagram illustrating a detailed example of a response generation operation based on log information of the response output system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.
[0010] Note that if the AI response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI response output device has an audio output function, it may be referred to as an audio output device. The AI response output device may simply be referred to as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI response output device or the display output of the AI response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.
[0011] Example 1 As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.
[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system in which the AI response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.
[0013] In the example of FIG. 1A, the AI response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Furthermore, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which the pixels emit light themselves. Furthermore, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.
[0014] 1A, the audio output unit 1140 provided in the AI response output device 10010 is configured with a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. By audio input from the microphone 1139 or operation input from the user via an operation input unit (described later), the AI response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.
[0015] The AI response output device 10010 may be provided with a local large-scale language model within the AI response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or an audio output from the audio output unit 1140.
[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.
[0017] Alternatively, the AI response output device 10010 may also include a local large-scale language model, and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.
[0018] The configuration when the AI response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions, or may go via a router or repeater. The AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as one system.
[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
[0020] 1A shows an example in which a display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of FIG. 1A, the prompt display area 10051 displays an icon 10052 representing a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of FIG. 1A, the artificial intelligence response display area 10061 displays an icon 10062 representing an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.
[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment. These large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in these artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called Transformer. Reference 1, for example, has been published as an example of the learning of these models.
[0022] [Reference 1] Long Ouyang, et. al. “Training language models t follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001 and the AI response output device 10010. In this way, instead of performing pre-learning for each application or terminal, replicating the large-scale language model, which is the base model generated by large-scale pre-learning, and using it on individual servers and terminals allows the resources used for learning to be shared, resulting in good resource efficiency.
[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.
[0025] Furthermore, large-scale language models can pre-train natural languages and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. FIG. 1A illustrates a large-scale language model server 20001 having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are AI models generated by large-scale pre-training on natural language and types of information other than natural language text information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. In addition, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.
[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.
[0028] The AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a nonvolatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, and an imaging unit 1180. The AI response output device 10010 may have a large screen, such as a monitor or television.
[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire the user input that serves as the basis for a prompt to the large-scale language model, which is the artificial intelligence.
[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN-type communication method. This allows the AI response output device 10010 to communicate with various servers connected to the Internet 19000.
[0031] The AI response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.
[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.
[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller or the like, and inputs a signal regarding an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. In addition to a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may also be used by, for example, an administrator to operate the AI response output device 10010. By the user's operation input via the operation input unit 1107, the AI response output device 10010 can acquire a user input that serves as the basis for a command sentence (prompt) to a large-scale language model, which is an AI. Note that a modified configuration in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107 is also possible.
[0034] The video signal input unit 1131 is connected to an external video output device and inputs video data. The video signal input unit 1131 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces, etc.
[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.
[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
[0037] The microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.
[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.
[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.
[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 controls video switching, such as which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also control image processing of the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling, which enlarges, reduces, or deforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.
[0042] The nonvolatile memory 1108 stores various data used by the AI response output device 10010. The data stored in the nonvolatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.
[0043] The local LLM processing unit 10028 has a memory capable of holding a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI response output device 10010.
[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also work in cooperation with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, which is a speaker or the like.
[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 of the AI response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and responses are obtained from these large-scale language models. All of this control can be performed by the control unit 1110.
[0046] The storage unit 1170 may also store a fixed response phrase database (which may be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction sentences from the AI response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI response output device 10010.
[0047] Furthermore, in the example of the fixed response phrases in the database shown in FIG. 1C, if multiple fixed response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the fixed response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the fixed response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.
[0048] Next, an example of condition number 5 shown in Fig. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I might not know about that." By responding in this way, the user can be prompted to input again, and the corrected user input can be waited for.
[0049] Next, an example of condition number 6 shown in Fig. 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in Fig. 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the fixed response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.
[0050] The AI response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).
[0051] 1C described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the response template database (response template DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the AI response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0052] In the above explanation, it has been explained that the AI response output device 10010 has a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.
[0053] 1A and 1B, an example has been described in which the AI response output device 10010 includes the display unit 10011. However, the AI response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI may be configured to accept input from a user to the AI via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the input from the user via the voice output unit 1140.
[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference of the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.
[0055] <Example 2> Next, as a second embodiment of the present invention, an example will be described in which the AI response output device 10010 described in the first embodiment is connected to the Internet, and is connected to a server equipped with a large-scale language model AI via the Internet to perform operation. In this embodiment, differences from the first embodiment will be described, and repeated explanations of the same configurations as those of these embodiments will be omitted.
[0056] An example of a connection state between an AI response output device 10010 and a large-scale language model server 19001 according to the second embodiment of the present invention will be described with reference to FIG. 2A. The AI response output device 10010 according to the second embodiment may be called a character conversation device. A system including the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment may be called a character conversation system. A video of a character 19051 is displayed on a display unit 10011 displayed by the AI response output device 10010. The video of the character 19051 is generated by rendering a 3D model of the character in a virtual space.
[0057] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can serve as an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this embodiment may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0058] In the example of FIG. 2A, the audio output unit 1140 provided in the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A, an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication is also acceptable. The communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions. The AI response output device 10010 can communicate with a large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with a second server 19002 different from the large-scale language model server 19001 via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 may be considered as one system.
[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2B. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that Figure 2B does not illustrate communication paths such as the Internet 19000 shown in Figure 2A. Figure 2B also illustrates a user 230 of the artificial intelligence response output device 10010.
[0060] Here, we will explain the sequence of operations of the AI response output device 10010. The AI response output device 10010 loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0061] First, the AI response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is in natural language. Note that the extraction of the text of the words spoken by the user 230 may be performed continuously for all words, or may start when the user utters a word within a predetermined period after a trigger keyword. For example, the trigger keyword may be when the user utters "hello" followed by the character's name. For example, if the name of the character 19051 is "Koto," then "Hello, Koto!" may be the trigger keyword.
[0062] The character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230 and transmits the prompt to the large-scale language model server 19001 using an API. Here, the prompt may be metadata containing information written in a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON. The prompt stores text information in a natural language as a main message. The prompts transmitted from the AI response output device 10010 to the large-scale language model server 19001 include setting prompts that store instructions such as initial settings, and user prompts that reflect instructions from the user. Type identification information identifying whether the prompt is a setting prompt or a user prompt may be stored in a portion of the prompt other than the main message. When the character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230, it creates the user prompt and sends it to the large-scale language model server 19001.
[0063] Next, the large-scale language model of the artificial intelligence of the large-scale language model server 19001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response including natural language text information based on the result. The large-scale language model server 19001 uses an API to send the response to the artificial intelligence response output device 10010. The response stores the natural language text information as a main message. Here, the response may be metadata storing information written in the same format as the instruction (e.g., a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON). If the response uses the same format as the instruction, type identification information may be stored outside the main message to indicate that the information is a different type from the initial setting instruction and the user instruction. For example, information indicating that the response is a response from the large-scale language model may be stored.
[0064] Next, the AI response output device 10010 receives a response from the large-scale language model server 19001 and extracts the natural language text information stored as the main message in the response. The character operation program of the AI response output device 10010 uses speech synthesis technology to generate a natural language voice as a response to the user based on the natural language text information extracted from the response, and outputs the voice from the speaker, i.e., the voice output unit 1140, so that it sounds as if it were the voice of the character 19051. This process may also be referred to as the character's "speaking."
[0065] Conversation examples 1 to 5 in Fig. 2C show specific examples of response voices from character 19051 in response to words from user 230, which are generated by the above-described processing of AI response output device 10010 and large-scale language model server 19001. In this way, user 230 can converse with character 19051 as if it were a real person.
[0066] 2B or a system including the AI response output device 10010, there is no need to install a large-scale language model, which requires a huge amount of data and computational resources for learning, in the AI response output device 10010. In addition, the advanced natural language processing capabilities of the large-scale language model can be utilized via an API, and when a user speaks to a character, a more appropriate response can be given to the user, enabling a more appropriate conversation to be held.
[0067] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2D. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2D shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0068] FIG. 2D also shows the exchange of instructions and responses in chronological order, from the display setting instruction, the first round of user instructions and their responses, to the fourth round of user instructions and their responses.
[0069] As shown in FIG. 2D , the setting directive can be used to initially instruct the large-scale language model of the artificial intelligence of the large-scale language model server 19001, such as the name of the large-scale language model itself, the role to be played, and conversation characteristics. The user's name can also be understood as the initial setting. This allows the large-scale language model to generate responses from the first round onward while respecting the role. Then, to a user who hears the voice of the character 19051 based on the responses from the first round onward, the character 19051 will seem to have the setting and personality of the person described in the setting directive. Furthermore, the large-scale language model server 19001 according to this embodiment is equipped with a memory that stores the content of the conversation until the end of the series of conversations, and is configured to store a series of user directives and their responses and then generate responses. This allows for a conversation such as that shown in FIG. 2D to be realized.
[0070] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2E shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.
[0071] 2E shows an example of a case where, after the series of conversations shown in FIG. 2D has ended, the user 230 speaks to the character 19051 again to start a new conversation. In FIG. 2E, the exchange of instructions and responses is shown in chronological order, from the first round of user instructions and their responses to the third round of user instructions and their responses.
[0072] Here, "termination" of the "continuation of a series of conversations" refers to a process in which, when a predetermined condition is met, the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversation that the large-scale language model server 19001 has maintained while the series of conversations was continuing. An example of the predetermined condition is, for example, when the AI response output device 10010 issues an instruction to the large-scale language model server 19001 to "terminate" the "continuation of a series of conversations." Another example of the predetermined condition is, for example, when a predetermined time or more has passed since the AI response output device 10010 stopped sending instruction statements to the large-scale language model server 19001 regarding the series of conversations (timeout). Another example of the predetermined condition is when, after authentication processing has been performed in the connection between the AI response output device 10010 and the large-scale language model server 19001, the authentication processing is terminated due to factors such as communication disconnection or the AI response output device 10010 being powered off while exchanging the instruction statements and responses.
[0073] Note that when the "continuation of a series of conversations" "ends," the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversations that it had maintained while the series of conversations was continuing. Therefore, even though the conversation shown in FIG. 2E takes place after the series of conversations shown in FIG. 2D, the server's response to the user's instruction is a response with content that does not include any memory of the character's name set in the large-scale language model, the role to be played, the characteristics of the conversation, or the user's name, which were included in the setting instruction shown in FIG. 2D. Similarly, the conversation shown in FIG. 2E is a response with content that does not include any memory of the series of conversations shown in FIG. 2D. In other words, with the "end" of the "continuation of a series of conversations" shown in FIG. 2D, the conversation in FIG. 2E starts from an initialized state of the large-scale language model of the artificial intelligence of the large-scale language model server 19001.
[0074] This causes the user 230 to feel as if the character 19051 has lost its memory or is a completely different person. From the user 230's perspective, the character's response feels very strange, resulting in a feeling of loneliness and disappointment. Such behavior poses a problem in that it is not possible to ensure the consistency of the settings and memories of the character 19051, such as its name, role, conversational characteristics, and personality, displayed on the AI response output device 10010.
[0075] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2F. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2F shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that is the response.
[0076] FIG. 2F illustrates an example of a case in which the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations illustrated in FIG. 2D has ended. Unlike the process illustrated in FIG. 2E, in the process illustrated in FIG. 2F, when starting a new conversation, the AI response output device 10010 transmits a setting instruction as the first instruction to the large-scale language model server 19001. The setting instruction stores the same natural language text as the initial setting instruction of FIG. 2D. This may be referred to as a reset text. The setting instruction is followed by natural language text describing the history of past conversations. This may be referred to as a conversation history text. The AI response output device 10010 may record the history of past conversations as natural language text information in the storage unit 1170 while the series of conversations illustrated in FIG. 2D is continuing, linking the history of the conversations to information on the date and time of the conversations. If there are conversations on different dates, each conversation may be recorded linked to date and time information, and the conversation history may be accumulated. When generating a setting instruction for the first instruction of a later conversation, such as that shown in Figure 2F, the natural language text information of the conversation recorded in storage unit 1170 and the information on the date and time when the conversation took place can be read out and used to generate the setting instruction.
[0077] When natural language text information of past conversation history is used to generate the setting instruction sentence, the format can be determined freely to a certain extent because it is data to be sent to a large-scale language model, but as shown in Figure 2F, it is sufficient to prepare prefixes and suffixes in natural language, such as "I talked about the following on ____ day of ____ month," or "You talked about the following on ____ day of ____ month," and combine these with the natural language text information of the recorded conversation to generate the text of the setting instruction sentence. Also, information on the date and time of the conversation read from storage unit 1170 may be combined with the above-mentioned "____ day of ____ month" portion to form part of the text of the setting instruction sentence.
[0078] Even if user 230 speaks to character 19051 again to start a new conversation after a series of conversations has ended, by performing the above-described generation process and transmission process of the setting instruction sentence in Fig. 2F, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation. This is preferable because it is perceived by the user as ensuring the consistency of the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation.
[0079] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Fig. 2G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2G shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0080] Figure 2G shows an example of a series of conversations shown in Figure 2F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 2G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 2F, so repeated description is omitted.
[0081] As shown in the natural language text of the server response in the table of Figure 2F, by using the setting directive shown in Figure 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 reflects the settings and conversation history of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation. This is more preferable because it allows the user to recognize that the settings and memories of the character, such as the role, name, conversational characteristics, or personality, at the time of the previous conversation are more closely matched. Note that this allows the characters to be viewed as the same from the user's perspective, and may therefore be referred to as a pseudo-identity of the characters from the user's perspective.
[0082] Furthermore, from the user's perspective, memories can be shared with the character, providing a more enjoyable character conversation experience.
[0083] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, FIG. 2H shows an example of the operation in which a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. A character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input input to the operation input unit 1107 or an operation detected by a touch operation input sensor of the display unit 10011.
[0084] In the example of FIG. 2H, in addition to character 19051 (named "Koto") used in the description of FIGS. 2A to 2G, character 19052 (named "Tom") and character 19053 (named "Necco") are shown. Character 19051 (named "Koto") and character 19052 (named "Tom") are human characters, and character 19053 (named "Necco") is a cat character. The display of characters displayed on display unit 10011 can be switched by switching the display on display unit 10011 to display images generated by rendering the characters in different virtual 3D spaces for each character.
[0085] Furthermore, it is preferable that the character operation program executed by control unit 1110 also changes the synthetic voice used for the "speech" of each character when switching the display of the characters displayed on display unit 10011. This can be achieved by storing synthetic voice data of voice tones associated with each character in storage unit 1170 in advance, and performing synthetic voice change processing when switching the display of the character.
[0086] In the example of Fig. 2H, the user 230 is configured to be able to converse with any of the characters. In the AI response output device 10010 of Fig. 2H, each of these characters is set with a different role, name, conversational characteristics, personality, etc. Also, the memories of each character based on the conversation history are managed as different for each character.
[0087] Therefore, the AI response output device 10010 constructs a database shown in FIG. 2I in the storage unit 1170, and manages the character settings and the character conversation history using this database.
[0088] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Fig. 2I. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Fig. 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.
[0089] A character operation program executed by the control unit 1110 of the AI response output device 10010 constructs the database 19200 in, for example, the storage unit 1170. The character ID is an identification number that identifies each of multiple characters that can be displayed on the AI response output device 10010, and may be a natural number or may use alphabets, etc. The name is data of the name of each of multiple characters that can be displayed on the AI response output device 10010.
[0090] The initial setting instruction is text information in a natural language that explains the settings such as the role, name, conversational features, or personality of each of multiple characters that can be displayed on the AI response output device 10010. The initial setting instruction is natural language text information that is the main data of the setting instruction sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the description can be read as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0091] The conversation histories, which continue as conversation histories 1, 2, ..., are records of conversations between each character and the user, and are recorded separately for each character. The conversation histories will be included in the natural language text information, which is the main data of the setting instruction statement sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the conversation histories be readable as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0092] When the character displayed on the display unit 10011 of the AI response output device 10010 is switched, the character operation program executed by the control unit 1110 of the AI response output device 10010 uses the database 19200 of Figure 2I to select and switch the initial setting instruction statement and conversation history used for the natural language text information that is the main data of the setting instruction statement transmitted from the AI response output device 10010 to the large-scale language model server 19001 so as to correspond to the character displayed on the display unit 10011 of the AI response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the conversation history in the conversation history area of the database 19200 of Figure 2I that corresponds to the character displayed on the display unit 10011.
[0093] By using the database 19200 in this manner, the character operation program executed by the control unit 1110 of the AI response output device 10010 establishes a conversation between the user 230 and the character using character utterances that utilize responses from the same AI large-scale language model of the same large-scale language model server 19001. From the user's perspective, the uniqueness of each character's personality and other settings is preserved, and it appears as if each character's unique conversational memories continue. This is more preferable because it appears as if the identity of each character's settings and memories, such as their role, name, conversational characteristics, or personality, from the time of the previous conversation has been more consistently maintained. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.
[0094] Therefore, even when the AI response output device 10010 is configured to switch the character displayed on the display unit 10011 from among multiple character candidates, the operation using the database 19200 described above reduces the sense of discomfort felt by the user when conversing with each character, and the user can share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.
[0095] Note that if the initial setting instructions for multiple characters cannot be edited by the user, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained close to the intentions of the provider of the AI response output device 10010 or the creator of the character content. Alternatively, the initial setting instructions for each character may be edited by the user in response to input from the operation input unit 1107 or the like. In this case, the character's role, name, conversational characteristics, or personality can be set to a preferred setting, allowing the user to converse with a character that they have individually set. In this case, the character's 3D model, its rendered image, and the type of synthesized voice for the character may be replaced accordingly.
[0096] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described with reference to Figure 2J. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service at a lower cost using the character conversation device based on the artificial intelligence response output device 10010 and the character conversation system based on the artificial intelligence response output device 10010 and the large-scale language model server 19001 will be described.
[0097] As explained in Figure 2B, it is extremely resource inefficient to train a large-scale language model at this level of artificial intelligence by limiting it to a specific application. Therefore, it is more resource efficient to perform large-scale training to generate a foundation model that can be applied to a variety of applications, and then use it on various devices via an API (Application Programming Interface). In such cases, providers of large-scale language models often recover the costs incurred in training the large-scale language model from device users as API usage fees. In natural language models, API usage fees are often charged based on the number of tokens, which are units of words that separate sentences, processed.
[0098] Therefore, in the artificial intelligence response output device 10010 of the second embodiment of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010 and a character conversation service using a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.
[0099] For example, by using the processing and configuration of Examples 1 to 3 shown in the table of Figure 2J, it is technically possible to reduce the number of tokens in the natural language text information transmitted between the AI response output device 10010 and the large-scale language model server 19001 using an API.
[0100] Example 1 is an example of a method for reducing the number of tokens in the conversation history text stored in the API setting directive and transmitted, in which the conversation history text is shortened using a document summarization process to reduce the number of tokens. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. The text summarization may be performed at the start of the next conversation, but it is more time-efficient to perform it at the end of the "series of conversations."
[0101] Furthermore, the text summarization process may be requested from the large-scale language model itself of the large-scale language model server 19001. However, in this case, the effect of saving the number of tokens is small. Therefore, for example, if the second server 19002 provides text summarization process for natural language via an API at a lower cost than the large-scale language model of the large-scale language model server 19001, the text summarization process can be requested from the second server 19002 via the API, and the text summary of the conversation history can be stored in a setting instruction statement for the large-scale language model server 19001 and transmitted.
[0102] Furthermore, if it is only text summarization processing, it can be performed on the terminal side, and text summarization can be performed by the control unit 1110 executing a document summarization program loaded in the memory 1109 of the AI response output device 10010. In this case, the effect of saving the number of tokens is high. Furthermore, even if the conversation history becomes long, if an upper limit on the number of characters after summary is specified in the text summarization processing, the upper limit on the text length of the conversation history is determined, so that an upper limit on the number of tokens can be set and tokens can be saved.
[0103] In addition, since the text information of the character's initial settings, such as the character's role, name, conversational characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information of the character's initial setting instructions and reduce the number of tokens of the text information in the conversation history.
[0104] The processing described in Example 1 may be performed by the character movement program executed by the control unit 1110 controlling each unit.
[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored in the API setting instruction and transmitted. For example, the number of tokens is reduced by deleting the oldest conversation history among the conversation histories with characters recorded in the storage unit 1170. Specifying an upper limit on the number of characters in the conversation history determines the upper limit on the length of the conversation history text, thereby setting an upper limit on the number of tokens and enabling token conservation. Alternatively, a predetermined period of the conversation history may be specified and conversation history exceeding that period may be deleted. This also enables token conservation. Note that, in Example 2, the text information for the character's initial setting, such as the character's role, name, conversation characteristics, or personality, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the text information in the character's initial setting instruction and reduce the number of tokens in the text information for the conversation history.
[0106] The processing described in Example 2 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0107] Example 3 is a method of reducing the number of tokens by reducing the frequency of sending setting instruction sentences using an API. Specifically, even after the device is powered on, after the displayed character is switched, or after the video settings and synthetic voice settings of the displayed character are completed, setting instruction sentences are not sent in advance, and only when the control unit 1110 determines that natural language text information included in the user's speech picked up by the microphone 1139 is text information that should use a large-scale language model of artificial intelligence is the setting instruction sentence sent to the large-scale language model server 19001, thereby reducing the frequency of sending setting instruction sentences to the large-scale language model server 19001 and reducing the number of tokens.
[0108] Specifically, for example, after the device is powered on or after an operation input to switch the displayed character is made, character 19051 (named "Koto") is displayed on display unit 10011 as shown in FIG. 2H by display processing of display unit 10011 controlled by a character operation program executed by control unit 1110. At this time, for example, if a synthetic voice for the character 19051 to appear is stored and prepared in storage unit 1170 or the like, a synthetic voice for the character's appearance such as "Good morning, I'm Koto," "Hello, I'm Koto," or "Good evening, I'm Koto" may be output from the speaker, which is audio output unit 1140. At this time, the image of character 19051 has already been set as the image of the character to be displayed on display unit 10011, and the synthetic voice output from the speaker, which is audio output unit 1140, is set to the synthetic voice corresponding to character 19051.
[0109] Here, the inference process of the large-scale language model of the artificial intelligence in the large-scale language model server 19001 already described also takes time as the instruction becomes longer. In particular, if the setting instruction includes text information related to past conversation history, the number of tokens in the instruction increases, resulting in a particularly long inference process time. The setting instruction itself and its response are not output to the user 230. Based on the response to the user instruction following the setting instruction, a synthesized voice is output as the character's "utterance" from the speaker, which is the voice output unit 1140. In this case, it may seem preferable at first glance to send the setting instruction from the AI response output device 10010 to the large-scale language model server 19001 in advance to complete the inference process of the large-scale language model for the setting instruction in advance, as this would result in a faster response in the output of synthesized voice of the character's 19051's "utterance" after the user 230 speaks to the character 19051.
[0110] However, if the setting instruction is transmitted to the large-scale language model server 19001 before the user 230 speaks and the inference processing of the large-scale language model for the setting instruction is completed in advance, for example, the user 230 may turn off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or the user 230 may switch the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. In this case, the number of tokens processed in the inference processing of the large-scale language model by transmitting the setting instruction to the large-scale language model server 19001 in advance becomes the number of processed tokens for which the usage fee is wasted. This hinders the provision of a character conversation device using the AI response output device 10010 and a character conversation service using a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 at a lower cost to users.
[0111] Therefore, after the AI response output device 10010 is powered on (ON) or after an operation input to switch the displayed character is made, it is desirable that the AI response output device 10010, under the control of the character operation program executed by the control unit 1110, sets the image of character 19051 as the image of the character to be displayed on the display unit 10011, and sets the synthetic voice output from the speaker, which is the audio output unit 1140, to the synthetic voice corresponding to character 19051, continue not to send a setting instruction statement to the large-scale language model server 19001 until the user 230 recognizes that he or she is speaking to character 19051.
[0112] Here, the point in time at which it is recognized that user 230 is speaking to character 19051 may be, for example, the point in time at which the trigger keyword described in Fig. 2B is detected, or the point in time at which the text of the words spoken by user 230 is extracted. In this way, the number of processing tokens that waste usage fees can be reduced, and a character conversation service by a character conversation device using artificial intelligence response output device 10010 or a character conversation system using artificial intelligence response output device 10010 and large-scale language model server 19001 can be provided to users at a lower cost.
[0113] Furthermore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051 has passed, it is desirable to continue not sending setting instructions to the large-scale language model server 19001, for example, if the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Specifically, examples of preset keywords include keywords such as "try jumping" and "try dancing," which are keywords by which the user 230 requests the character 19051 to react, such as by animating the character 19051 to move or emitting synthetic voice. In this case, the character operation program executed by the control unit 1110 reads out motion data, animation video, and / or synthetic voice data corresponding to the reaction stored in the storage unit, corresponding to the character 19051, and / or synthetic voice data corresponding to the reaction, and uses these data to generate video to be displayed on the display unit 10011 and output synthetic voice from the speaker, which is the audio output unit 1140.
[0114] Such processing does not necessarily require inference processing of the large-scale language model of the large-scale language model server 19001. After the processing, if the user 230 turns off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or if the user 230 switches the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, if the setting instruction statement is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be the number of processing tokens that unnecessarily consumes the usage fee.
[0115] Therefore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051, it is desirable to continue not sending the setting instruction statement to the large-scale language model server 19001 until it is determined whether or not the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Only when it is determined that inference processing of a large-scale language model is required is it desirable to send the setting instruction statement to the large-scale language model server 19001 and proceed with the inference processing of the large-scale language model.
[0116] The processing described in Example 3 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0117] According to the method for reducing (saving) the number of processing tokens for a large-scale language model using the examples of Figure 2J described above, a character conversation device using the AI response output device 10010 and a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.
[0118] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2K. The example of FIG. 2K shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which has been described in each of FIGS. 2A to 2J, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063, which is a response from the large-scale language model, together with an image of a character 19051 on the display unit 10011. The text 10063, which is a response from the large-scale language model, may be displayed superimposed on top of the image of the character 19051, as shown in FIG. 2K. Alternatively, the text 10063, which is a response from the large-scale language model, may be displayed together with the image of the character 19051 without being superimposed on the image of the character 19051.
[0119] The display in Figure 2K is an example, but for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by audio.
[0120] Therefore, in this case, the control unit 1110 may perform control to start a display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, as shown in Fig. 2K. In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more suitably use the character conversation device (artificial intelligence response output device 10010). Note that the user 230 may be configured to manually switch ON / OFF the display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011.
[0121] Next, an example of a fixed response phrase database (fixed response phrase DB) in the character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in FIGS. 2H and 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition numbers and condition contents are the same as those in FIG. 1C. In the example of FIG. 2L, individual fixed response phrases are set for each of the multiple characters in response to these conditions. For example, fixed response phrases for each condition are stored for each of the three characters, Character 1: Koto, Character 2: Tom, and Character 3: Necco, as described in FIGS. 2H and 2I. The output control of fixed response phrases is the same as that in FIG. 1C, and therefore will not be repeated.
[0122] In the example of FIG. 2L, the control unit 1110 selects a corresponding fixed response phrase from a fixed response phrase database (fixed response phrase DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses the selected fixed response phrase for output control as a response uttered by the character. For example, in the example of the fixed response phrase database (fixed response phrase DB) in FIG. 2L, even under the same conditions, the fixed response phrases are changed to expressions or contents that correspond to the individuality of the characters. This allows the character conversation device (artificial intelligence response output device 10010) to provide the user with conversations that correspond to the individuality of the displayed characters. The user can feel that each character has a more consistent personality. This makes it possible to realize a character conversation device (artificial intelligence response output device 10010) that gives multiple characters a more realistic presence.
[0123] The fixed response phrase database (fixed response phrase DB) of FIG. 2L described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 2L may also be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in each server. In this way, even if the AI response output device 10010 does not have a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0124] The character conversation device and character conversation system according to the second embodiment described above can reduce the sense of discomfort felt by the user from the conversation with the character displayed on the AI response output device 10010. Furthermore, the character conversation device and character conversation system according to the second embodiment can provide the character conversation service to the user at a lower cost.
[0125] In the above description of the second embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 19001 is used as the large-scale language model. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B, and the large-scale language model held by the local LLM processing unit 10028 may be used instead of the large-scale language model held by the large-scale language model server 19001. In this case, in the above description of the second embodiment, the large-scale language model held by the large-scale language model server 19001 may be read as the large-scale language model held by the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).
[0126] In this case, too, it is possible to further reduce the sense of discomfort felt by the user from conversations with characters displayed on the AI response output device 10010. Note that when a large-scale language model held by the local LLM processing unit 10028 is used instead of the large-scale language model held by the large-scale language model server 19001, there is less need to consider usage fees according to the number of processed tokens, but by reducing the number of processed tokens even for the large-scale language model held by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power required for inference. In this case, it is possible to provide users with a character conversation service that consumes less power.
[0127] In the above description of the second embodiment, an example has been described in which a conversation history with a character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). Alternatively, a conversation history with a character may be recorded and stored in a second server 19002 or another cloud server connected to the Internet 19000. In this case, when a user and a character start a new conversation, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or another cloud server, acquires (downloads) a past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create an instruction for a large-scale language model. Specific methods for using past conversation history to create an instruction for a large-scale language model are as described in the figures of the second embodiment, and therefore repeated description will be omitted.
[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character's conversation history up to a predetermined point in time, such as every time a conversation between the user and the character takes place or when the conversation between the user and the character ends, to the second server 19002 or another cloud server. That is, the character conversation device (artificial intelligence response output device 10010) may upload the conversation history with the character to the second server 19002 or another cloud server at a predetermined timing, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or another cloud server and use it to generate an instruction sentence for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) that the user used the previous day and the character conversation device (artificial intelligence response output device 10010) that the user is about to use are different individual devices and can display the same character, and when the user has multiple conversations with the same character between the different individual devices at different times, it is possible to realize a conversation in which the character's memories are pseudo-carried over from the previous conversation, which is more convenient for the user.
[0129] The process described above in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with a character to the second server 19002 or another cloud server, thereby pseudo-taking over the character's memory, is also effective when handling the database 19200 including the conversation histories of multiple characters described in Figures 2H and 2I. In other words, if the database 19200 described in Figure 2I is configured to be uploaded and downloaded to the second server 19002 or another cloud server, not only for one character but for multiple characters, and between different individual devices, when a user has multiple conversations with each of the multiple characters at different times, it is possible to realize a conversation in which the memory of each character is pseudo-taken over from the previous conversation, which is more convenient for the user.
[0130] Example 3 Next, the third embodiment of the present invention is an improvement of the character conversation device (artificial intelligence response output device 10010) and the character conversation system explained in the drawings of the second embodiment. In this embodiment, differences from the second embodiment will be explained, and repeated explanations of the same configurations as those of the second embodiment will be omitted.
[0131] As in Example 2, the character in Example 3 can provide the user with the service of a large-scale language model, which is an artificial intelligence, and can be of assistance to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this example may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0132] An example of a character conversation device and a character conversation system according to a third embodiment of the present invention will be described with reference to Fig. 3A. The character conversation system of the third embodiment is provided with a large-scale language model server 20001 instead of the large-scale language model server 19001 in Fig. 2A, and is connected to the Internet 19000.
[0133] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence, and is a multimodal large-scale language model artificial intelligence that can process not only the natural language text information that the large-scale language model server 19001 could process, but also types of information other than natural language text information.
[0134] Moreover, the AI response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (AI response output device 10010) of the second embodiment, as an example.
[0135] In the third embodiment as well, the AI response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the Internet 19000 using an API.
[0136] The character conversation system of the third embodiment includes a mobile information processing terminal 20010 used by a user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.
[0137] 3B, an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011, which is a touch operation input panel, a control unit 20012, an external power supply input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.
[0138] The display panel 20011 is equipped with a touch operation input sensor and can accept touch operation input by the finger of the user 230. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel. The display panel 20011 may also be referred to as a display unit.
[0139] The communication unit 20020 may be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and a memory, and the control unit controls the display panel 20011 and the communication unit 20020. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000. This allows the mobile information processing terminal 20010 to communicate with various servers connected to the Internet 19000.
[0140] The power supply 20014 converts AC current input from the outside via the external power supply input interface 20013 into DC current and supplies the required DC current to each unit of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Furthermore, the secondary battery 20015 supplies power to each unit requiring power via the external power supply input interface 20013 when power is not supplied from the outside.
[0141] The video signal input unit 20023 is connected to an external video output device and inputs video data. The video signal input unit 20023 can be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be various USB interfaces.
[0142] The audio signal input unit 20024 is connected to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 20024 may also be various USB interfaces. In the case of an HDMI-standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface in which a terminal and a cable are integrated.
[0143] The audio output unit 20021 can output audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 can also output audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured with a speaker. The audio output unit 20021 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard.
[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 20012, which will be described later, may perform voice recognition processing on the generated audio signal to obtain text information from the audio signal.
[0145] The imaging unit 20025 is a camera having an image sensor. A camera may be provided on the front side of the mobile information processing terminal 20010, on the display panel 20011 side, or on the back side of the display panel 20011 side. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.
[0146] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 20016 before product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices, external servers, etc., via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may also be output to external devices, external servers, etc., via the communication unit 20020.
[0147] The video control unit 20017 performs various controls related to the video signal input to the display panel 20011. The video control unit 20017 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be referred to as a video processing unit or image processing unit. The video control unit 20017 controls video switching, such as which video signal to input to the display panel 20011 between the video signal to be stored in the memory 20026 and the video signal (video data) input to the video signal input unit 20023. The video control unit 20017 may also control image processing of the video signal input from the video signal input unit 20023 and the video signal to be stored in the memory 20026. Examples of image processing include scaling, which enlarges, reduces, or transforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0148] The attitude sensor 20018 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each unit connected thereto.
[0149] The nonvolatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the nonvolatile memory 20027 includes, for example, data for various operations to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 20026 stores video data to be displayed on the display panel 20011, data for controlling the device, etc. The control unit 20012 may read various software from the storage unit 20016 and expand and store it in the memory 20026.
[0150] The control unit 20012 controls the operation of each unit connected to it. The control unit 20012 may also work in conjunction with a program stored in the memory 20026 to perform arithmetic processing based on information acquired from each unit within the mobile information processing terminal 20010.
[0151] Next, an example of the operation of the character conversation device (AI response output device 10010) according to the third embodiment of the present invention will be described with reference to Figure 3C. This can also be said to be an example of the operation of a character conversation system including the AI response output device 10010 and the large-scale language model server 20001. In the third embodiment as well, the character conversation device (AI response output device 10010) loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0152] In the second embodiment, the actions performed by the user 230 with respect to the character conversation device (artificial intelligence response output device 10010) were mainly calls made by the user 230 using his / her voice. The character conversation device (artificial intelligence response output device 10010) of the second embodiment performed a series of operations starting from the process of collecting the voice of the user 230 with a microphone. In contrast, the character conversation device (artificial intelligence response output device 10010) of the third embodiment is also capable of performing the series of operations performed by the character conversation device (artificial intelligence response output device 10010) starting from the process of collecting the voice of the user 230 with a microphone, as described in the second embodiment. In addition, in the character conversation device (artificial intelligence response output device 10010) of the third embodiment, the user 230 can perform an action with respect to the character conversation device (artificial intelligence response output device 10010) by user operation via the operation input unit 1107 of FIG. 1B. Here, examples of the operation input unit 1107 of FIG. 1B include a mouse, a keyboard, and a touch panel.
[0153] Furthermore, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by performing a touch operation by the user that can be detected by the touch operation input sensor of the display unit 10011 of Figure 1B.
[0154] In addition, the user 230 can operate the mobile information processing terminal 20010 and communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting the user's 230 operation input to the character conversation device (artificial intelligence response output device 10010).
[0155] Alternatively, an information storage image such as a two-dimensional code storing information that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as a two-dimensional code captured by the imaging unit 1180, and obtain the information. Alternatively, an image that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the image of the display may be captured by the imaging unit 1180 of FIG. 1B that the character conversation device (artificial intelligence response output device 10010) has. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180, and obtain the results of the image recognition processing.
[0156] As described above, the character conversation device (artificial intelligence response output device 10010) of the third embodiment has a greater variety of actions that the user 230 can take toward the character conversation device (artificial intelligence response output device 10010) than the character conversation device (artificial intelligence response output device 10010) described in the second embodiment. As a result, the character conversation device (artificial intelligence response output device 10010) of the third embodiment can acquire the results of actions taken by the user 230 other than the user's voice, and generate an instruction sentence (prompt) to be sent to the large-scale language model server 20001 based on the results. As a result, the instruction sentence to be sent to the large-scale language model server 20001 can more preferably include information of a type other than text information in a natural language extracted from the user's voice. Examples of information of a type other than text information in a natural language extracted from the user's voice include images, videos, and sounds.
[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment uses an API to send an instruction to the large-scale language model server 20001. In this embodiment, the instruction may be metadata containing information written in a notation using tags, such as the markup format of a markup language, a notation using predetermined symbols, such as Markdown, or an object notation of a predetermined script, such as JSON. In this embodiment, the instruction may be classified into two types: a setting instruction that stores instructions, such as initial settings, and a user instruction that reflects instructions from a user. Type identification information identifying whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. In this case, the instruction includes text information in a natural language as the main message. Furthermore, in this embodiment, the main message of the instruction may include, in addition to the text information in natural language, a non-natural language information source, such as an image, video, or audio, as information of a type other than the natural language text information. A specific method for including a non-natural language information source in the instruction will be described later.
[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. The large-scale language model server 20001 receives an instruction from a character conversation device (artificial intelligence response output device 10010). Based on the instruction, the multimodal large-scale language model performs inference and generates a response including natural language text information that is the result of the inference. Here, because the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.
[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or sounds stored as the main message in the response. The character operation program of the character conversation device (artificial intelligence response output device 10010) may use a voice synthesis technology to generate a natural language voice as a reply to the user based on the natural language text information extracted from the response, and output the voice from the voice output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.
[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display characters in a natural language that serve as a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on text information in a natural language extracted from the above-mentioned response. At this time, the characters may be displayed together with the character 19051, may be displayed superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the information of the image of the non-natural language information source extracted from the above-mentioned response. At this time, the image may be displayed together with the character 19051, may be superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the video information of the non-natural language information source extracted from the above-mentioned response. At this time, the video may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed in place of the video of the character 19051. These specific processes may be executed by the video control unit 1160.
[0163] In addition, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output a voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140, which is a speaker.
[0164] As described above, the character conversation device (artificial intelligence response output device 10010) of FIG. 3C or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 does not require the large-scale language model itself, which requires vast amounts of data and computational resources for learning, to be installed in the character conversation device (artificial intelligence response output device 10010). Furthermore, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via an API. In response to a user's action toward a character, a response based on a non-natural language information source can be provided in addition to a response based on natural language text, enabling a more appropriate conversation to be held.
[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3D . This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3D shows examples of non-natural language information sources, such as natural language text and images, of the main message of an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001, and examples of non-natural language information sources, such as natural language text and images, of the main message of a server response that is the response. In this embodiment, the non-natural language information source can be an image, video, audio, or the like, but FIG. 3D shows an example of an image as the non-natural language information source.
[0166] 3D also shows an exchange of instructions and responses in chronological order, from a setting instruction, a first round of user instructions and their responses, to a second round of user instructions and their responses. The instructions and responses shown in FIG. 3D include non-natural language information source 20061 and non-natural language information source 20062, which were not shown in FIG. 2D of Example 2. In the example of FIG. 3D, both non-natural language information source 20061 and non-natural language information source 20062 are images.
[0167] Here, for ease of explanation, FIG. 3D shows an image of the non-natural language information source 20061 pasted into the instruction. However, there are multiple methods for transmitting or specifying data from the non-natural language information source 20061 in an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these multiple methods or switch between them. An example of each method will be described below.
[0168] The first method for transmitting or specifying non-natural language information source data in a directive is used, for example, when the non-natural language information source to be specified is a non-natural language information source located in a location such as a server connected to a network such as the Internet. A specific example of the first method is to use information such as tags and symbols in the directive to specify a non-natural language information source file located on a network such as the Internet by using location information (such as a URL) on the network such as the Internet and a file name.
[0169] For example, it is a tag that specifies an image in a markup language. <img src=""****”"> You can also specify an image on a network such as the Internet by writing the location information and file name information of the image file in the **** part using the tag. <video src=""****”">You can also specify a video file on a network such as the Internet by writing the location information and file name information of the video file in the **** part using the tag. <audio src=""****”">By using the above and entering the location information and file name information of the audio file in the **** part, audio that exists on a network such as the Internet can be specified. Furthermore, if the notation is JSON, an image that exists on a network such as the Internet can be specified by preparing a key such as img_src and entering the location information and file name information of the image file as the value. For video files and audio files, it is sufficient to prepare the respective keys and values. This specific example of a format is just one example, and other unique formats may also be used. In either case, information that specifies the location information and file name information of the non-natural language information source file can be stored in the directive.
[0170] When information specifying the location information and file name information of a non-natural language information source file is stored in a directive, as in the first method, the directive itself does not need to store the data of the non-natural language information source file itself. Therefore, the amount of data in the directive can be reduced. In the first method, the large-scale language model server 20001 that receives a directive specifying non-natural language information source data simply uses the location information and file name information of the non-natural language information source file stored in the directive to obtain the non-natural language information source file located in a location such as a server connected to a network such as the Internet.
[0171] Here, how location information and file name information are input when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence using the first method will be described. In FIG. 3C, it has been explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are increased beyond the voice of the user 230 compared to the second embodiment. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, and the like, by user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in FIG. 1B.
[0172] Furthermore, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a web browser program and display a GUI of the web browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the web browser program may be accepted by a user operation via the operation input unit 1107 (e.g., a mouse, keyboard, or touch panel) or a user touch operation detectable by a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, video, or audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction. In this case, the web browser program may acquire location information and file name information of the non-natural language information source data and pass them to the character operation program.
[0173] Furthermore, user 230 may operate mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from mobile information processing terminal 20010, thereby inputting location information such as a URL for specifying non-natural language information source data to the character conversation device (artificial intelligence response output device 10010). Alternatively, location information such as a URL for specifying non-natural language information source data, file name information, and the like may be input in a manner such as displaying an information storage image such as a two-dimensional code on display panel 20011 of mobile information processing terminal 20010, performing image recognition processing on an image captured by imaging unit 1180 of character conversation device (artificial intelligence response output device 10010), and acquiring the results of the image recognition processing, as described in FIG.
[0174] Note that the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file is previously present in a location such as a server connected to a network such as the Internet. For example, if non-natural language information source data such as images, videos, or audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) is to be included in the instruction statement, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to the second server 19002 via the Internet 19000, and include in the instruction statement the location information on the Internet (such as a so-called URL) and file name of the non-natural language information source data on the uploaded second server 19002. In this case, the second server 19002 functions as a so-called intermediate server.
[0175] Similarly, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in an instruction statement, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the location information on the Internet (such as a URL) and the file name of the non-natural language information source data on the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include the acquired location information on the Internet (such as a URL) and the file name of the non-natural language information source data uploaded to the second server 19002 in the instruction statement.
[0176] Furthermore, a media server may be constructed within the character conversation device (artificial intelligence response output device 10010) so that the character operation program of the character conversation device (artificial intelligence response output device 10010) can cooperate with memory 1109 and storage unit 1170 to be accessible from other servers via the Internet 19000. In this case, when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction statement using the first method, the character conversation device (artificial intelligence response output device 10010) may store location information on the Internet (such as a URL) indicating the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data in the instruction statement.
[0177] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction sentence is, for example, a method in which the non-natural language information source data itself is simply stored (attached) to the instruction sentence (prompt) and transmitted. Generally, non-natural language information source data such as images, videos, and audio has a larger data volume than text information in natural language. Therefore, in this case, the data volume of the instruction sentence (prompt) itself is larger than that in the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) temporarily stores the non-natural language information source data to be stored (attached) to the instruction sentence (prompt) in the memory 1109, and when transmitting the instruction sentence (prompt), stores (attaches) the non-natural language information source data in the memory 1109 via the communication unit 1132 and outputs the data to the large-scale language model server 20001. The non-natural language information source data itself that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read from the storage unit 1170 and stored in memory 1109.
[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data using instruction sentences.
[0179] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. Therefore, in the first pass of the user instruction shown in the example of Figure 3D, it can obtain images of the swimming pool and poolside, which are non-natural language information source 20061, and text information in natural language, and output the text information in natural language as a response to the first pass of the user instruction as an inference result, as shown in the figure.
[0180] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example of the response to the second round of user instructions in Figure 3D, the large-scale language model server 20001 can include in the response a non-natural language information source 20062 generated by inference from the multimodal large-scale language model and send it to the character conversation device (artificial intelligence response output device 10010). In Figure 3D, the non-natural language information source 20062 is an example of an image in which a circle image is added to the image of the swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 3D, but may be video or audio.
[0181] When the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, a method similar to the first method or the second method in which the character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence can be used.
[0182] Specifically, as a method similar to the first method described above, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The non-natural language information source 20062, such as an image, video, or audio, itself may be stored in the large-scale language model server 20001, or the non-natural language information source 20062 may be transferred to and stored in a second server 19002 functioning as an intermediate server. In either case, the large-scale language model server 20001 may store, in a response instruction, information specifying the location information and file name information of the non-natural language information source file. The character conversation device (artificial intelligence response output device 10010) that receives the response may access the large-scale language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction to acquire the non-natural language information source 20062.
[0183] Specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction sentence and use it for various outputs to the user 230.
[0184] According to the operation of the character conversation device (artificial intelligence response output device 10010) and the character conversation system of the third embodiment described above with reference to Fig. 3D, instructions and responses are sent and received between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230 to realize a conversation using non-natural language information such as images, videos, and sounds. This makes it possible to realize a more advanced and natural conversation as shown in each message in Fig. 3D.
[0185] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3E shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0186] 3E shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in FIG. 3D has ended. In the example of FIG. 3E, processing using the conversation history as described in FIGS. 2F, 2G, and 2I of Example 2 is not performed. Therefore, like FIG. 2E of Example 2, FIG. 3E shows a response with content that does not remember at all the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, the user's name, the conversation history, and the like that were included in the setting directive.
[0187] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3F. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3F shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response.
[0188] Figure 3F shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in Figure 3D has ended. Here, Figure 3F shows an example in which the method of storing a message explaining the history of past conversations in the setting instruction sentence, which was explained in Figure 2F of the second embodiment, is also applied to the character conversation device (artificial intelligence response output device 10010) of the third embodiment. Specifically, in Figure 3F, the message that is the content of the setting instruction sentence in Figure 3D is stored as a reset message, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.
[0189] The large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. Therefore, non-natural language information source data may have been transmitted or specified in past instructions and responses. Therefore, in the example of Figure 3F, the conversation history message reflects not only the natural language text information in past instructions and responses, but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method for transmitting or specifying non-natural language information source data in the instruction in Figure 3F is similar to the transmission or specification of non-natural language information source data described in Figure 3D, and therefore a repeated description will be omitted.
[0190] In the example of Figure 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself in the instruction statement, and storing (attaching) the non-natural language information source data in the instruction statement without storing (attaching) the non-natural language information source data in the instruction statement. This also applies to the instruction statement of Figure 3F.
[0191] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Fig. 3G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Fig. 3G shows an example of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0192] Figure 3G shows an example of a series of conversations shown in Figure 3F, from the first user instruction and its response following the first setting instruction to the third user instruction and its response. Figure 3G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 3F, so repeated description is omitted.
[0193] As described above, even when using the large-scale language model server 20001 having a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information according to the third embodiment, even when the user 230 speaks to the character 19051 again to start a new conversation after a series of conversations has ended, if the generation and transmission processes of the setting instruction sentence in Fig. 3F are performed, the response to the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation, as shown in Fig. 3G. This is preferable because it allows the user to recognize that the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation are more closely matched.
[0194] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, FIG. 3H uses the example described in FIG. 2H of the second embodiment for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.
[0195] Furthermore, database 20200 for managing character settings and character conversation history, shown in Figure 3H, has the same format as database 19200 shown in Figure 2I of Example 2, and in Figure 3H, only the differences from database 19200 shown in Figure 2I will be explained. Also, the content of the character "Koto" in the database will be explained, and the content of other characters will be omitted.
[0196] As described above, the large-scale language model server 20001 of the third embodiment is a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information. Therefore, both the instruction statements from the character conversation device (artificial intelligence response output device 10010) and the responses from the large-scale language model server 20001 include not only natural language text information but also the transmission or specification of non-natural language information source data. Therefore, the database 20200 shown in FIG. 3H records, in the conversation history data, not only the natural language text information included in these instruction statements and responses, but also the transmission or specification of non-natural language information source data. The specific method for transmitting or specifying non-natural language information source data in the conversation history recording is the same as the transmission or specification of non-natural language information source data described in FIG. 3D, and therefore a repeated description will be omitted.
[0197] In the example of FIG. 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself to the instruction statement and not storing (attaching) the non-natural language information source data to the instruction statement. This is also true for the conversation history of FIG. 3H. However, in the conversation history of FIG. 3H, if the method of specifying non-natural language information source data is to specify location information and file name information of a non-natural language information source file on a server on a network such as the Internet (such as the second server 19002 functioning as an intermediate server or another cloud server), there is a possibility that the non-natural language information source file on the server may be deleted over a long period of time in the conversation history. In this case, even if the location information and file name information are used, the non-natural language information source file cannot be retrieved at a later date, and information in the conversation record may be lost.
[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts the instruction sentence and the response message into a conversation history and records them, it can use the location information and file name information to obtain the non-natural language information source file itself specified in the instruction sentence and the response from a server on the network and store it in the storage unit 1170. Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) can rewrite the location information and file name of the non-natural language information source file to Internet location information (such as a URL) indicating the location of the non-natural language information source file within the media server built within the character conversation device (artificial intelligence response output device 10010), and then record the non-natural language information source in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself deletes the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the conversation record information, which is more suitable for preserving the conversation record.
[0199] By using the database of Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between multiple character candidates to display on the display unit 10011, the user will feel less uncomfortable when conversing with each character, and will be able to share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.This effect can be achieved as shown in Figure 2I of Example 2, and can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information.
[0200] In addition, in the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, the large-scale language model server 20001 uses a multimodal large-scale language model artificial intelligence that can process not only natural language text information but also non-natural language information other than natural language text information.
[0201] Here, an API is used for communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001. In a multimodal large-scale language model, the API usage fee may be charged according to the number of processed natural language text information, which are units of words that separate sentences called tokens, as well as the amount of data from non-natural language information sources.
[0202] Therefore, in order to provide users with a character conversation service by the character conversation system according to this embodiment at a lower cost, the following modified example may be used.
[0203] In a first variation, the database conversation history record of FIG. 3H also records the transmission or specification of non-natural language information source data. However, the character and the user exchange conversations about the non-natural language information source data using text information in natural language, and the content of the conversations is recorded as text information in natural language. Even if the transmission or specification of the non-natural language information source data is omitted from the database conversation history record of FIG. 3H, the conversation itself about the non-natural language information source data will still be recorded as text information in natural language to some extent. Therefore, if a certain amount of information reduction is acceptable, the transmission or specification of the non-natural language information source data may be omitted from the database conversation history record of FIG. 3H. In this case, the transmission or specification of the non-natural language information source data is also omitted from the conversation history message of the setting instruction sentence of FIG. 3F. This reduces the amount of data from non-natural language information sources communicated using the API.
[0204] Next, as a second modification, in the recording of the conversation history of the database in FIG. 3H, natural language text information explaining the content of the non-natural language information source data is recorded instead of transmitting the non-natural language information source data or recording specified information. The natural language text information explaining the content of the non-natural language information source data may be obtained, for example, by starting a conversation between the large-scale language model of the large-scale language model server 20001 and a character conversation device (artificial intelligence response output device 10010) separately from the conversation as a character, and having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the natural language text information may be obtained by having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit through a conversation with another large-scale language model of another server that is available at a lower cost than the large-scale language model of the large-scale language model server 20001. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be used as natural language text information explaining the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is a tag in a markup language. <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">This is the text information written in the **** part of the above.
[0205] Furthermore, in the case of JSON notation, in an object that is stored in association with the location information and file name information of the non-natural language information source data, which are keys and values indicating the location information of the non-natural language information source data, a key corresponding to the alternative text can be further linked and stored as a value that is the alternative text data itself.
[0206] In this case, too, the transmission of the non-natural language information source data or the recording of the specified information can be omitted in the recording of the conversation history in the database of Fig. 3H, and the transmission of the non-natural language information source data or the specified information is also omitted from the conversation history message of the setting instruction sentence of Fig. 3F, thereby reducing the amount of data from the non-natural language information source communicated using the API.
[0207] Next, a third variation is an example in which, at the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data is not stored in the user instructions, but is replaced with natural language text information explaining the content of the non-natural language information source data. For example, in the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data 20061 may be replaced with the user instructions, and a description such as "This image is of a swimming pool with seats and parasols by the pool. There is water in the swimming pool. There are drinks on the table next to the seats" may be stored as natural language text information. In this case, the description may be obtained by having another large-scale language model on another server, which is more inexpensive to use than the large-scale language model on the large-scale language model server 20001, explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the description may be obtained from a server of various other services that can obtain summaries and descriptions of non-natural language information source data such as images, videos, and audio. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be natural language text information that explains the content of the non-natural language information source data.
[0208] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3I. The example of FIG. 3I shows an example of displaying a response from a large-scale language model to an instruction sentence from a user, which is described in each of FIGS. 3A to 3H, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this is an example of displaying text 10063 of natural language information source data, image 10064 of non-natural language information source data, and / or video 10065 of non-natural language information source data, which are responses from the large-scale language model, together with a video of a character 19051, on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed superimposed on top of the video of the character 19051, as shown in FIG. 3I.
[0209] Furthermore, text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on the video of character 19051. The display in FIG. 3I is an example, but for example, if user 230 adjusts the volume of the audio output of audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating operation input unit 1107 or the touch operation input sensor of display unit 10011, user 230 will not be able to hear the response from the large-scale language model by audio. Therefore, in this case, control unit 1110 may control to start a display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in FIG. 3I.
[0210] In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more preferably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the image of the character 19051, may be manually switched on / off by the user 230 operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. According to the display example of FIG. 3I, the character conversation device (artificial intelligence response output device 10010) that supports multimodal display can more preferably output a response from the large-scale language model.
[0211] According to the character conversation device and character conversation system of Example 3 described above, in addition to the effects of the character conversation device and character conversation system of Example 2, it is possible to provide users with a more advanced conversation experience that includes information in non-natural languages in addition to information in natural languages by using a multimodal large-scale language model. Furthermore, according to the character conversation device and character conversation system of Example 3, it is possible to provide users with a character conversation service at a lower cost.
[0212] In the above description of the third embodiment, an example has been described in which the large-scale language model held by the large-scale language model server 20001 is used as the large-scale language model. In contrast to this, the character conversation device (artificial intelligence response output device 10010) may be equipped with the local LLM processing unit 10028 shown in FIG. 1B and may use the multimodal large-scale language model held by the local LLM processing unit 10028. In this case, the multimodal large-scale language model held by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model held by the large-scale language model server 20001.
[0213] In this case, in the above description of the third embodiment, the multimodal large-scale language model held by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model held by the local LLM processing unit 10028 of the character conversation device (the AI response output device 10010). In this case, too, a more sophisticated conversation experience that includes non-natural language information in addition to natural language information can be provided to the user using the multimodal large-scale language model. Note that, when the multimodal large-scale language model held by the local LLM processing unit 10028 is used instead of the multimodal large-scale language model held by the large-scale language model server 20001, there is less need to consider usage fees based on the number of processed tokens and the amount of data in the non-natural language information source. However, even with the multimodal large-scale language model held by the local LLM processing unit 10028, the number of processed tokens and the amount of data in the non-natural language information source can be reduced, thereby reducing the consumption of resources such as power required for inference. In this case, a character conversation service that consumes less power can be provided to the user.
[0214] The configuration described in Example 2 for uploading and downloading the conversation history with a character or the database data including the conversation history with a character to the second server 19002 or another cloud server can also be used in the example using a multimodal large-scale language model described in Example 3. In this case, too, when a user has multiple conversations with each of the multiple characters at different times between different individual devices for one character or multiple characters, it is possible to realize a conversation in which the memory of each character is pseudo-carried over from the previous conversation, which is more convenient for the user.
[0215] Example 4 Next, Example 4 of the present invention is an improvement of the AI response output device 10010, the character conversation device, or these systems described in the drawings of Example 2 or Example 3. In this example, differences from Example 2 or Example 3 will be described, and repeated explanations of the same configurations as those examples will be omitted.
[0216] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0217] An example of operation using a database in a character conversation device (artificial intelligence response output device 10010) according to a fourth embodiment of the present invention will be described with reference to Fig. 4A. The database according to the fourth embodiment shown in Fig. 4A is an extension of the database described with reference to Fig. 2I or Fig. 3I. Specifically, the database shown in Fig. 4A assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores in the database initial setting instructions and conversation histories corresponding to each user and character.
[0218] In the example of Fig. 4A, for user 1 with a user ID of 1, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, the initial setting command sentences and conversation histories for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored.
[0219] These initial setting instruction sentences and conversation history data are stored as separate data in different areas for each combination of user and character. For the sake of explanation, in Fig. 4A, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (AI response output device 10010) uses the initial setting instruction sentences and conversation history stored in different areas for each combination of user and character based on the user currently using (logging in to) the character conversation device (AI response output device 10010) or its system, thereby making it possible to more appropriately maintain the consistency of the character's personality and the continuity of memory for each different user.
[0220] Specifically, consider a situation in which User 1 has previously had a conversation with character Tom using a character conversation device (AI response output device 10010), and User 2 is unaware of that conversation, and then User 2 subsequently has a conversation with character Tom. In this case, if AI response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify users, the response output from AI response output device 10010 will be based on a conversation history that User 2 does not remember, and the conversation between User 2 and the character of AI response output device 10010 may become inconsistent.
[0221] In contrast, even in a similar situation, if the database shown in Figure 4A is used, the control unit 1110 of the character conversation device (AI response output device 10010) identifies users by ID, stores the initial setting instructions and conversation history in a different area for each user, and uses the initial setting instructions and conversation history stored in a different area for each user to generate an AI response. As a result, the initial setting instructions and conversation history used to generate an AI response for each user are based on the operation or conversation history of that user and are managed separately from the operation or conversation history of other users. This makes it possible to more appropriately maintain consistency in the conversation history between each user and each character of the AI response output device 10010.
[0222] The database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in the storage unit 1170 of the AI response output device 10010 and used by the control unit 1110. Furthermore, without being limited to this, the database of initial setting directives and / or conversation histories may be stored in a server on the network. For example, if the AI response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 in generating an AI response, the database of initial setting directives and / or conversation histories described in FIG. 4A may be stored in these servers themselves. In this way, it is possible to omit the process of transmitting the initial setting directives and conversation histories again from the AI response output device 10010 to these servers by including them in directives, thereby reducing the number of transmission tokens for use of the large-scale language model.
[0223] When storing the database of the initial setting instruction sentence and / or the conversation history described in FIG. 4A, the AI response output device 10010 can transmit the user ID, the character ID, and the user instruction sentence for the subsequent conversation to these servers. The large-scale language models on these servers can use the user ID and character ID obtained from the AI response output device 10010 to obtain the corresponding initial setting instruction sentence and the conversation history from the database of the initial setting instruction sentence and / or the conversation history of FIG. 4A. The large-scale language models on these servers can perform inference using the initial setting instruction sentence and the conversation history, and the user instruction sentence for the subsequent conversation transmitted from the AI response output device 10010, to generate an AI response and transmit it to the AI response output device 10010. In this way, the effect of more appropriately maintaining the consistency of the character's personality and the continuity of memory for each different user can be obtained while saving the number of transmitted tokens when using the large-scale language model.
[0224] Next, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) according to the fourth embodiment of the present invention will be described with reference to Fig. 4B. The database according to the fourth embodiment shown in Fig. 4B is an extension of the database described in Fig. 1C or Fig. 2L. Specifically, the database shown in Fig. 4B assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data on standard response phrases corresponding to each user and character in the database.
[0225] In the example of Fig. 4B, for user 1, whose user ID is 1, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3. In addition, for user 2, whose user ID is 2, and user 3, whose user ID is 3, response template data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3.
[0226] These standard response phrase data are stored as separate data in different areas for each user-character combination. For ease of explanation, in FIG. 4B, the data stored in each area is represented as standard response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, standard response phrase data 101 is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 1: Koto shown in FIG. 2L. Data 201 in FIG. 4B is stored as a database such as a table corresponding to standard response phrases corresponding to condition numbers 1 to 7 for character 2: Tom shown in FIG. 2L.
[0227] Data 301 in FIG. 4B is stored as a database such as a table corresponding to the fixed response phrases corresponding to condition numbers 1 to 7 of character 3: Necco shown in FIG. 2L. Data 102, 202, and 302 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 2. Data 103, 203, and 303 in FIG. 4B store fixed response phrases in a similar format that have been modified for user 3. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses fixed response phrase data stored in different areas for each combination of user and character, based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.
[0228] In this way, even the same character can respond with different template responses for each user. That is, even for the same character, it may be more appropriate to vary the content of the template responses depending on the relationship between the character and the user. For example, depending on the relationship between the character's set age and the age of the user registered in the AI response output device 10010 or the system, the user may be older, the same age, or younger than the character. In this case, different content for the template responses of the character to older users, the template responses of the character to users of the same age, and the template responses of the character to younger users can make the conversation between the user and the character more appropriate or natural. That is, by performing operations using the database of FIG. 4B and varying the content of the template responses depending on the relationship between the character and the user, it is possible to create more appropriate or natural conversations.
[0229] 4B described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the response template database (response template DB) shown in FIG. 4B may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the response template database (response template DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the response template database (response template DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the AI response output device 10010 is not equipped with a fixed response phrase database (fixed response phrase DB), it is possible to generate a response using the fixed response phrase database (fixed response phrase DB).
[0230] According to the character conversation device and character conversation system of Example 4 described above, it is possible to produce more suitable or more natural conversation depending on the relationship between the character and the user, the conversation history, etc.
[0231] <Example 5> Next, a fifth embodiment of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of the first, second, and third embodiments. Specifically, this is an example in which the response generation process of the AI response output device 10010 is switched from response generation process using a large-scale language model on a network to response generation process using a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, or response generation process using a fixed response phrase database. In this embodiment, differences from these embodiments will be described, and repeated description of configurations similar to those of these embodiments will be omitted.
[0232] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0233] An example of response generation process switching processing in the AI response output device 10010 of the fifth embodiment of the present invention will be described using FIG. 5A. The table in FIG. 5A shows examples 1 to 9 of response generation process switching processing in the AI response output device 10010. In the table in FIG. 5A, the column "Switching Overview" shows an overview of the switching processing for each example. The column "State before switching of LLM on the network (API connected LLM)" shows the state before response generation processing by a large-scale language model on the network (large-scale language model connected using an API), such as the large-scale language model provided in the large-scale language model server 19001 in FIG. 1 and the multimodal large-scale language model provided in the large-scale language model server 20001, is switched to another response generation process. The column "Switching Occurrence Condition" shows the conditions under which switching processing of the response generation process occurs. The column "Switching destination from LLM on the network (API-connected LLM)" indicates the switching destination to which the response generation process of the AI response output device 10010 is switched from a large-scale language model on the network (a large-scale language model connected using an API), such as a large-scale language model provided in the large-scale language model server 19001 and a multimodal large-scale language model provided in the large-scale language model server 20001. When the condition indicated in "Switching occurrence condition" occurs in the state of "State before switching of LLM on the network (API-connected LLM)" shown in Figure 5A, the control unit 1110 of the AI response output device 10010 may perform control to switch to the large-scale language model, database, or correspondence indicated in "Switching destination from LLM on the network (API-connected LLM)."
[0234] Below, each example shown in the table of FIG. 5A will be described. Example 1, as shown in the "Switching Overview," is an example in which switching is performed depending on the network connection status of the AI response output device 10010. In Example 1, the "state before switching of the LLM (API-connected LLM) on the network" indicates that the network connection status of the AI response output device 10010 is connectable. Here, in Example 1, the "switching occurrence condition" indicates "when network connection becomes unavailable." That is, this means when the connection via the network between the AI response output device 10010 and a large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable. Specifically, the connection failure may be due to a communication failure state on the connection path from the AI response output device 10010 to the Internet 19000. Alternatively, the connection failure may be due to a communication failure state on the Internet 19000. Alternatively, the connection failure may be due to a situation in which the large-scale language model on the network (a large-scale language model connected using an API) itself cannot connect to the Internet 19000. Also, in Example 1, "local LLM" is shown as the "switching destination from LLM on the network (API-connected LLM)." This specifically means that switching processing is performed to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. That is, in Example 1, even if connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and response generation processing by the large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable, response generation processing is switched to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. This makes it possible to continue response generation processing using the large-scale language model, despite differences in performance as a large-scale language model.
[0235] Next, Example 2 in FIG. 5A will be described. In Example 2, the "switching destination from the networked LLM (API-connected LLM)" in Example 1 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 2, if a connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and the response generation process using the large-scale language model on the network (a large-scale language model connected using an API) cannot be used, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response through simpler processing and output the response to the user.
[0236] Next, Example 3 of FIG. 5A will be described. In Example 3, the "switching destination from LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "no-response handling." The "no-response handling" means that even if a user input requesting a response from a large-scale language model is received from the user via the touch panel, microphone 1139, or operation input unit 1107, no response to this input is generated, or even if a user input requesting a response from a large-scale language model is received, no response to this is output. In other words, Example 3 makes it easier to handle a situation where connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason, making it impossible to use the response generation process by the large-scale language model on the network (a large-scale language model connected using an API).
[0237] Next, Example 4 of FIG. 5A will be described. As shown in the "Switching Overview," Example 4 is an example in which switching is performed due to a response delay of an LLM on the network. In Example 4, the "state before switching an LLM on the network (API-connected LLM)" indicates a state in which a response from an LLM on the network is obtained within a predetermined time. Here, in Example 4, the "switching occurrence condition" indicates a case in which a response from an LLM on the network is not obtained within the predetermined time and exceeds the predetermined time. Also, in Example 4, the "switching destination from an LLM on the network (API-connected LLM)" indicates a "local LLM." The "local LLM" of the switching destination is the same as in Example 1, so repeated explanation will be omitted. That is, in Example 4, even if the response from an LLM on the network (a large-scale language model connected using an API) exceeds the predetermined time for some reason and the response generation process by the LLM on the network (a large-scale language model connected using an API) cannot be used smoothly, the response generation process is switched to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This makes it possible to continue the response generation process using a large-scale language model, even if there are differences in performance as a large-scale language model.
[0238] Next, Example 5 in FIG. 5A will be described. Example 2 is an example in which the "switching destination from the networked LLM (API-connected LLM)" in Example 4 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so a repeated description will be omitted. That is, in Example 5, if for some reason a response from the networked LLM (large-scale language model connected using an API) exceeds a predetermined time and the response generation process using the networked LLM (large-scale language model connected using an API) cannot be used smoothly, by switching to response generation process using the preparatory response phrase database, it becomes possible to generate a response using simpler processing and output the response to the user.
[0239] Next, Examples 6 to 9 of FIG. 5A will be described. As shown in "Switching Overview," Examples 6 to 9 are examples in which switching is performed when the upper limit of API usage or usage fee is reached. As described in Example 2, providers of large-scale language models often collect the costs used to train the large-scale language model from terminal users as API usage fees for the terminal. In such cases, with natural language models, API usage fees are often charged based on the number of processing of units of words that separate sentences, called tokens. Here, various methods of charging and limiting API usage fees are conceivable. One possible example is to define the upper limit of the amount of large-scale language model usage services that a user can receive under normal circumstances using the number of tokens processed.
[0240] In this case, the user may be able to receive services using large-scale language models at a specified API usage fee until the usage volume (or corresponding usage fee) is reached, and once the upper limit of the usage volume (or corresponding usage fee) is reached, certain restrictions may be imposed, such as the user being unable to receive services using large-scale language models in their normal state (performance or frequency).
[0241] Examples 6 to 9 in FIG. 5A are examples of response generation process switching control by the control unit 1110 of the AI response output device 10010 when such restrictions occur in a service using a large-scale language model. Specifically, in Example 6, the "state before switching the on-network LLM (API-connected LLM)" is a state in which the API usage volume and API usage fee have not reached a predetermined upper limit. This means that the usage volume of the on-network LLM (API-connected LLM) has not reached a predetermined upper limit. In this case, the user can use the on-network LLM (API-connected LLM) in a normal state.
[0242] Here, in Example 6, the "switching trigger condition" is when the API usage volume or API usage fee reaches a predetermined upper limit. This means when the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit. Also, in Example 6, the "switching destination from the network-based LLM (API-connected LLM)" is a second LLM on the network that is different from the LLM (which may be referred to as the first LLM) used in normal operation. An example of a second LLM on the network is an LLM with a lower fee than the first LLM used in normal operation. Since it is a lower-fee service, the performance of the second LLM may be lower than that of the first LLM. Even in this case, there is still a significant advantage if large-scale language models can be used inexpensively even after the usage volume / fee limit of the first LLM is reached.
[0243] Next, Example 7 in Figure 5A will be described. In Example 7, the "switching destination from the network-based LLM (API-connected LLM)" in Example 6 is changed from a second LLM on a network different from the LLM (which may be referred to as the first LLM) used in the normal state to a "local LLM." In Example 7, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the network-based LLM (API-connected LLM) reaches a predetermined upper limit, it is possible to continue performing response generation processing using a large-scale language model by switching to response generation processing using a local LLM that is not subject to restrictions such as the usage volume of the network-based LLM, the API usage volume, or the API usage fee.
[0244] Next, Example 8 in FIG. 5A will be described. In Example 8, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "prepared response phrase DB (database)." The response generation process using the "prepared response phrase DB (database)" is similar to the process described in FIG. 1C, FIG. 2L, or FIG. 4B, and therefore will not be described again. In Example 8, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the networked LLM (API-connected LLM) reaches a predetermined upper limit, the process switches to response generation using a preparatory response phrase database, which is not subject to restrictions such as the usage volume of the networked LLM, the API usage volume, or the API usage fee. This makes it possible to generate a response using simpler processing and output the response to the user.
[0245] Next, Example 9 in FIG. 5A will be described. In Example 9, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "no-response handling." The "no-response handling" refers to a handling in which a response to the user is not generated or a response to the user is not output. Example 9 makes it easier to handle a situation in which response generation processing using a large-scale language model on the network (a large-scale language model connected using an API) is unavailable because the API usage or API usage fee has reached a predetermined upper limit, i.e., the usage of the networked LLM (API-connected LLM) has reached a predetermined upper limit.
[0246] According to the switching control of the response generation process of the AI response output device 10010 shown in Examples 1 to 9 of Figure 5A as described above, even in situations where the response generation process using an LLM (large-scale language model connected using an API) on the network cannot be used as usual, more appropriate switching or response can be made according to each situation.
[0247] 5A may be performed by combining a plurality of examples. For example, the switching control of Examples 1 to 3 may be combined with any of the controls of Examples 4 to 9. Similarly, the control of Example 4 or Example 5 may be combined with any of the controls of Examples 1 to 3 or Examples 6 to 9. Similarly, the control of Examples 6 to 9 may be combined with any of the controls of Examples 1 to 5.
[0248] Next, an example of display of an AI assistant or a character when the AI response output device 10010 of the fifth embodiment is configured as an AI assistant device or a character conversation device will be described with reference to FIGS. 5B to 5D.
[0249] First, Figure 5B is an example of the display of an AI assistant or character on the AI response output device 10010 when performing the switching control of Example 3 in Figure 5A. In the example of Figure 5B, the display state of the AI assistant or character is changed depending on whether the network connection status of the AI response output device 10010 is network connectable or network unconnectable. The network connectable and network unconnectable states of the AI response output device 10010 are as explained in Figure 5A, so a repeated explanation will not be given.
[0250] In the example of FIG. 5B, the AI response output device 10010 (1) displays the AI assistant or character in a normal, awake state when a network connection is possible, but (2) displays the AI assistant or character in a "sleeping" state when a network connection is not possible. In the switching control of example 3 of FIG. 5A, when the AI response output device 10010 cannot connect to the network, it does not generate or output a response even if a command is input from the user. In this case, if the AI assistant or character displayed by the AI response output device 10010 is in a normal, awake state, the user will feel uncomfortable. However, if the AI assistant or character displayed by the AI response output device 10010 is displayed in a sleeping state, the user will understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user.
[0251] In the case of Figure 5B(2), it is desirable that the user understand that "the reason the AI assistant or character is not responding is because it is asleep" before making a user input requesting a response using a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 1001. Therefore, it is desirable that the start timing of the state in Figure 5B(2) where the AI assistant or character is displayed in a "sleeping" state when network connection is not possible, is immediately after the control unit 1110 of the AI response output device 10010 determines that network connection is not possible, before the user makes a user input requesting a response using a large-scale language model.
[0252] Next, as another display example, the display example of FIG. 5C will be described. The display example of FIG. 5C is an example in which the display state of the AI assistant or character is changed depending on the state of the "switching destination from the LLM on the network (API-connected LLM)" in the table during the switching control of FIG. 5A. Specifically, FIG. 5C shows (1) a display example of the AI assistant or character in a state in which the AI response output device 10010 can connect to a large-scale language model on the network (a large-scale language model connected using an API) and is able to use a response generation process using the large-scale language model on the network (referred to as the normal state in this figure); (2) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to a response generation process using an LLM or a template response database with lower performance than the large-scale language model on the network (a large-scale language model connected using an API); and (3) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A.
[0253] In the example of FIG. 5C , for example, (1) when the AI response output device 10010 is in a "normal state," the AI response output device 10010 displays the AI assistant or character in a state where there are no particular problems. Note that the "normal state" in FIG. 5C may be considered a state other than states (2) and (3). Also, for example, (2) when the AI response output device 10010 has switched to a response generation process using an LLM or a response template database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), the AI response output device 10010 displays the AI assistant or character in a "sleepy" state. Note that "displaying the AI assistant or character in a "sleepy" state" may also be expressed as "a display indicating that the AI assistant or character is feeling drowsy."
[0254] The response generation process (2) has lower performance than the response generation process using a large-scale language model on the network (a large-scale language model connected using an API) in the normal state (1). Therefore, by displaying the AI assistant or character in a "sleepy" state, it is possible to implicitly convey to the user that the response performance of the AI assistant or character is low. This makes it possible to further reduce the sense of discomfort felt by the user due to a low-performance response. Note that the switching conditions under which the AI response output device 10010 switches to response generation processing using an LLM or a fixed response phrase database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), are as described in FIG. 5A, and therefore a repeated explanation will be omitted.
[0255] 5C(2), it is desirable to implicitly inform the user that the response performance of the AI assistant or character is low before the user makes a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the start timing of the state in FIG. 5C(2) where the AI assistant or character is displayed in a "sleepy" state be immediately after the AI response output device 10010 switches to a response generation process using an LLM or a template response database, which has lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user makes a user input requesting a response from a large-scale language model.
[0256] Also, for example, in (3) the state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A, the AI response output device 10010 displays the AI assistant or character in a "sleeping" state. As also described in FIG. 5B, by displaying the AI assistant or character displayed by the AI response output device 10010 in a "sleeping" state, the user can understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user. Note that the conditions under which the AI response output device 10010 switches to the no-response mode described in FIG. 5A are the same as those described in Example 3 or Example 9 of FIG. 5A, and therefore a repeated explanation will be omitted. Note that in the case of FIG. 5C(3), it is desirable for the user to understand that "the AI assistant or character is not responding because it is sleeping" before making a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the timing for starting the state (3) in Figure 5C in which the AI assistant or character is displayed in a "sleeping" state be immediately after the AI response output device 10010 switches to the no-response response described in Figure 5A, before the user input requesting a response from a large-scale language model.
[0257] 5C, the AI response output device 10010 displays a display that implicitly reflects a change in the state of the AI assistant or character without directly providing the user with a technical explanation of the state of the AI response output device 10010 related to the response generation process. This can reduce the sense of discomfort felt by the user compared to when a technical explanation of the state of the AI response output device 10010 related to the response generation process is directly provided to the user. Furthermore, this can reduce the sense of discomfort felt by the user compared to when the display state of the AI assistant or character remains the same as its normal state despite a change in the state of the AI response output device 10010 related to the response generation process.
[0258] However, some users may wish to know a more precise explanation of the technical state of each state. Therefore, a display example for such users will be described with reference to FIG. 5D. Among the rows of the table shown in FIG. 5D, the rows for explaining the device state and display state are identical to those in FIG. 5C, and therefore, repeated explanations will be omitted. Furthermore, the display example of the AI assistant or character shown in the row for the display example of the AI assistant or character is almost identical to that in FIG. 5C, except that a question mark (?) is displayed in the display example. The question mark (?) is a mark that the user operates when requesting an explanation from the AI response output device 10010, and may also be referred to as a help mark.
[0259] In the example of FIG. 5D , when a user selects the question mark (?) through a user operation, such as via the touch panel of the operation input unit 1107 or the display unit 10011 in FIG. 1B , the display of the AI assistant or character on the AI response output device 10010 changes to the display example shown in the row of the display example after user operation. Specifically, regardless of whether the device state is (1), (2), or (3), an explanation of the technical state of each state is displayed. For example, in the example of FIG. 5D , if the device state is (1) normal, a display explaining that the device is in a normal state with no particular technical limitations can be displayed, such as "Normal state." Furthermore, if the device state is (2) using a low-performance LLM or a standard response phrase database, a display technically explaining the low-performance state can be displayed, such as "Low-performance mode." This display can also be considered a display explaining the reason why the AI assistant or character is displaying a "sleepy" state.
[0260] In this case, a more detailed technical explanation may be provided. Specifically, a message such as "Low-performance LLM usage mode" or "Canned response mode" may be displayed. If the device status is (3) unresponsive, a message such as "Network connection unavailable" may be displayed, providing a technical explanation of the reason for switching to unresponsive mode. If the reason for switching to unresponsive mode is that a response from an LLM (a large-scale language model connected via an API) on the network exceeds a specified time, a message such as "Response from LLM is delayed" may be displayed. If the reason for switching to unresponsive mode is that the network LLM usage, API usage, or API usage fee has reached its limit, a message such as "LLM usage limit reached," "API usage limit reached," or "API usage fee has reached a specified amount" may be displayed. These messages may be considered to explain the reason why the AI assistant or character is displayed in a "sleeping" state.
[0261] According to the display example of FIG. 5D described above, even if there are technical constraints in the response generation process in the AI response output device 10010, first, instead of providing a direct explanation to the user, the state of the device is implicitly indicated by a change in the display state of the AI assistant or character, thereby further reducing the sense of discomfort felt by the user. This display is more suitable for users who do not need technical explanations. Furthermore, by displaying an operation mark to explain the technical state, a display is provided to users who operate the mark that technically explains the state of the response generation process in the AI response output device 10010 (normal state or state with technical constraints). This makes it possible to provide a more suitable display for users who want to know the technical state accurately.
[0262] In the examples of Figures 5B, 5C, and 5D, a "sleeping" state is shown as an example of the display state of the AI assistant or character when the AI assistant or character is "unresponsive," but this is only an example and the embodiment is not limited to this. Instead of the "sleeping" state, another display state that implies a situation where the AI assistant or character is unable to respond, such as "taking a break," may be used. In the examples of Figures 5C and 5D, a "sleepy" state is shown as an example of the display state of the AI assistant or character when a low-performance LLM or a standard response phrase database is being used, but this is only an example and the embodiment is not limited to this. Alternatively, another display state that implies that the AI assistant or character has low response performance, such as "hungry," may be used.
[0263] According to the AI response output device and the AI response output system according to the fifth embodiment described above, it is possible to more appropriately switch the response generation process used by the AI response output device depending on the connection state between the large-scale language model on the network and the AI response output device, the response delay state from the large-scale language model on the network, the usage amount of the large-scale language model on the network, etc. Furthermore, when the AI response output device according to the fifth embodiment is configured as an AI assistant device or a character conversation device, it is possible to perform a display that is less strange to the user.
[0264] Example 6 Next, Example 6 of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in the drawings of Examples 1 to 5. Specifically, this is an example in which the response generation process of the AI response output device 10010 is more suitably combined with a response generation process using a large-scale language model on a network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, and a response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated explanations of configurations similar to those examples will be omitted.
[0265] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0266] An example of a response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 shows an example of a flowchart of the response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention. Specifically, a time axis in which time progresses from top to bottom, a processing flow, and a response output example are shown. The response shown in the response output example may be output via display on the display unit 10011 of the AI response output device 10010 or audio output by the audio output unit 1140.
[0267] In the example of FIG. 6, first, at time t0, a user input is made via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010, requesting a response based on a large-scale language model, and the control unit 1110 of the AI response output device 10010 acquires the user input (step 600). Next, at time t1, the control unit 1110 starts preparations for response output using a fixed response phrase database stored in the storage unit 1170, and starts response output using the fixed response phrase database (step 601). In the example of FIG. 6, response output using the fixed response phrase database starts at time t2, and as shown in the figure, the fixed response is being output but has not yet been completed. "Good morning" in the figure indicates the output of part of the sentence that continues "Good morning..."
[0268] At time t3, before the response output using the fixed response phrase database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and transmits the generated instruction sentence to a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, thereby starting a request for a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 starts acquiring a response from the large-scale language model (step 603).
[0269] An example of a response output in which the response output using the fixed response phrase database is completed at time t5 is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is the ____ day of the month, isn't it?" is completed at time t5 using the fixed phrases stored in the fixed response phrase database and date information stored in memory. Here, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 has already started acquiring a response from the large-scale language model. Therefore, at time t6, which follows time t5 when the response output display using the fixed response phrase database is completed, the control unit 1110 starts outputting a response from the large-scale language model following the response output using the fixed response phrase database (step 604). Thereafter, at time t7, a response from the large-scale language model is output following the response output using the fixed response phrase database. When the response output from the large-scale language model is completed, the response output according to the processing flow shown in FIG. 6 is completed (step 605).
[0270] Next, the effect of the processing flow shown in Fig. 6 of the present invention will be described. Processing a large-scale language model requires a large amount of computational resources. Generally, even if inference, which requires fewer computational resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several to several tens of seconds from when the control unit starts to request a response from the large-scale language model until it can obtain a response from the large-scale language model. This period corresponds to the period from time t3 to time t4 shown in Fig. 6. Furthermore, from time t0, when a user input is made, until time t4, the control unit 1110 is unable to obtain a response output from the large-scale language model, and therefore is unable to output a response from the large-scale language model to the user.
[0271] 6, there is no start of preparation for response output using the fixed response phrase database and no start of response output using the fixed response phrase database, the user may continue to wait for a period of several seconds to more than ten seconds from time t0 when the user input is made to time t4 without receiving a response from the AI response output device 10010. For example, when the AI response output device 10010 is configured as an AI assistant device or a character conversation device, the waiting time may give the user a sense of discomfort.
[0272] In contrast, in the processing flow according to the sixth embodiment of the present invention shown in Fig. 6, the control unit 1110 starts the process of outputting a response using a fixed response phrase database, which requires fewer computational resources than the process of a large-scale language model, before starting to acquire a response from the large-scale language model. This prevents the user from having to wait from time t0 to time t4 without receiving a response from the AI response output device 10010. From the user's perspective, whether the response output is using a fixed response phrase database or a response output from a large-scale language model, it is the same as receiving a response from the AI response output device 10010.
[0273] 6, by providing step 601 before step 603, the response of the AI response output device 10010 to the user can be artificially accelerated. This can further reduce the sense of discomfort felt by the user due to long waiting times. Furthermore, by outputting a response from a large-scale language model following a response using the template response database in step 604, the user can perceive these outputs as if they were a series of more natural outputs.
[0274] According to the AI response output device and AI response output system of Example 6 described above, the waiting time for a response from the AI response output device can be shortened, and the sense of discomfort felt by the user can be further reduced.
[0275] Example 7 The seventh embodiment of the present invention is an improvement of the AI response output device 10010 or the AI response output system described above. In this embodiment, differences from the first embodiment will be described, and repeated explanations of the same configurations as those embodiments will be omitted.
[0276] The AI response output system according to the seventh embodiment is an example in which, when an instruction sentence is sent from the AI response output device 10010 or the AI response output system to a large-scale language model, a process of disambiguating a user's question sentence, media information that adds or supplements the question sentence, etc. is performed as needed. More specifically, when creating an instruction sentence (prompt) for a large-scale language model AI from a user's question sentence, media information that supplements the question sentence, etc., the system checks the ambiguity of the user's question sentence and the media information that adds or supplements it, and performs a process of disambiguating the user's question sentence, media information that supplements the question, etc. as needed, so that a more appropriate answer to the user's question sentence can be obtained from the AI.
[0277] An example of a response output system according to a seventh embodiment of the present invention will be described with reference to Fig. 7. As shown in Fig. 7, the configuration of the response output system according to the seventh embodiment is basically the same as that of the first embodiment. In the seventh embodiment, an AI response output device 10010 is communicably connected to a large-scale language model server 19001 or a large-scale language model server 20001 equipped with a large-scale language model, and a second server 19002 different from these servers, via the Internet 19000. However, the above servers may be provided with a check database (check DB), and in this embodiment, an example will be described in which the second server 19002 is provided with a check database (check DB) 19010.
[0278] In the seventh embodiment, a configuration will be described in which the large-scale language model server 19001 is equipped with a large-scale language model (LLM). In the following description, the large-scale language model equipped in the large-scale language model server 19001 may be simply referred to as the large-scale language model or LLM. Incidentally, the large-scale language model may be a large-scale language model on a network, or a local large-scale language model (such as the local LLM processing unit 10028) equipped in the AI response output device 10010.
[0279] In the above configuration, when a user asks a question or makes a request to a large-scale language model, which is an artificial intelligence, and attempts to obtain a response or answer, the process is, for example, as follows: First, the user inputs the question or request to the large-scale language model to the AI response output device 10010. As an example, the user operates the operation input unit 1107 (FIG. 1B), which is an input unit, to input a question or request, which is text information in a natural language. The question / request input to the large-scale language model may be input by operating the operation input unit 1107 (FIG. 1B) or may be input by voice. Furthermore, if the large-scale language model is a multimodal LLM capable of handling data in various formats, the user may also input various media information, such as images, videos, and audio, that the large-scale language model can handle. The AI response output device 10010 generates a response instruction sentence (prompt) for the large-scale language model based on the question, request, and media information input by the user, and transmits the generated response instruction sentence (prompt) to the large-scale language model server 19001 equipped with the large-scale language model. When the large-scale language model generates a response to the response instruction sentence (for example, an answer sentence to a user's question), the generated response is sent from the large-scale language model server 19001 to the AI response output device 10010.
[0280] Upon receiving this response, the AI response output device 10010 processes the response of the large-scale language model (LLM) for the user. As an example of processing the response of the large-scale language model, the AI response output device 10010 displays a response sentence (also called an answer sentence) generated by the large-scale language model on the display unit 10011. In this manner, the user can obtain a response from the large-scale language model to their question or request.
[0281] However, users may not always be able to obtain the desired response from the large-scale language model. For example, depending on how a question is written, the content of the question may be interpreted differently from the user's intention, and the desired response may not be obtained from the large-scale language model. In another example, if the resolution of an input image is low, the content of the image may be misinterpreted by the multimodal LLM, and the desired response may not be obtained.
[0282] Therefore, in the response output system according to the seventh embodiment, when a user inputs a question or request to the AI response output device 10010, the AI response output device 10010 checks the ambiguity of the question or request, performs processing to resolve the ambiguity of the question or request as necessary, and generates a response instruction sentence based on the question sentence after the processing to resolve the ambiguity (hereinafter, sometimes referred to as the processed question sentence). This allows the large-scale language model to more appropriately interpret the question or request, making it easier to obtain the response desired by the user from the large-scale language model. Furthermore, if the question or request input by the user is ambiguous and can be interpreted in multiple ways, a response instruction sentence is generated for each interpretation. The large-scale language model responds with the response instruction sentence generated based on each interpretation. This increases the likelihood that the response desired by the user will be included in the multiple responses from the large-scale language model, even when the user inputs a question or request that can be interpreted in multiple ways.
[0283] The method for inputting a question or request by a user to the AI response output device 10010 may be character input using the operation input unit 1107 as an input unit, or a method for converting a voice-input question using the microphone 1139 into text. Media information such as an image may be input by selecting a media information file using the operation input unit 1107. The process of checking the ambiguity of the question or request, resolving the ambiguity, checking the ambiguity (unclearness) of the media information, generating an instruction statement, and outputting an answer from the large-scale language model to the display unit 10011 may be executed by, for example, the control unit 1110. The control unit 1110 in this embodiment generates a response instruction statement for the large-scale language model based on the question statement and acquires a response generated by the large-scale language model in response to the response instruction statement. Specifically, the control unit 1110 checks the ambiguity of the input question, request, or media information, executes a process of resolving the ambiguity as necessary, and generates a response instruction statement based on the processed question or request.
[0284] Hereinafter, a response output processing flow in the response output system according to the seventh embodiment will be described in detail, mainly a processing flow for obtaining a response from a large-scale language model to a user's question or request when the question or request contains ambiguity. FIG. 8A is a diagram showing an example of a processing flow of the response output system according to the seventh embodiment. FIG. 8B is a diagram showing an example of a processing flow of the AI response output device according to the seventh embodiment. FIG. 9 is a diagram showing a detailed example of the ambiguity resolution processing of the response output system according to the seventh embodiment. Furthermore, FIGS. 10A, 10B, 10C, and 10D are diagrams showing an example of the operation of the response output system according to the seventh embodiment. In addition, in the diagram showing the processing flow of the response output system, the same steps are assigned the same symbols, and duplicated explanations of each step may be omitted.
[0285] First, an outline of the processing flow of the response output system according to the seventh embodiment will be described with reference to Fig. 8A. As shown in Fig. 8A, in step S1000, when an AI response output device 10010 receives an input such as a question or request from a user, it checks whether the input content is ambiguous, and if ambiguity is found, performs ambiguity resolution processing. When there is no ambiguity in the user's input content, it generates an instruction statement requesting a response based on the user's input content, sends the generated instruction statement to the large-scale language model, receives a response from the large-scale language model, and displays it on the display unit 10011.
[0286] 8A and 8B, the detailed operation of the response output system according to the seventh embodiment will be described. FIG. 8B is a processing flow illustrating the operation of step S1000 in detail. As shown in FIG. 8B, first, a user operates, for example, the operation input unit 1107 to input a question or request to the large-scale language model. In step S1001, the control unit 1110 receives the user input (question or request). In step S1002, the control unit 1110 executes a check process to determine whether the question or request input by the user is ambiguous. In step S1003, if the check process determines that the question or request needs to be disambiguated, the control unit 1110 generates a disambiguated question or request (hereinafter referred to as a disambiguated question or disambiguated request). Then, based on the disambiguated question (or disambiguated request), the control unit 1110 generates a prompt for the large-scale language model (LLM) included in the large-scale language model server 19001. In this example, the control unit 1110 generates a response prompt instructing the creation of an answer sentence in response to the user's question. These checking and disambiguation processes are described in more detail below.
[0287] Next, this response instruction sentence is sent to the large-scale language model (step S1004). Note that the response instruction sentence may also be simply called an instruction sentence.
[0288] Next, when the large-scale language model (LLM) receives an instruction sentence, for example, a response instruction sentence, sent by the AI response output device 10010 (step S1101), it generates a response based on this instruction sentence (step S1102). In the example shown in Fig. 8, since the instruction sentence is a response instruction sentence, the large-scale language model creates an answer sentence or response sentence to the user's question as the response generation in step S1102. Then, in step S1103, the answer sentence or response sentence created by the large-scale language model is sent to the AI response output device 10010.
[0289] When the control unit 1110 of the AI response output device 10010 receives a response (e.g., an answer sentence) created by the large-scale language model (step S1005), it executes processing on the response of this large-scale language model in step S1006. As an example of processing on the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question.
[0290] As described above, in the response output system according to the seventh embodiment, when performing the response output process, a process is performed to check whether or not there is ambiguity in the question or request input by the user, and a process is performed to resolve the ambiguity of the question or request as necessary. This makes it easier to obtain the user's intended answer to the user's question or request from the large-scale language model.
[0291] Next, the ambiguity confirmation process for the user input content in step S1002 and the ambiguity resolution process in step S1003 will be described in more detail. For example, as shown in Fig. 8B, when control unit 1110 receives a user input in step S1001, control unit 1110 then performs a process in step S1002 to confirm whether or not there is ambiguity in the input content, such as a question or request entered by the user, and determines whether or not ambiguity resolution process is required for the input content.
[0292] Processing according to the determination result of this confirmation process is performed in a disambiguation process group in step S1003. For example, if it is determined that the input content is not ambiguous, the disambiguation process group in step S1003 is not performed, and an instruction sentence is generated based on the input content entered by the user. On the other hand, if it is determined that the input content is ambiguous, the disambiguation process group in step S1003 is performed, and an instruction sentence is generated based on the disambiguated content of the user input content. Specific examples of the disambiguation process group will be described later. For example, the disambiguation process group may be a process for disambiguating text or sentences, where the meaning of a question or request statement entered by the user can be interpreted in multiple ways, or a process for disambiguating media information, where the content of an image entered by the user as media information is difficult to accurately determine due to low resolution (in other words, unclearness). The disambiguation process may be repeated. For example, the disambiguated content may be rechecked to see if there is any ambiguity, and if there is any ambiguity, the disambiguation process may be performed again.
[0293] To explain the ambiguity confirmation process of step S1002 in more detail, the control unit 1110 executes a process to confirm whether a question, request, or the like entered by the user in text corresponds to a predetermined item as the confirmation process. As an example, the control unit 1110 executes a process to confirm whether a sentence entered by the user contains a specific ambiguous expression set in advance. Examples of the specific ambiguous expression include ambiguous expressions relating to numbers (cheap, expensive, many, young, etc.), ambiguous expressions relating to degrees (very, quite, very, very slowly, etc.), and ambiguous expressions relating to demonstratives (this, that, which, which, before, next time, etc.). If the user-input sentence contains such a specific ambiguous expression, it is determined that disambiguation processing is necessary. Note that the contents of the confirmation items are not particularly limited, and may be any items that can determine whether or not to execute disambiguation processing for the user-input sentence.
[0294] In this example, the specific ambiguous expressions refer to words, grammar, and the like that are pre-registered in a check database (check DB) 19010 of the second server 19002. That is, it is assumed that a list of specific ambiguous expressions is stored in the check DB 19010 of the second server 19002. The control unit 1110 then executes ambiguity confirmation processing for the user-input sentence using the check DB 19010. Specifically, as the ambiguity confirmation processing for the user-input sentence in step S1002, the control unit 1110 first compares words and grammar in the user-input sentence with the specific ambiguous expressions registered in the check DB 19010. Next, it is determined whether or not any words and grammar in the user-input sentence match any of the specific ambiguous expressions registered in the check DB 19010.
[0295] As another example of the ambiguity confirmation process of step S1102, a media information ambiguity confirmation process will be described. As the confirmation process, when a user inputs media information such as an image, video, or audio, the control unit 1110 executes a process of confirming whether the media information satisfies a preset media information quality standard. As an example, the control unit 1110 checks the quality of the media information input by the user, and determines that there is no ambiguity if the preset quality standard is met, or that there is ambiguity if the preset quality standard is not met. The media information quality standard may be, for example, a separate index set for each type of media information, such as resolution for images or video, bit rate for audio, or the presence or absence of noise. The method for confirming the ambiguity of media information and the index used for the confirmation are not particularly limited, and may be any index that can determine whether or not to perform ambiguity resolution processing on the media information input by the user.
[0296] As a process for checking the ambiguity of media information entered by a user, the control unit 1110 acquires the media type and quality information of the media information entered by the user and performs a process for checking whether the information satisfies a predetermined quality standard. Specifically, as a process for checking the ambiguity of the user input content (here, the input of media information) in step S1002, the control unit 1110 first acquires the media type and quality information of the media information entered by the user, for example, from a header area in the media information, compares the acquired quality information with a quality standard for the acquired media information among predetermined quality standards for each piece of media information, and determines, based on the comparison result, whether the media information entered by the user satisfies the predetermined quality standard. Note that the method for acquiring the media type and quality information of the media information entered by the user is not particularly limited, and may be, for example, acquired by executing a software tool within the device or on an arbitrary server that analyzes the media information and returns the media information and quality information.
[0297] Next, details of the disambiguation process in step S1003 will be described with reference to FIG. 9. Example 1 in FIG. 9 is an example of the disambiguation process, in which if there is ambiguity in the user input content, the system notifies the user of the ambiguity and requests resolution of the ambiguity. The disambiguation process is realized by the AI response output device 10010 and the large-scale language model operating in cooperation with each other, and the operation of both will be described in detail here. First, if the control unit determines that the content entered by the user is ambiguous through the above-mentioned ambiguity confirmation process, for example, if the text input contains a specific ambiguous expression registered in the check DB 19010, or if the media information, such as an image, is unclear and does not meet the quality standard, it displays a message to the user to notify the user. Here, unclear media information refers to, for example, quantitative criteria such as an image with a resolution (unit: dpi) lower than a predetermined value or a video with a bit rate (unit: bps) lower than a predetermined value, or qualitative criteria such as an image or video in which it is impossible to determine the identity and number of objects depicted in the image or video using image recognition technology. Specific examples of notifying the user that the input text contains ambiguity or that the input media information contains unclear information include displaying in the display area of the AI response output device 10010 which parts or expressions in the text are ambiguous, or that the quality of the media information is so low that the content cannot be accurately determined. At this time, the ambiguous expressions may be highlighted by underlining or changing the color, the color of the media information file name may be changed, or an icon such as "?" or "!" may be added to the media information file icon to highlight the deficiency in the media information. The system then requests the user to correct the text input and re-enter it, or to replace the media information with the same content but of higher quality and re-enter it, and waits for the user to re-enter. Note that if there is ambiguity in the text input, the re-entry may be performed by voice rather than by text. When entering text, ambiguity due to sentence or grammar is likely to occur, but switching to another input method, voice input, may resolve the ambiguity.When re-inputting is repeated multiple times, if the same text as previously inputted is inputted again, an error may be returned, or the text may be displayed in a different color or underlined.
[0298] When a user re-enters text or media information, the ambiguity check process is performed again on the re-entered information. The details of the ambiguity check process are as described above, and the process content does not differ significantly between initial input and re-entry. However, for example, by remembering where the ambiguous parts of the text input were and what kind of ambiguous expressions were used, and by highlighting or changing the color of the parts that contained ambiguous expressions the previous time, the process can be performed more efficiently. On the other hand, if the ambiguity is not resolved even after re-entry, a third re-entry is requested. If no improvement is seen even after the user has repeatedly entered the information a predetermined number of times, a prompt sentence may be generated based on the user's input, a response may be requested from the large-scale language model, and the obtained response may be presented to the user. However, in this case, since ambiguity remains in the user's input, the response obtained from the large-scale language model may not be the answer the user desires, so the above process is repeated. During the repeated process, if the re-entered information also contains ambiguous expressions, the ambiguous expressions may be displayed in red or highlighted to make them easier for the user to identify.
[0299] On the other hand, if the ambiguity is resolved by re-input, the control unit 1110 generates a command statement for the large-scale language model (LLM) based on the disambiguated user input, then sends the generated command statement to the large-scale language model and waits for a response from the large-scale language model.
[0300] Now, moving on to the explanation of the operation of the large-scale language model, when the large-scale language model receives an instruction sent from the control unit, for example, a response instruction, it generates a response based on this instruction. The instruction may be composed of text and media information, and may request the large-scale language model to answer a text question accompanied by media information, such as "Where is the location indicated by the media information (image)?". The response created by the large-scale language model is then sent to the AI response output device 10010.
[0301] When the control unit 1110 of the AI response output device 10010 receives a response (e.g., an answer sentence) created by the large-scale language model, it executes processing on the response of this large-scale language model. As an example of processing on the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question.
[0302] As described above, in the response output system according to Example 1 of the seventh embodiment, when performing the response output process, a process of checking the ambiguity of the question sentence or request sentence input by the user is performed, and a process of resolving the ambiguity as necessary is performed, in which the ambiguous part is pointed out to the user and the user is prompted to re-input after the ambiguity has been resolved. This makes it possible to generate a question or request with the ambiguity resolved, making it easier to obtain the answer the user intended for the user's question from the large-scale language model.
[0303] Example 2 in FIG. 9 is an example in which, like Example 1, there is ambiguity in the user input content, and if the ambiguity can be easily resolved, the control unit 1110 of the AI response output device 10010 resolves the ambiguity and generates an instruction statement for the large-scale language model. Examples of ambiguities that can be easily resolved include typos, omissions, incorrect conversion of kanji, spelling mistakes in English, and improperly placed punctuation. If such ambiguities are included in the question or request statement input by the user, rather than prompting the user to re-input as in Example 1, the control unit 1110 corrects the question or request statement input by the user to resolve the ambiguity, and then generates an instruction statement for the large-scale language model. The process of sending the generated instruction statement to the large-scale language model, waiting for a response, and, if a response is received, displaying it on the display unit 10011 as an answer to the user's question is the same as in Example 1, and therefore a detailed description thereof will be omitted.
[0304] As described above, in the response output system according to Example 2 of the seventh embodiment, when performing the response output process, a process of checking the ambiguity of the question sentence or request sentence input by the user is performed, and a process of resolving the ambiguity as necessary, in which the control unit 10010 resolves ambiguity that can be easily resolved. This makes it possible to generate a question or request with resolved ambiguity, making it easier to obtain the answer intended by the user from the large-scale language model in response to the user's question.
[0305] 9 shows an example in which, when a user input includes media information such as an image and the media information is ambiguous (e.g., unclear), the large-scale language model is requested to replace the media information with clearer information, and the large-scale language model is instructed to respond to the user's question or request again based on the clear image generated by the large-scale language model. First, when a user inputs ambiguous media information (hereinafter referred to as first media information), the control unit 1110 of the AI response output device 10010 generates an instruction statement requesting that the ambiguity be resolved without changing the content of the first media information, that the first media information be remade to have a higher resolution if it is an image, or to have less noise if it is audio, or that the first media information be found by search, and sends this instruction statement together with the first media information to the large-scale language model, and waits for a response from the large-scale language model.
[0306] Now, moving on to the explanation of the operation of the large-scale language model, when the large-scale language model receives an instruction sent from the control unit, for example, a response instruction, it generates a response based on this instruction. Here, by correcting or processing the transmitted first media information, such as removing noise, the content of the media information itself is not changed, and second media information that has been resolved with ambiguity is generated. The second media information generated by the large-scale language model is then transmitted to the AI response output device 10010.
[0307] When the control unit 1110 of the AI response output device 10010 receives a response (here, second media information) created by the large-scale language model, it generates an instruction sentence for the large-scale language model based on the disambiguated second media information, rather than the ambiguous first media information, and the question or request input by the user (originally input by the user along with the first media information). The control unit 1110 then transmits the generated instruction sentence to the large-scale language model and waits for a response from the large-scale language model. After receiving the second media information, the control unit 1110 may confirm with the user that the content of the second media information is consistent with the first media information originally input by the user and that the content has not changed. In this case, for example, the control unit 1110 may present the user with a message that the first media information was unclear and therefore has been corrected, along with the content of the second media information, and may also display a message asking the user, "Is the content consistent?", allowing the user to respond and select "Yes" or "No" to the message. By doing this, the large-scale language model can be instructed to generate a response to the user's original question after confirming that the content of the second media information is identical to the content of the first media information and is simply a disambiguation, which is expected to make it even easier for the user to obtain the answer they want.
[0308] Then, when the large-scale language model receives an instruction sent from the control unit, for example, a response instruction and media information, it generates a response based on the instruction and media information.The answer sentence created by the large-scale language model is then sent to the AI response output device 10010.
[0309] The subsequent operation of the control unit 1110 of the artificial intelligence response output device 10010 is not significantly different from Examples 1 and 2. When a response (e.g., an answer sentence) created by a large-scale language model is received, the control unit 1110 processes the response of this large-scale language model, for example, by displaying the answer sentence to the user's question on the display unit 10011.
[0310] As described above, in the response output system according to Example 3 of the seventh embodiment, when performing a response output process, a process of checking the ambiguity of a question sentence or a request sentence input by a user is performed, and a process of disambiguating the ambiguity as necessary is performed. In this case, when unclear and ambiguous media information is input, the input media information is corrected to clear and unambiguous media information. This allows questions and requests to be generated based on the disambiguated media information, making it easier to obtain the user's intended answer to the user's question from the large-scale language model. Note that the operation of Example 3 is particularly effective when the input first media information is clear but unclear. When the input first media information is so unclear that the content is unintelligible or when the media information file is corrupted, it is preferable for the user to manually change the media information to clearer information as in Example 1, rather than performing corrections, etc., as in Example 3.
[0311] Example 4 in Figure 9 shows the operation when there is no ambiguity in the user input content, in which case the control unit 1110 of the AI response output device 10010 generates an instruction statement for the large-scale language model based on the user input content and sends this to the large-scale language model. When the large-scale language model receives the instruction statement sent from the control unit, it generates a response based on this instruction statement, and sends the generated answer statement to the AI response output device 10010. The control unit 1110 of the AI response output device 10010 receives the response and displays it on the display unit 10011 as an answer to the user's question.
[0312] As described above, in the response output system according to the seventh embodiment, when performing the response output process, the system checks the ambiguity of the question sentence and media information input by the user and also performs the process of eliminating the ambiguity as necessary. This makes it easier to obtain the answer intended by the user from the large-scale language model in response to the user's question.
[0313] A specific example of the operation of the response output system according to the seventh embodiment will be described with reference to FIGS. 10A to 10D. FIG. 10A illustrates an example in which the user input is ambiguous. The user inputs the text "What is the name of this red fruit?" and an image of a "red fruit" (first image: RedFruit1.png) that the user previously photographed and saved to the AI response output device 10010. The AI response output device 10010 performs an ambiguity check process for the text "What is the name of this red fruit?" and the input first image. For example, the input first image may depict a "red fruit," but it may be difficult to identify what kind of fruit it is. Specifically, the input first image may be a close-up image of a "red fruit," and the entire image may not be clear, making it difficult to identify what kind of fruit it is. In this case, the control unit 1110 displays or outputs a voice message requesting the user to "add information about 'this red fruit.'" An example of a user's response to this is to take a new photograph of "this red fruit" with a camera and input the captured image as a second image (image file: RedFruit2.png) to the AI response output device 10010. The control unit 1110 is able to grasp the full picture of "this red fruit" by inputting the second image file, and determines that the ambiguity has been resolved. It then generates an instruction sentence for the large-scale language model and sends this to the large-scale language model. It then receives a response from the large-scale language model and displays "It's an apple." (The red fruit the user asked about) on the display unit 10011 as an answer to the user's question.
[0314] FIG. 10B shows another example of a case where the user input is ambiguous. The user inputs a text question, "What is the name of this red fruit?" and an image file (RedFruit_0.png) representing a "red fruit" that the user previously photographed and saved, as media information, to the AI response output device 10010. Receiving this input, the AI response output device 10010 performs an ambiguity check process on the input. As in FIG. 10A, there are cases where the type of fruit cannot be identified from the image input by the user. That is, the control unit 1110 determines that the text and image input by the user are insufficient to derive a clear answer; in other words, the user's question is ambiguous. Therefore, the control unit 1110 generates an instruction, such as "Please provide three images of fruits that correspond to the red fruit shown in the image file RedFruit0.png, along with the names of those fruits, in order of likelihood," along with the image file (RedFruit_0.png) input by the user, and transmits this instruction to the large-scale language model. The control unit 1110 then receives three image files representing "red fruits" and the names of the fruits depicted in the images from the large-scale language model (RedFruit_1.png: apple, RedFruit_2.png: strawberry, RedFruit_3.png: cherry). The three received image files are then presented to the user along with a question such as, "Are there any of these three images that correspond to the 'red fruit' you (the user) want to know about?" The user is then prompted to select the corresponding image file. The user's selection results in the selection of a clear image of the "red fruit" the user is interested in. If the user selects, for example, RedFruit_1.png, the control unit 1110 displays the name of the red fruit received from the large-scale language model that corresponds to RedFruit_1.png, i.e., "apple," on the display unit 10011 as the answer to the user's question. Specifically, for example, the display unit 10011 may display a message such as "The red fruit the user is asking about is an 'apple.'" Alternatively, the response may be output by voice.
[0315] FIG. 10C shows another example of a case where the user input is ambiguous. The user inputs a text question, "What is the name of this song?" and an audio file representing "this song," which the user previously recorded, purchased, downloaded, or saved, into the AI response output device 10010. Upon receiving this input, the AI response output device 10010 performs an ambiguity check on the input and determines that the "this song" represented by the input audio file cannot be identified because the input audio file is extremely short or has poor sound quality. In other words, the input is ambiguous. The control unit 1110 then displays a message requesting the user to "add more information about 'this song.'" Furthermore, the control unit 1110 may also display advice to the user, such as "Please record the audio data as long as possible." One example of a user response to this is to hum the melody of "this song," record it as an audio file (melody.mp3), and input this file into the AI response output device 10010. The control unit 1110 determines that the input of this audio file has clarified what "this song" refers to and resolved the ambiguity, and generates an instruction statement for the large-scale language model. The content of the instruction statement is composed of, for example, the input audio file and text such as "Please tell me the name of the song in the audio file I'm sending," and sends this to the large-scale language model. After that, it receives a response from the large-scale language model, and displays "The name of the song the user asked about is 'Jingle Bells'" on the display unit 10011 as an answer to the user's question.Other possible questions and requests from users regarding audio include knowing the lyrics but not remembering the melody, knowing the melody but not remembering the lyrics, or knowing the lyrics and melody but not knowing the name of the song. In any case, the control unit 1110 requests the information the user has, and if that information alone is ambiguous and insufficient to identify the song, it requests the user for additional information to resolve the ambiguity, and generates instructions for the large-scale language model based on that additional information, making it easier for the user to obtain the answer they want compared to when the ambiguous information entered by the user is sent directly to the large-scale language model to obtain an answer.
[0316] FIG. 10D shows another example of an ambiguous user input. The user inputs a text question to the AI response output device 10010: "An error occurred while turning pages on my smartphone. Please tell me how to deal with this." The AI response output device 10010, upon receiving this input, performs an ambiguity check on the input and determines that the "page turning" in the user's "smartphone page turning" operation can be achieved by multiple operations, and that the operation referred to in the user's question cannot be identified; that is, the input is ambiguous. Therefore, the AI response output device 10010 displays a message requesting the user to "Please explain the 'page turning' operation in detail." Furthermore, the AI response output device 10010 may also display specific examples of information the user would like to add, such as "It would be good to have a video of the page turning operation" or "If an error message is displayed, it would be good to have the error display screen." One example of a user's response to this is to prepare a video file (PageTurning.mov) of the page turning operation and input it into the AI response output device 10010. The control unit 1110 determines that the input of this video file clarifies the meaning of the "page turning operation" performed by the user and resolves the ambiguity. An instruction statement for the large-scale language model is generated, consisting of the video file (PageTurning.mov) input by the user and text such as "Please tell me how to prevent errors from occurring when performing the operation shown in the video file," and is sent to the large-scale language model. A response is then received from the large-scale language model, and the answer to the user's question, "Please operate slowly with only one finger touching the screen," is displayed on the display unit 10011.
[0317] Next, a modified example of the response output system of the seventh embodiment will be described. In the description so far, when a user makes a question or request to a large-scale language model, it is confirmed whether there is ambiguity in the media information, such as text or images, input by the user. If it is determined that there is ambiguity, a disambiguation process is performed, and an instruction is issued to the large-scale language model to request a response based on the question sentence, request sentence, and media information that have been disambiguated. In contrast, in the modified example described below, when the question sentence or request sentence input by the user to the large-scale language model is ambiguous and can be interpreted in multiple ways, a response instruction sentence is generated for each interpretation. The large-scale language model responds with the response instruction sentence generated based on each interpretation. As a result, even when a question or request that can be interpreted in multiple ways is input by the user, the large-scale language model responds for each interpretation of the meaning, increasing the possibility that the response desired by the user will be included in the multiple responses from the large-scale language model.
[0318] A modified example of the seventh embodiment will be described in detail below with reference to FIGS. 11 and 12. FIG. 11 shows a modified example of the disambiguation process of step S1003, in which, when the user input content can be interpreted in multiple ways, multiple answers corresponding to the multiple interpretations are presented to the user. In other words, since the content input by the user can be interpreted in multiple ways as is, the range of answers from the large-scale language model becomes large, making it less likely that the user will obtain the answer they desire. Therefore, for example, if instruction 1, instruction 2, and instruction 3 are generated for each of the three interpretations, interpretation 1, interpretation 2, and interpretation 3, each of the three generated instruction statements is unambiguous, and therefore the response of the large-scale language model to these instruction statements will have a smaller range of answers, making it easier to obtain a more accurate answer.
[0319] Example 1 in FIG. 11 illustrates an example in which, when a user input can be interpreted in multiple ways, answers from a large-scale language model based on each interpretation are presented to the user one by one in sequence. First, if the control unit 1110 determines that the user input is ambiguous through the ambiguity confirmation process described above, for example, if the text input contains a specific ambiguous expression registered in the check DB 19010 and, in particular, if one input can be interpreted in multiple ways, the control unit 1110 analyzes how many interpretations are possible. Here, an example is described in which three interpretations are possible. Next, the control unit 1110 generates instruction statements (instruction statement 1, instruction statement 2, and instruction statement 3) that request an answer from the large-scale language model for each of the three interpretations. Next, the control unit 1110 sends the generated instruction statements, for example, instruction statement 1 first, to the large-scale language model and waits for a response from the large-scale language model.
[0320] Now, we will explain the operation of the large-scale language model. When the large-scale language model receives an instruction sent from the control unit, for example, a response instruction, it generates a response based on this instruction. After that, the answer sentence created by the large-scale language model is sent to the AI response output device 10010.
[0321] When the AI response output device 10010 receives a response (e.g., an answer sentence) created by a large-scale language model, it executes processing on the response of this large-scale language model. As an example of processing the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question. In the processing up to this point, the answer of the large-scale language model to one instruction sentence has been presented to the user, but if there are still instruction sentences for which an answer has not yet been obtained from the large-scale language model, the above processing is repeated until all of them are gone. That is, answer 1 is obtained from the large-scale language model to instruction sentence 1 and presented to the user, then answer 2 is obtained from the large-scale language model to instruction sentence 2 and presented to the user, and then answer 3 is obtained from the large-scale language model to instruction sentence 3 and presented to the user.
[0322] As described above, in the response output system according to Example 1 of the seventh embodiment, when performing the response output process, a process is performed to check the ambiguity of the question or request entered by the user, and when the user input content can be interpreted in multiple meanings, answers from the large-scale language model based on each interpretation are presented to the user one by one in sequence. As a result, even if the true meaning of the user's question or request cannot be identified, it is highly likely that one of the multiple answers obtained from the large-scale language model is what the user wants, making it easier for the user to feel satisfied.
[0323] Example 2 in FIG. 11 illustrates a case in which, when a user input content can be interpreted in multiple ways, answers from a large-scale language model based on each interpretation are presented to the user all at once. First, if the control unit 1110 determines that the user input content is ambiguous through the ambiguity confirmation process described above, for example, if the text input content contains a specific ambiguous expression registered in the check DB 19010 and, in particular, if one input content can be interpreted in multiple ways, it analyzes how many interpretations are possible. Here, an example is described in which three interpretations are possible. Next, the control unit 1110 generates one instruction statement requesting the large-scale language model to provide an answer for each of the three interpretations. Specifically, when the user input content is a question, the control unit 1110 generates an instruction statement requesting the large-scale language model to output answers to question interpretation 1, question interpretation 2, and question interpretation 3 all at once. Next, the control unit 1110 sends the generated instruction statement to the large-scale language model and waits for a response from the large-scale language model.
[0324] Now, moving on to the explanation of the operation of the large-scale language model, when the large-scale language model receives an instruction sent from the control unit, for example, a response instruction, it generates a response based on this instruction. In this example, since one instruction contains multiple requests, it generates answers to each of them, i.e., multiple answers, for example, Answer 1, Answer 2, and Answer 3, and then generates a combined answer sentence. The answer sentence created by the large-scale language model is then sent to the AI response output device 10010.
[0325] When the control unit 1110 of the AI response output device 10010 receives a response (e.g., an answer sentence) created by the large-scale language model, it executes processing on the response of this large-scale language model. As an example of processing on the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question. The answer sentence displayed here includes three answers to each of three interpretations of the question sentence entered by the user, so the user is presented with three answers to the question sentence entered by the user.
[0326] As described above, in the response output system according to Example 2 of the seventh embodiment, when performing the response output process, a process is performed to check the ambiguity of the question or request entered by the user, and when the user input content can be interpreted in multiple meanings, answers from the large-scale language model based on each interpretation are presented to the user all at once. As a result, even if the true meaning of the user's question or request cannot be identified, there is a high possibility that one of the answers obtained from the large-scale language model is what the user wants, making it easier for the user to feel satisfied.
[0327] Example 3 in FIG. 11 illustrates a case in which a user input can be interpreted in multiple ways. Priorities are assigned to each interpretation, and answers are obtained from a large-scale language model in descending order of priority and presented to the user. First, the control unit 1110 analyzes how many possible interpretations are possible when the user input is ambiguous due to the ambiguity confirmation process described above and the user input can be interpreted in multiple ways. Here, an example is described in which three possible interpretations are possible. Next, the control unit 1110 assigns a priority to each of the three possible interpretations. For example, interpretation 1 of the user-input question is assigned a "high" priority, interpretation 2 a "medium" priority, and interpretation 3 a "low" priority. When a user-input question can be interpreted in multiple ways, the highest priority is assigned to the interpretation that is determined to be closest to the user's intent. The definition of how many levels of priority evaluation values to assign (in this embodiment, three levels: "high," "medium," and "low") is stored, for example, in the nonvolatile memory 1108 of the population response output device 10010. Furthermore, the control unit 1110 executes a program stored in, for example, the nonvolatile memory 1108 to interpret the user input content and determine which interpretation to assign a priority to if the user input content can be interpreted in multiple ways. This program has the functions of analyzing how many possible interpretations the user input content can have and determining how close each interpretation is to the user's intention. Furthermore, the program may also incorporate information about each user's text input tendencies and habits, as well as information about the accuracy of past interpretations, so that accurate judgments can be made according to each user's characteristics. The control unit 1110 generates a command statement requesting a response from the large-scale language model for interpretation 1, which has the highest priority among the three interpretations. The control unit 1110 then sends the generated command statement to the large-scale language model and waits for a response from the large-scale language model.
[0328] Now, we will explain the operation of the large-scale language model. When the large-scale language model receives an instruction sent from the control unit, for example, a response instruction, it generates a response based on this instruction. After that, the answer sentence created by the large-scale language model is sent to the AI response output device 10010.
[0329] When the control unit 1110 of the AI response output device 10010 receives a response (e.g., an answer sentence) created by the large-scale language model, it executes processing on the response of this large-scale language model. As an example of processing the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question. In the processing up to this point, the answer of the large-scale language model to one instruction sentence has been presented to the user. However, if there are still instruction sentences for which an answer has not yet been obtained from the large-scale language model, the above processing is repeated until all instruction sentences have been obtained. That is, answer 1 is obtained from the large-scale language model to interpretation 1 with a "high" priority and presented to the user, then answer 2 is obtained from the large-scale language model to interpretation 2 with a "medium" priority and presented to the user, and then answer 3 is obtained from the large-scale language model to interpretation 3 with a "low" priority and presented to the user.
[0330] As described above, in the response output system according to Example 3 of Example 7, when performing a response output process, a process of checking the ambiguity of a question or request input by a user is performed, and when the user input content can be interpreted into multiple meanings, a priority is set for each interpretation, and answers are obtained from the large-scale language model in descending order of priority, i.e., in descending order of those determined to be closest to the intent of the user's question or request, and are presented to the user. As a result, even when the user's question or request is ambiguous and can be interpreted into multiple meanings, an answer from the large-scale language model for an interpretation that is likely to match the intent of the user's question or request can be presented to the user, allowing the user to quickly obtain an answer that matches the intent of their question.
[0331] A specific example of Example 3 of the seventh embodiment will be described with reference to FIG. 12. When a user inputs a request such as "Please write a description of a new product that is low cost and low power, and is an improvement over the current model," the input content can be interpreted in three ways: A, B, and C. In other words, A is an interpretation that both cost and power have been improved, B is an interpretation that only cost has been improved, and C is an interpretation that only power has been improved. Here, for example, priorities are set for the three interpretations. In the example of FIG. 12, the question is whether the phrase "improved" relates to both "low cost and low power," "only low cost," or "only low power." While all of these interpretations are grammatically correct, it is clear that, generally, most people interpret it as relating to both "low cost and low power," while only a few people interpret it as relating to "low power" alone. For example, such tendencies can be registered as evaluation values in the check DB19010, and the control unit 1110 can refer to the tendencies and evaluation values registered in the check DB19010 when determining which interpretation is most likely to represent the true meaning of a sentence entered by a user that can be interpreted in multiple ways.
[0332] The control unit 1110 generates an instruction A' that removes ambiguity from the user input, such as "Please write an introduction about a new product that has been improved to be lower cost and consume less power than the current model," as instruction A' that requests an answer from the large-scale language model for the high-priority interpretation A. This is sent to the large-scale language model, and a response a is obtained from the large-scale language model, such as "The new product has a cost that is 5% lower and a power that is 10% lower than the current model," which is displayed on the display unit 10011.
[0333] Next, for interpretation B with a "medium" priority, an instruction sentence B' is generated that clarifies the instruction content by eliminating ambiguity and is sent to the large-scale language model, and a response b is received from the large-scale language model, which is displayed on display unit 10011. A response is also obtained from the large-scale language model using the same procedure for interpretation C with a "low" priority, and is displayed on display unit 10011.
[0334] In this way, when the user input content can be interpreted in multiple ways, the system requests responses from the large-scale language model in order of the interpretation that is most likely to match the intent of the user's question or request, and presents the received responses to the user, allowing the user to obtain an answer that is most likely to match the intent of their question or request.
[0335] As described above, in the response output system of the seventh embodiment, when a user asks a question or makes a request to a large-scale language model, it checks whether there is ambiguity in the media information such as text or images input by the user, and if it determines that there is ambiguity, it performs a process to resolve the ambiguity. Then, it issues an instruction to request a response from the large-scale language model based on the question sentence, request sentence, and media information that have been subjected to the ambiguity resolution process.
[0336] This makes it easier to obtain a more appropriate answer to a user's question from the large-scale language model. By performing the above-described disambiguation process on a question sentence, it is expected that, for example, when a question sentence is input in which it is unclear what an object referred to by a demonstrative pronoun is, the question sentence can be clarified. The above-described disambiguation process can also be performed on media information such as images. For example, by replacing a blurry image with a clear image, it is expected that the content indicated by the image can be clarified. Therefore, by requesting a response from the large-scale language model based on a question sentence or request sentence that has been disambiguated, it becomes easier to obtain the answer intended by the user from the large-scale language model.
[0337] Furthermore, when a user inputs a question or request that can be interpreted in multiple ways, by requesting a large-scale language model to provide multiple answers corresponding to the multiple interpretations, it is highly likely that one of the multiple answers obtained from the large-scale language model is the answer intended by the user, making it easier to present the answer intended by the user without bothering the user. Also, by setting priorities for each of the multiple interpretations, and requesting answers from the large-scale language model preferentially based on those that are most likely to match the user's intention, and presenting the answers obtained from the large-scale language model to the user, it is easier to present answers that are likely to match the answer intended by the user without bothering the user.
[0338] Example 8 Example 8 of the present invention is an improvement of the AI response output device 10010 or AI response output system described above. In this example, differences from Example 1 will be explained, and repeated explanations of configurations similar to those examples will be omitted.
[0339] The AI response output system according to Example 8 refers to a history of multiple user inputs to the AI response output device 10010. If related content has been input multiple times, the system requests the large-scale language model AI to provide additional information related to the content, and presents the additional information provided by the large-scale language model to the user. More specifically, the user inputs a question or the like to the AI response output device 10010. Upon receiving the input, the AI response output device 10010 generates an instruction for the large-scale language model AI based on the user input, obtains a response from the large-scale language model AI, and presents the response to the user. The system checks the relevance of the user input content every predetermined number of times, and if a relevance is confirmed, presents the user with additional information related to the content entered by the user, separate from the response to the user input. This allows the system to respond to the information the user wants to know while also presenting additional information related to the information the user wants to know, thereby enabling the user to easily acquire more knowledge.
[0340] An example of a response output system according to an eighth embodiment of the present invention will be described with reference to FIG. 7. As shown in FIG. 7, the configuration of the response output system according to the eighth embodiment is basically the same as that according to the first embodiment. In the eighth embodiment, an AI response output device 10010 is also communicatively connected to a large-scale language model server 19001 or a large-scale language model server 20001 equipped with a large-scale language model, and a second server 19002 different from these servers, via the Internet 19000. The above servers may include a check database. In FIG. 7, the second server 19002 includes a check database (check DB) 19010. However, in the eighth embodiment, the second server 19002 includes a database other than the check database (check DB) 19010, such as a vocabulary database or a specialized field vocabulary database specialized in a specialized field.
[0341] In the eighth embodiment, a configuration will be described in which the large-scale language model server 19001 is equipped with a large-scale language model (LLM). In the following description, the large-scale language model equipped in the large-scale language model server 19001 may be simply referred to as the large-scale language model or LLM. Incidentally, the large-scale language model may be a large-scale language model on a network, or a local large-scale language model (such as the local LLM processing unit 10028) equipped in the AI response output device 10010.
[0342] In the above configuration, when a user asks a question or makes a request to a large-scale language model, which is an artificial intelligence, and attempts to obtain a response or answer, the process is, for example, as follows: First, the user inputs the question or request to the large-scale language model to the AI response output device 10010. As an example, the user operates the operation input unit 1107, which is an input unit, to input a question or request, which is text information in a natural language. The question / request input to the large-scale language model may be operated using the operation input unit 1107 or may be input by voice. Furthermore, if the large-scale language model is a multimodal LLM that can handle data in various formats, the user may also input various media information that the large-scale language model can handle, such as images, videos, and audio. The AI response output device 10010 generates a response instruction sentence (prompt) for the large-scale language model based on the question, request, and media information input by the user, and transmits the generated response instruction sentence (prompt) to the large-scale language model server 19001 that includes the large-scale language model. When the large-scale language model generates a response to the response instruction sentence (for example, an answer sentence to a user's question), the generated response is sent from the large-scale language model server 19001 to the AI response output device 10010.
[0343] Upon receiving this response, the AI response output device 10010 processes the response of the large-scale language model (LLM) for the user. As an example of processing the response of the large-scale language model, the AI response output device 10010 displays a response sentence (also called an answer sentence) generated by the large-scale language model on the display unit 10011. In this manner, the user can obtain a response from the large-scale language model to their question or request.
[0344] Here, when a user wants to gain knowledge about a certain theme, the user may input questions about that theme in multiple installments into the AI response output device 10010. For example, after asking a question about word A related to a theme and receiving a response to that question from the large-scale language model AI, the user may then ask a question about word B related to the same theme, and then ask about word C, and so on, gradually acquiring knowledge about the theme. This is because the user is able to input these as questions because he or she already knows that words A, B, and C are related to the theme. If the user does not know other words related to the theme (for example, word D), he or she will not input a question about word D, and will not be able to acquire knowledge about word D.
[0345] Therefore, in the response output system according to Example 8, when a user inputs a question or request to the AI response output device 10010 a predetermined number of times in succession, the system checks whether the user input content is relevant. If relevant, the system generates an instruction statement instructing the large-scale language model AI to present words related to the user input content that the user has not yet input. By presenting the response obtained from the large-scale language model AI to the user in this way, the user can obtain additional information related to the input word. Furthermore, by applying this mechanism, it is possible to not only present the user with words related to the content input by the user, but also to generate a single question that encompasses the contents of the three questions, along with individual responses to the three questions, when the contents of Question 1, Question 2, and Question 3 input by the user are related to each other, and present the response to that question to the user. This allows the user to obtain a comprehensive response that covers the three things they wanted to know, thereby gaining deeper knowledge than if they had received three individual responses. Furthermore, when the contents of Question 1, Question 2, and Question 3 entered by the user are related to each other, a single question can be generated that includes the three questions and related information not included in them, and a response to that question can be presented to the user. This allows the user to obtain a response that not only contains what they wanted to know, but also adds information they did not know, thereby gaining a wider and deeper knowledge.
[0346] The method for inputting a question or request from a user to the AI response output device 10010 may include text input using the operation input unit 1107 as an input unit, or a method of converting a voice-input question using the microphone 1139 into text. Media information such as an image may be input by selecting a media information file using the operation input unit 1107. The process of checking the relevance of a question or request, the process performed when there is relevance, the process of generating an instruction statement, and the process of outputting an answer from the large-scale language model to the display unit 10011 are executed by, for example, the control unit 1110. In this embodiment, the control unit 1110 generates a response instruction statement for the large-scale language model based on the question statement and acquires a response generated by the large-scale language model in response to the response instruction statement. Specifically, the control unit 1110 generates a response instruction statement based on the input question or request, displays the response obtained from the large-scale language model, and further checks the relevance of the input question or request. If it is determined to be relevant, the control unit 1110 generates a response instruction statement requesting additional information related to the question or request, and displays the additional information as a response obtained from the large-scale language model.
[0347] Hereinafter, a detailed description will be given of a response output process in a response output system according to an eighth embodiment, mainly a process flow for checking the relevance of a question or request input by a user multiple times, performing a process depending on the presence or absence of the relevance, and obtaining a response from a large-scale language model to the question or request. FIG. 13A is a diagram illustrating an example of a process flow of the response output system according to the eighth embodiment. FIG. 13B is a diagram illustrating an example of a process flow of an AI response output device according to the eighth embodiment. FIG. 14 is a diagram illustrating the concept of the response output system according to the eighth embodiment. FIG. 15 is a diagram illustrating a detailed example of a relevance checking process and a process step when there is relevance in the process flow of the response output system according to the eighth embodiment. Furthermore, FIGS. 16 and 17 are diagrams illustrating an example of the operation of the response output system according to the eighth embodiment. In addition, in the diagrams illustrating the process flow of the response output system, the same steps are assigned the same symbols, and redundant explanations of each step may be omitted.
[0348] First, an outline of the processing flow of the response output system according to Example 8 will be described with reference to Fig. 13A. As shown in Fig. 13A, in step S1009, the AI response output device 10010 receives input such as a question or request from a user, transmits a response instruction sentence for the input to a large-scale language model, receives a response from the large-scale language model and displays it on a display unit 10011, and also performs a relevance confirmation process when input is received from the same user a predetermined number of times or more in succession.
[0349] 13A and 13B, the detailed operation of the response output system according to the eighth embodiment will be described. FIG. 13B is a processing flow illustrating the operation of step S1009 in detail. As shown in FIG. 13B, first, a user operates, for example, the operation input unit 1107 to input a question or request to the large-scale language model. In step S1001, the control unit 1110 receives the user input (question or request). In step S1007, the control unit 1110 responds to the user input. Specifically, a response instruction sentence based on the user input is generated, a response request for the instruction sentence is sent to the large-scale language model, the response received from the large-scale language model is displayed on the display unit 10011, and the next user input is awaited. In addition, when the same user inputs a question or request consecutively a predetermined number of times or more within a predetermined period, a relevance confirmation process is executed. The predetermined period may be, for example, the same day (between midnight and midnight), a single login period (from login to logoff), or the elapsed time since the first question or request was input (for example, within 30 minutes of the first question or request being input). Next, in step S1008, the control unit 1110 generates a request statement (hereinafter referred to as a related information request statement) for requesting additional information related to the question or request entered by the user as a result of this relevance confirmation process, if necessary.
[0350] Then, based on the related information request statement, an instruction statement (prompt) is generated for the large-scale language model (LLM) included in the large-scale language model server 19001. In this example, a response instruction statement is generated to instruct the creation of an answer statement as a response to the user's question. Next, this response instruction statement is sent to the large-scale language model (step S1004). Note that the above response instruction statement may also be simply referred to as a prompt.
[0351] Next, as shown in FIG. 13A (steps S1101 to S1103), when the large-scale language model (LLM) receives an instruction sentence, for example, a response instruction sentence, sent by the AI response output device 10010 (step S1101), it generates a response based on this instruction sentence (step S1102). In the example shown in FIG. 13B, since the instruction sentence is a response instruction sentence, the large-scale language model creates a response sentence to the user's question as the response generation in step S1102. Thereafter, in step S1103, the response sentence created by the large-scale language model is sent to the AI response output device 10010.
[0352] 13B, when the AI response output device 10010 receives a response (e.g., an answer sentence) created by the large-scale language model (step S1005), it executes processing on the response of this large-scale language model in step S1006. As an example of processing on the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question.
[0353] As described above, in the response output system according to the eighth embodiment, when a question or request is input by the same user a predetermined number of times or more in a predetermined period of time, a relevance check process is performed to display a response to the question or request input by the user, along with additional information related to the question or request. This allows the user to receive not only the answer to the question but also the additional information about the question, thereby enabling the user to understand the question in more detail.
[0354] Next, the process of checking the relevance of user input contents in step S1007 and the process of checking when there is a relevance in step S1008 will be described in more detail. For example, as shown in FIG. 13B, when control unit 1110 receives user input in step S1001, control unit 1110 then checks in step S1007 whether the user has consecutively input questions, requests, etc., and checks whether the consecutively input contents are relevant, and determines whether it is necessary to execute the process when there is a relevance. Specific examples of the relevance checking process will be described later, but they may include checking the relevance of questions or requests made by text consecutively input by the user, or checking the relevance of the content of media information consecutively input by the user.
[0355] The processing according to the determination result of this confirmation processing is performed as processing when there is correlation. For example, if the user input is performed consecutively a predetermined number of times or more, but it is determined that the input content at that time is not related, no special processing is performed as processing when there is correlation in step S1008. On the other hand, if the user input is performed consecutively a predetermined number of times or more and it is determined that the input content at that time is related, processing when there is correlation is performed in step S1008, and a command statement requesting additional information related to the user input content is generated.
[0356] Next, the user input response and relevance confirmation process of step S1007 will be specifically described with reference to FIG. 14. When the control unit 1110 receives, for example, User question (1) 10053A as user input, it generates an instruction statement instructing the generation of a response to the question, sends this to the large-scale language model, obtains a response, and displays it as AI answer (1) 10063A. The user then inputs User question (2) 10053B, and the control unit 1110 generates an instruction statement instructing the generation of a response to the question, sends this to the large-scale language model, and outputs AI answer (2) 10063B as the response. Next, the user inputs User question (3) 10053C, and the control unit 1110 generates an instruction statement instructing the generation of a response to the question, sends this to the large-scale language model, and outputs AI answer (3) 10063C as the response. During this process of user input and response, the control unit 1110 performs relevance confirmation processing of the user input content based on a set of input content relevance confirmation parameters.
[0357] The input content relevance confirmation parameter group 20010 is stored, for example, in the nonvolatile memory 1108...
Claims
1. an input unit into which a question is input by a user; a control unit that generates a response instruction sentence for a large-scale language model based on the question sentence and acquires a response sentence generated by the large-scale language model in response to the response instruction sentence; an output unit that performs output based on the response sentence acquired by the control unit, The control unit performing a clarification process for the question sentence, and generating the response instruction sentence based on the question sentence after the clarification process; Response output system.
2. 2. The response output system according to claim 1, The control unit performing a check process to confirm whether the question contains a predetermined specific expression; In the checking process, if the specific expression is included in the question sentence, a clarification process is performed on the question sentence. Response output system.
3. 3. The response output system according to claim 2, The control unit In the check process, if the question sentence can be interpreted into a plurality of meanings, the response instruction sentence is generated for each interpretation of the question sentence. Response output system.
4. The response output system according to claim 1, media information can be input to the input unit; The control unit When the media information is input by the user, performing a media information check process to check whether the media information satisfies a preset media information evaluation value; In the media information check process, when the media information does not satisfy the preset media information evaluation value, a clarification process is performed on the media information. Response output system.
5. an input unit into which a question is input by a user; a control unit that generates a response instruction sentence for a large-scale language model based on the question sentence and acquires a response sentence generated by the large-scale language model in response to the response instruction sentence; an output unit that performs output based on the response sentence acquired by the control unit, The control unit When the question sentence is input multiple times, a process of checking the relevance of the question sentences is performed; generating additional information related to the question sentences when the question sentences are related to each other in the relevance check process; Response output system.
6. 6. The response output system according to claim 5, The control unit If the question is entered a predetermined number of times or more, The relevance check process is performed. Response output system.
7. 6. The response output system according to claim 5, The control unit In the relevance check process, Refer to groups of words that are grouped into multiple groups, Check whether the question contains any words belonging to a specific group among the plurality of groups; If at least two of the multiple input question sentences contain a phrase that belongs to the same specific group, the question sentences are determined to be related to each other. Response output system.
8. The response output system according to claim 7, The control unit If at least two of the multiple input question sentences contain a phrase belonging to the specific same group, As additional information related to the above question, performing a process of generating a phrase that is different from the phrase included in the question sentence from among the phrases that belong to the same specific group; Response output system.
9. The response output system according to claim 7, The control unit For at least two questions that are determined to be related among the questions, A process to generate a single question that encompasses the content of the question and a response to it. Response output system.
10. an input unit into which a question is input by a user; a log information input unit to which log information is input; a control unit that generates a response instruction sentence for a large-scale language model based on the question sentence and acquires a response sentence generated by the large-scale language model in response to the response instruction sentence; an output unit that performs output based on the response sentence acquired by the control unit, The control unit recording a history of instructions and responses between the user and the large-scale language model as log information; When the user and the large-scale language model ask and answer questions, Check whether the log information relating to the user is recorded; If the log information relating to the user is recorded, a process of inputting the log information into the log information input unit is performed. Response output system.
11. The response output system according to claim 10, A memory for temporarily recording information; a non-volatile memory for recording information; The control unit recording in the memory a history of instructions and responses between the user and the large-scale language model; performing a process of classifying the contents of the history; assigning a classification label based on the classification of the history; recording the history including the classification label as log information in a nonvolatile memory; Response output system.
12. The response output system according to claim 11, The control unit confirming whether the question contains a word or phrase that matches a classification label assigned to the log information; If the question contains a phrase that matches a classification label assigned to the log information, inputting log information to which a classification label that matches the phrase has been assigned to the log information input unit; Response output system.
13. The response output system according to claim 11, The control unit When recording the log information, creating and recording log management information for managing the log information for each user; Response output system.
14. The response output system according to claim 10, The control unit When the user and the large-scale language model ask and answer questions, Check whether the log information relating to the user is recorded; If the log information relating to the user is recorded, At least one piece of log information is selected by the user from the log information, and the selected piece of log information is input to the log information input unit. Response output system.
15. The response output system according to claim 10, The control unit A state to be set in the log information, a state for changing a summarization rate of the log information based on the priority of the contents of the log information; Response output system.
Citation Information
Patent Citations
Structural unit for tank construction
JP1977008512A