Response output device
The response output device improves user interaction with AI by generating response instruction sentences through a conversion process, enhancing the effectiveness of AI response generation.
Patent Information
- Application Number
- PCT/JP2024/039136
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2024-11-01
- Publication Date
- 2025-07-17
AI Technical Summary
Existing response output technologies using artificial intelligence do not adequately consider user interaction and configuration for optimal response generation.
A response output device equipped with an input unit, control unit, and output unit, which includes a conversion process to generate response instruction sentences for a large language model, allowing for improved interaction and response generation.
Enhances the suitability of response output technology by providing more effective and user-friendly interactions with artificial intelligence systems.
Smart Images

Figure JP2024039136_17072025_PF_FP_ABST
Abstract
Description
Response output device
[0001] The present invention relates to a response output device.
[0002] A response output technology using artificial intelligence such as a language model is disclosed in, for example, Patent Document 1.
[0003] Special table 2019-528512 publication
[0004] However, the disclosure of Patent Document 1 does not sufficiently consider a configuration for more suitably providing a response output technology using artificial intelligence to a user.
[0005] An object of the present invention is to provide a more suitable response output technique.
[0006] In order to solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problem, and one example thereof may be a response output device including: an input unit into which a question sentence is input by a user; a control unit that generates a response instruction sentence for a large-scale language model based on the question sentence and acquires the response sentence generated by the large-scale language model in response to the response instruction sentence; and an output unit that performs output based on the response sentence acquired by the control unit, wherein the control unit executes a conversion process for the question sentence and generates the response instruction sentence based on the question sentence after the conversion process.
[0007] According to the present invention, a more suitable response output technique can be provided. Other problems, configurations, and effects will become clear in the following description of the embodiments.
[0008] FIG. 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention. FIG. 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 2 is a diagram illustrating an example of the operation of the artificial intelligence response output device and system according to an embodiment of the present invention. FIG. 3 is an explanatory diagram of an example of a character conversation device and character conversation system according to an embodiment of the present invention. FIG. 4 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. FIG. 5 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. FIG. 6 is an explanatory diagram of an example of a conversation in the character conversation device and character conversation system according to an embodiment of the present invention. FIG. 7 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. FIG. 8 is an explanatory diagram of an example of the operation of the character conversation device and character conversation system according to an embodiment of the present invention. Fig. 1 is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention; Fig. 2 is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention; Fig. 3 is an explanatory diagram of an example of a conversation in a character conversation device and a character conversation system according to an embodiment of the present invention; Fig. 4 is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention.FIG. 1 is an explanatory diagram of an example of the operation of a character conversation device and a character conversation system according to an embodiment of the present invention. ... an artificial intelligence response output device according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of a display of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of a response generation process of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 1 is a diagram of an example of a response output system according to an embodiment of the present invention. FIG. 1 is a diagram of an example of a processing flow in a response output system according to an embodiment of the present invention. FIG. 1 is a diagram of an example of a processing flow in a response output system according to an embodiment of the present invention. FIG. 1 is a diagram of an example of a check database according to an embodiment of the present invention. FIG. 1 is a diagram of an example of a processing flow in a response output system according to an embodiment of the present invention. FIG. 1 is a diagram of an example of a database matching priority table according to an embodiment of the present invention. FIG. 1 is a diagram showing an example of a processing flow in a response output system according to an embodiment of the present invention. FIG. 2 is a diagram showing an example of a processing flow in a response output system according to an embodiment of the present invention. FIG. 3 is a diagram showing an example of a processing flow in a response output system according to an embodiment of the present invention. FIG. 4 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 5 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 6 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention.FIG. 1 is a diagram showing an example of a processing flow of a response output system according to an embodiment of the present invention. FIG. 1 is a diagram showing an example of a processing flow of a response output system according to an embodiment of the present invention. FIG. 1 is a diagram showing an example of a display screen of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 1 is a diagram showing an example of a processing flow of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 1 is a diagram showing an example of a display screen of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 2 is a diagram showing an example of a display screen of an artificial intelligence response output device according to an embodiment of the present invention. FIG. 3 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 4 is an explanatory diagram of an example of an operation of a response output system according to an embodiment of the present invention. FIG. 5 is a diagram showing an example of a processing flow of a response output system according to an embodiment of the present invention. FIG. 6 is a diagram showing an example of a processing flow of a response output system according to an embodiment of the present invention.
[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical ideas disclosed in this specification. Furthermore, in all drawings used to explain the present invention, components having the same functions are given the same reference numerals, and repeated explanations thereof may be omitted.
[0010] Note that if the AI response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. If the AI response output device has an audio output function, it may be referred to as an audio output device, or simply as an information processing device. A system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI response output system. Furthermore, if the AI response output device provides a user with a response service based on a large-scale language model, which is an AI, and is helpful to the user, the AI response output device or the display output of the AI response output device can serve as an AI (AI) assistant for the user. Therefore, in this case, the AI response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI assistant system or an AI assistant display system. Furthermore, in this case, the AI response output device serves as an interface between the user and the AI, and therefore may be referred to as an AI interface device. In this case, a system including an AI response output device and a large-scale language model server that stores a large-scale language model may be referred to as an AI interface system.
[0011] First Embodiment As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.
[0012] 1A, an example of an AI response output device 10010 of the present invention will be described. In addition, in the case where the AI response output device 10010 cooperates with a large-scale language model server 19001 via communication or the like, an example of a system in which the AI response output device 10010 includes the large-scale language model server 19001 and / or a multimodal large-scale language model server 20001 will be described.
[0013] In the example of FIG. 1A, the AI response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat display, a screen that projects an image from the rear, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. The display unit 10011 may also be provided with a touch operation input sensor and configured as a touch panel.
[0014] 1A, the audio output unit 1140 of the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139 that can pick up the user's voice. By audio input from the microphone 1139 or user operation input via an operation input unit (described later), the AI response output device 10010 can acquire user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.
[0015] The AI response output device 10010 may be provided with a local large-scale language model within the AI response output device 10010 itself. In this case, the response of the large-scale language model may be output as a display output from the display unit 10011 and / or as an audio output from the audio output unit 1140.
[0016] In addition, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and / or as an audio output on the audio output unit 1140.
[0017] Alternatively, the AI response output device 10010 may also include a local large-scale language model and may be configured to communicate with an external large-scale language model server 19001 having the large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the AI response output device 10010 may switch between a response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001, and output either one as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140. Alternatively, a response generated based on both the response from the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and / or a voice output from the voice output unit 1140.
[0018] The configuration when the AI response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A , the communication between the communication unit 1132 and the communication device 19011 is shown as being wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. Furthermore, the communication path from the communication unit 1132 to the Internet 19000 may include wired and wireless portions, or may go via a router or repeater. The AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.
[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" may be considered to refer to the local large-scale language model provided by the AI response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
[0020] The example of Figure 1A shows an example in which the display unit 10011 displays elements in two display areas: a prompt display area 10051 in which a user inputs a prompt to a large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which a response from the large-scale language model is displayed. In the example of Figure 1A, the prompt display area 10051 displays an icon 10052 indicating a user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt. In the example of Figure 1A, the artificial intelligence response display area 10061 displays an icon 10062 indicating an artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence. The display example of the display unit 10011 of the AI response output device 10010 shown in Fig. 1A is merely an example. Depending on the implementation example in which the AI response output device 10010 is used, a display different from the example shown in Fig. 1A may be performed.
[0021] Here, large-scale language models will be described. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models have been published, such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT. These technologies may be used in this embodiment as well. Note that these large-scale language models are artificial intelligence models generated by large-scale pre-training on the natural language contained in numerous documents and texts existing in the human world. The number of parameters in artificial intelligence models exceeds 100 million. In addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of a base model is a model called a Transformer. Reference 1, for example, has been published as an example of learning these models.
[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. Advanced models are capable of natural language question answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation. Because these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for a specific application is extremely resource-inefficient. Therefore, models are generated through large-scale pre-training as foundation models applicable to various applications. For example, the large-scale language model server 19001 shown in FIG. 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface). Furthermore, the AI response output device 10010 shown in FIG. 1A may be equipped with a local large-scale language model and configured to be used by the AI response output device 10010 itself. The learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided in the large-scale language model server 19001, the AI response output device 10010, etc. In this way, instead of performing pre-learning for each application or each terminal, replicating the large-scale language model that is the base model generated by large-scale pre-learning and using it on individual servers or terminals allows the resources used for learning to be shared, resulting in good resource efficiency.
[0024] Furthermore, even if a large-scale language model is used as a base model generated through large-scale pre-training, it may be configured so that additional learning such as transfer learning is performed on individual servers or devices depending on the application or purpose.
[0025] Furthermore, large-scale language models can be pre-trained on natural languages and perform input / output processing targeting natural languages. Furthermore, multimodal large-scale language model artificial intelligence capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention. In FIG. 1A, a large-scale language model server 20001 is shown, which is a server having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model artificial intelligence include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies may also be used in this embodiment. These multimodal large-scale language models are artificial intelligence models generated by large-scale pre-training on natural language and types of information other than natural language text information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. In addition, there are also models that undergo reinforcement learning based on human feedback. Hereinafter, types of information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.
[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0027] Next, using Figure 1B, we will explain an example configuration of an artificial intelligence response output device 10010 that accepts input from a user to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as a large-scale language model to the input from the user.
[0028] The AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, etc. The AI response output device 10010 may have a large screen, such as a monitor or television.
[0029] The display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be referred to as a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the artificial intelligence.
[0030] The communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 1132 of the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000. Note that the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the case of a wired connection, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN communication method. This allows the AI response output device 10010 to communicate with various servers connected to the Internet 19000.
[0031] The AI response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.
[0032] The power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the DC current required by each component of the AI response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each component requiring power via the external power supply input interface 1111 when power is not supplied from the outside.
[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller, and inputs a signal for an operation different from a user's touch operation on the touch operation input sensor of the display unit 10011. Separate from a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the AI response output device 10010. By the user's operation input via the operation input unit 1107, the AI response output device 10010 can acquire user input that serves as the basis for instructions (prompts) for the large-scale language model, which is the AI. A modified configuration is also possible in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107.
[0034] The video signal input unit 1131 connects to an external video output device and inputs video data. The video signal input unit 1131 may be configured with various digital video input interfaces. For example, the video signal input unit 1131 may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be configured with various USB interfaces.
[0035] The audio signal input unit 1133 is connected to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 1133 may also be various USB interfaces, etc. In the case of an HDMI-standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface in which a terminal and a cable are integrated.
[0036] The audio output unit 1140 can output audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 can also output audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured with a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
[0037] The microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 1110, which will be described later, performs voice recognition processing on the generated audio signal to acquire text information from the audio signal. By using audio input from the microphone 1139, the AI response output device 10010 can acquire user input that serves as the basis for instructions (prompts) to the large-scale language model, which is the AI.
[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 of the AI response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.
[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor device memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 1170 before product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from an external device, an external server, or the like, via the communication unit 1132. The video data, image data, and the like recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, and the like recorded in the storage unit 1170 may also be output to an external device, an external server, or the like via the communication unit 1132.
[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 performs, for example, video switching control, such as determining which video signal to input to the display unit 10011 between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform control to perform image processing on the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109. Examples of image processing include scaling processing, which enlarges, reduces, or deforms an image; brightness adjustment processing, which changes the brightness; contrast adjustment processing, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0041] The attitude sensor 1113 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the AI response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each unit connected thereto.
[0042] The non-volatile memory 1108 stores various data used by the AI response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, etc. The control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.
[0043] The local LLM processing unit 10028 includes a memory capable of storing a large-scale language model (LLM) and can execute inference of the large-scale language model under the control of the control unit 1110. The hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the AI response output device 10010.
[0044] The control unit 1110 controls the operation of each connected unit. The control unit 1110 may also cooperate with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit in the AI response output device 10010. The control state of the control unit 1110 includes, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028 or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, such as a speaker.
[0045] When a user inputs via the touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and transmitted to the local large-scale language model of the local LLM processing unit 10028 provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. All of these controls for obtaining a response from these large-scale language models can be performed by the control unit 1110.
[0046] The storage unit 1170 may also store a fixed response phrase database (which may also be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction statements from the AI response output device 10010. The control unit 1110 may control the generation of responses to be output using data stored in the fixed response phrase database. FIG. 1C shows an example of the fixed response phrase database. In the example of FIG. 1C, fixed responses to be output by the AI response output device 10010 are stored for each condition assigned a condition number. For example, as in condition number 1, when the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, a response may be output using a fixed response phrase such as "Good morning" or "Today is ____ day of ____ month, isn't it?" The ____ part of "____ day of ____ month" may be generated using information stored in the memory 1109 or the like of the AI response output device 10010.
[0047] Furthermore, in the example of the standard response phrases in the database shown in FIG. 1C, if multiple standard response phrases separated by / are stored, the control unit 1110 may control the output of a response by randomly selecting one of the standard response phrases using a random number or the like. This can eliminate or improve the situation where responses under the same conditions become monotonous. The explanation for the example of condition number 1 is the same for the examples of condition numbers 2, 3, and 4. The control unit 1110 may control the output of the standard response phrase of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.
[0048] Next, an example of condition number 5 shown in FIG. 1C will be described. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of a user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language or when the user input contains an obvious grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch what you said" or "I might not know about that." By responding in this manner, the user can be prompted to input again, and the system can wait for a corrected user input.
[0049] Next, an example of condition number 6 shown in Figure 1C will be described. Condition number 6 is an example of a case where the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in Figure 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the standard response phrase "It seems to be working poorly." By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is malfunctioning, and to prompt the user to take action on the error, etc.
[0050] The AI response output device 10010 may output a response using the fixed response phrase database (fixed response phrase DB) described with reference to Fig. 1C instead of a response from a large-scale language model such as the local large-scale language model provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001. Alternatively, the AI response output device 10010 may output a response that combines the responses from these large-scale language models with a response using the fixed response phrase database (fixed response phrase DB).
[0051] 1C described above may be stored in the storage unit 1170 and used by the control unit 1110 of the AI response output device 10010. However, the fixed response database (fixed response DB) shown in FIG. 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response database (fixed response DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the fixed response database (fixed response DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective server. In this way, even if the artificial intelligence response output device 10010 is not equipped with a standard response phrase database (standard response phrase DB), it is possible to generate a response using the standard response phrase database (standard response phrase DB).
[0052] In the above description, the AI response output device 10010 has been described as having a display panel with a display screen using fixed pixels. This concept may also include a projection type image display device (projector) in which a projection optical system is provided behind the display panel with a display screen using fixed pixels, and an optical image of the image on the display panel of the display screen is projected onto a screen or wall.
[0053] 1A and 1B, an example has been described in which the AI response output device 10010 includes the display unit 10011. However, the AI response output device 10010 according to an embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if the display unit 10011 is not included, the AI may be configured to accept input from a user to the AI via the voice signal input unit 1133 or the microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the user input via the voice output unit 1140.
[0054] According to the artificial intelligence response output device and artificial intelligence response output system of the first embodiment of the present invention described above, it is possible to accept input from a user to an artificial intelligence such as a large-scale language model, and output a response to the input from the user that is generated by inference by the artificial intelligence, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the artificial intelligence response output device itself.
[0055] Next, as a second embodiment of the present invention, an example will be described in which the AI response output device 10010 described in the first embodiment is connected to the Internet, and is connected to a server equipped with a large-scale language model AI via the Internet to perform operation. In this embodiment, differences from the first embodiment will be described, and repeated explanations of configurations similar to those of the first embodiment will be omitted.
[0056] An example of the connection state between the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment of the present invention will be described with reference to Figure 2A. The AI response output device 10010 according to the second embodiment may be called a character conversation device. Furthermore, a system including the AI response output device 10010 and the large-scale language model server 19001 according to the second embodiment may be called a character conversation system. An image of a character 19051 is displayed on a display unit 10011 displayed by the AI response output device 10010. The image of the character 19051 is an image generated by rendering a 3D model of the character in a virtual space.
[0057] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can serve as an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this embodiment may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0058] In the example of FIG. 2A , the audio output unit 1140 of the AI response output device 10010 is composed of a speaker. The AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. The AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A , an example of wireless communication between the communication unit 1132 and the communication device 19011 is shown, but wired communication is also acceptable. The communication path from the communication unit 1132 to the Internet 19000 may include both wired and wireless portions. The AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with a second server 19002 different from the large-scale language model server 19001 via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 may be considered as a single system.
[0059] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2B. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Note that Figure 2B omits illustration of communication paths such as the Internet 19000 shown in Figure 2A. Figure 2B also illustrates a user 230 of the artificial intelligence response output device 10010.
[0060] Here, we will explain the sequence of operations of the AI response output device 10010. The AI response output device 10010 loads a character movement program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character movement program, thereby realizing the various processes described below.
[0061] First, the AI response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is in natural language. Note that the extraction of the text of the words spoken by the user 230 may be performed continuously for all words, or may start when the user utters a word within a predetermined period after a trigger keyword. For example, the trigger keyword may be when the user utters "hello" followed by a character name. For example, if the name of the character 19051 is "Koto," then "Hello, Koto!" may be the trigger keyword.
[0062] The character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230 and transmits the prompt to the large-scale language model server 19001 using an API. Here, the prompt may be metadata containing information written in a notation using tags, such as the markup format of a markup language, a notation using predetermined symbols, such as Markdown, or an object notation of a predetermined script, such as JSON. The prompt stores natural language text information as a main message. The types of prompts transmitted from the AI response output device 10010 to the large-scale language model server 19001 include setting prompts that store instructions such as initial settings, and user prompts that reflect instructions from the user. Type identification information identifying whether the prompt is a setting prompt or a user prompt may be stored in a portion of the prompt other than the main message. When the character operation program of the AI response output device 10010 creates a prompt based on the text of the words spoken by the user 230, it creates the user prompt and sends it to the large-scale language model server 19001.
[0063] Next, the large-scale language model of the artificial intelligence of the large-scale language model server 19001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response including natural language text information based on the result. The large-scale language model server 19001 uses an API to send the response to the artificial intelligence response output device 10010. The response stores natural language text information as a main message. Here, the response may be metadata storing information written in the same format as the instruction (e.g., a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as Markdown, or an object notation of a predetermined script such as JSON). If the response uses the same format as the instruction, type identification information may be stored outside the main message to indicate that the information is a different type from the initial setting instruction and the user instruction. For example, information indicating that the response is a response from the large-scale language model may be stored.
[0064] Next, the AI response output device 10010 receives a response from the large-scale language model server 19001 and extracts the natural language text information stored as the main message in the response. The character operation program of the AI response output device 10010 uses speech synthesis technology to generate a natural language voice as a response to the user based on the natural language text information extracted from the response, and outputs the voice from the speaker, i.e., the voice output unit 1140, so that it sounds as if it were the voice of the character 19051. This process may also be referred to as the character's "speaking."
[0065] Conversation examples 1 to 5 in Fig. 2C show specific examples of the response voice of the character 19051 in response to words from the user 230, as a result of the processing by the AI response output device 10010 and the large-scale language model server 19001 described above. In this way, the user 230 can have a conversation with the character 19051 as if it were a real person.
[0066] 2B or a system including the AI response output device 10010, there is no need to install a large-scale language model, which requires vast amounts of data and computational resources for learning, in the AI response output device 10010. In addition, the advanced natural language processing capabilities of the large-scale language model can be utilized via an API, and when a user speaks to a character, a more appropriate response can be given to the user, enabling a more appropriate conversation to be held.
[0067] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2D. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2D shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0068] FIG. 2D also shows the exchange of instructions and responses in chronological order, from the display setting instruction, the first round of user instructions and their responses to the fourth round of user instructions and their responses.
[0069] As shown in FIG. 2D , the setting instruction can be used to initially instruct the large-scale language model of the artificial intelligence of the large-scale language model server 19001 on the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, and so on. The user's name can also be understood as the initial setting. This allows the large-scale language model to generate responses from the first round onward while respecting the role. Then, to a user who hears the voice of the character 19051 based on the responses from the first round onward, the character 19051 will seem to have the setting and personality of the person described in the setting instruction. Furthermore, the large-scale language model server 19001 according to this embodiment is equipped with a memory that stores the content of the conversation until the end of the series of conversations, and is configured to store a series of user instructions and their responses before generating responses. This allows for the realization of a conversation such as that shown in FIG. 2D .
[0070] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2E shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0071] 2E shows an example of a case where, after the series of conversations shown in FIG. 2D has ended, the user 230 speaks to the character 19051 again to start a new conversation. In FIG. 2E, the exchange of instructions and responses is shown in chronological order, from the first round of user instructions and their responses to the third round of user instructions and their responses.
[0072] Here, "termination" of the "continuation of a series of conversations" refers to a process in which, when a predetermined condition is met, the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversation that was maintained while the series of conversations was continuing. An example of the predetermined condition is, for example, when the AI response output device 10010 issues an instruction to the large-scale language model server 19001 to "terminate" the "continuation of a series of conversations." Another example of the predetermined condition is, for example, when a predetermined time or more has passed since the AI response output device 10010 stopped sending instruction statements for the series of conversations to the large-scale language model server 19001 (timeout). Another example of the predetermined condition is when, after authentication processing is performed in the connection between the AI response output device 10010 and the large-scale language model server 19001, the authentication processing is terminated due to factors such as communication disconnection or the AI response output device 10010 being powered off while exchanging the instruction statements and responses.
[0073] Note that when the "continuation of a series of conversations" is "ended," the large-scale language model server 19001 erases from the large-scale language model server 19001 the memory of the conversations that it had maintained while the series of conversations was continuing. Therefore, even though the conversation shown in FIG. 2E is after the series of conversations shown in FIG. 2D , the server response to the user instruction is a response with content that does not include any memory of the character name set in the large-scale language model, the role to be played, the characteristics of the conversation, the user's name, etc., that were included in the setting instruction shown in FIG. 2D . Similarly, the conversation shown in FIG. 2E is a response with content that does not include any memory of the series of conversations shown in FIG. 2D . In other words, with the "end" of the "continuation of a series of conversations" shown in FIG. 2D , the conversation in FIG. 2E starts from a state in which the large-scale language model of the artificial intelligence of the large-scale language model server 19001 is initialized.
[0074] This causes the user 230 to feel as if the character 19051 has lost its memory or is a completely different person. From the user 230's perspective, the character's response feels very strange, resulting in a feeling of loneliness and disappointment. Such behavior poses a problem in that it is not possible to ensure the consistency of the settings and memories of the character 19051, such as its name, role, conversational characteristics, and personality, displayed on the AI response output device 10010.
[0075] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2F. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2F shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0076] FIG. 2F illustrates an example of a case in which the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations illustrated in FIG. 2D has ended. Unlike the process of FIG. 2E , in the process of FIG. 2F , when starting a new conversation, the AI response output device 10010 sends a setting instruction as the first instruction to the large-scale language model server 19001. The setting instruction stores the same natural language text as the initial setting instruction of FIG. 2D . This may be referred to as reset text. The setting instruction is followed by natural language text describing the history of past conversations. This may be referred to as conversation history text. The AI response output device 10010 may record the history of past conversations as natural language text information in the storage unit 1170 while the series of conversations illustrated in FIG. 2D is continuing. If there are conversations on different dates, each conversation may be recorded in association with date and time information, and the conversation history may be accumulated. When generating a setting instruction for the first instruction of a later conversation such as that shown in Figure 2F, the natural language text information of the conversation recorded in storage unit 1170 and the information on the date and time when the conversation took place can be read out and used to generate the setting instruction.
[0077] When natural language text information of a past conversation history is used to generate the setting instruction sentence, the format can be determined freely to a certain extent because the data is to be sent to a large-scale language model, but as shown in Figure 2F, natural language prefixes and suffixes such as "I talked about the following on ____ day of ____ month" and "You talked about the following on ____ day of ____ month" can be prepared and merged with the natural language text information of the recorded conversation to generate the text of the setting instruction sentence.In addition, information on the date and time of the conversation read from storage unit 1170 can be merged with the above-mentioned "____ day of ____ month" portion to form part of the text of the setting instruction sentence.
[0078] Even if the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations has ended, if the above-described generation process and transmission process of the setting instruction sentence in Figure 2F are performed, the response of the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation. This is more preferable because it is perceived by the user as ensuring the consistency of the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation.
[0079] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2G shows an example of the natural language text of the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 19001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response that serves as the response.
[0080] Figure 2G shows an example of a series of conversations shown in Figure 2F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 2G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 2F, so repeated description will be omitted.
[0081] As shown in the natural language text of the server response in the table of Figure 2F, by using the setting directive shown in Figure 2F, the server response by the large-scale language model artificial intelligence of the large-scale language model server 19001 reflects the character's settings and conversation history, such as role, name, conversational characteristics, or personality, at the time of the previous conversation. This is more preferable because it allows the user to recognize that the character's settings and memories, such as role, name, conversational characteristics, or personality, at the time of the previous conversation, are more identical. Note that this allows the characters to be viewed as identical from the user's perspective, so it may also be referred to as a pseudo-identity of the characters from the user's perspective.
[0082] Furthermore, from the user's perspective, they can share memories with the character, providing a more enjoyable character conversation experience.
[0083] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2H. This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010) and the large-scale language model server 19001. Specifically, FIG. 2H shows an example of the operation in which a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. A character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input input to the operation input unit 1107 or an operation detected by a touch operation input sensor of the display unit 10011.
[0084] In the example of FIG. 2H , in addition to character 19051 (named "Koto") used in the description of FIGS. 2A to 2G , character 19052 (named "Tom") and character 19053 (named "Necco") are shown. Character 19051 (named "Koto") and character 19052 (named "Tom") are human characters, and character 19053 (named "Necco") is a cat character. Character display switching on display unit 10011 can be achieved by switching to and displaying on display unit 10011 images generated by rendering characters in different virtual 3D spaces for each character. The process for displaying rendered images of the 3D models of each character can be, for example, any of the first to third processing examples described in FIG. 15A . Furthermore, depending on the character, a dynamic 2D image may be displayed.
[0085] Furthermore, it is preferable that the character operation program executed by control unit 1110 also changes the synthetic voice used for the "speech" of each character when switching the display of the characters displayed on display unit 10011. This can be achieved by storing synthetic voice data of voice tones associated with each character in storage unit 1170 in advance, and performing synthetic voice change processing when switching the display of the character.
[0086] In the example of Fig. 2H, the user 230 is configured to be able to converse with any of the characters. In the AI response output device 10010 of Fig. 2H, each of these characters is set with a different role, name, conversational characteristics, personality, etc. Also, the memories of each character based on the conversation history are managed as different for each character.
[0087] Therefore, the AI response output device 10010 constructs a database shown in FIG. 2I in the storage unit 1170, and manages the character settings and the character conversation history using this database.
[0088] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to Figure 2I. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, Figure 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.
[0089] A character operation program executed by the control unit 1110 of the AI response output device 10010 constructs the database 19200 in, for example, the storage unit 1170. The character ID is an identification number that identifies each of multiple characters that can be displayed on the AI response output device 10010, and may be a natural number or an alphabet, etc. The name is data of the name of each of multiple characters that can be displayed on the AI response output device 10010.
[0090] The initial setting instruction is text information in a natural language that explains the settings such as the role, name, conversational features, or personality of each of multiple characters that can be displayed on the AI response output device 10010. The initial setting instruction is natural language text information that is the main data of the setting instruction sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the description can be read as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0091] The conversation histories, which continue as conversation histories 1, 2, ..., are records of conversations between each character and the user, and are recorded separately for each character. The conversation histories will be included in the natural language text information, which is the main data of the setting instruction statement sent from the AI response output device 10010 to the large-scale language model server 19001, so it is desirable that the content of the conversation histories be readable as is by the large-scale language model of the AI in the large-scale language model server 19001.
[0092] When the character to be displayed on the display unit 10011 of the AI response output device 10010 is switched, the character operation program executed by the control unit 1110 of the AI response output device 10010 uses the database 19200 of Figure 2I to select and switch the initial setting instruction statement and conversation history used for the natural language text information that is the main data of the setting instruction statement transmitted from the AI response output device 10010 to the large-scale language model server 19001 so as to correspond to the character to be displayed on the display unit 10011 of the AI response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the conversation history in the conversation history area of the database 19200 of Figure 2I that corresponds to the character to be displayed on the display unit 10011.
[0093] By using the database 19200 in this manner, the character operation program executed by the control unit 1110 of the AI response output device 10010 establishes a conversation between the user 230 and the character using character utterances that utilize responses from the same AI large-scale language model of the same large-scale language model server 19001. From the user's perspective, this preserves the uniqueness of each character's personality and other settings, and it feels as if each character's unique conversational memories continue. This is more preferable because it appears to the user that the identity of each character's settings and memories, such as their role, name, conversational characteristics, or personality, from the time of the previous conversation, is more clearly maintained. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.
[0094] Therefore, even when the AI response output device 10010 is configured to switch the character displayed on the display unit 10011 from among multiple character candidates, the operation using the database 19200 described above reduces the sense of discomfort felt by the user when conversing with each character, and the user can share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience.
[0095] Note that if the initial setting instructions for multiple characters cannot be edited by the user, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained close to the intentions of the provider of the AI response output device 10010 or the creator of the character content. Alternatively, the initial setting instructions for each character may be edited by the user in response to input from the operation input unit 1107 or the like. In this case, the character's role, name, conversational characteristics, or personality can be set to preferred settings, allowing the user to converse with a character that they have individually set. In this case, the character's 3D model, its rendered image, and the type of synthesized voice for the character may be replaced accordingly.
[0096] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 2 of the present invention will be described using Figure 2J. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 19001. Specifically, a method for providing a character conversation service at a lower cost using the character conversation device based on the artificial intelligence response output device 10010 and the character conversation system based on the artificial intelligence response output device 10010 and the large-scale language model server 19001 will be described.
[0097] As explained in FIG. 2B , limiting a large-scale language model to a specific application and training it at this level of artificial intelligence is extremely resource inefficient. Therefore, it is more resource efficient to generate a model through large-scale training as a foundation model that can be applied to various applications, and then use it on various terminals via an API (Application Programming Interface). In such cases, providers of large-scale language models often recover the costs incurred in training the large-scale language model from terminal users as API usage fees. In such cases, with natural language models, API usage fees are often charged based on the number of processing times of units of words that separate sentences, called tokens.
[0098] Therefore, in the artificial intelligence response output device 10010 of Example 2 of the present invention, by reducing the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large-scale language model server 19001 using an API, it is possible to provide users with a character conversation device using the artificial intelligence response output device 10010 and a character conversation service using a character conversation system using the artificial intelligence response output device 10010 and the large-scale language model server 19001 at a lower cost.
[0099] For example, by using the processing and configuration of Examples 1 to 3 shown in the table of Figure 2J, it is possible to technically reduce the number of tokens in the natural language text information transmitted between the AI response output device 10010 and the large-scale language model server 19001 using an API.
[0100] Example 1 is an example of a method for reducing the number of tokens in the conversation history text stored in the API setting directive and transmitted, in which the conversation history text is shortened by using a document summarization process to reduce the number of tokens. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. While the text summarization may be performed at the start of the next conversation, it is more time-efficient to perform it at the end of the "series of conversations."
[0101] Furthermore, the text summarization process may be requested from the large-scale language model itself of the large-scale language model server 19001. However, in this case, the effect of saving the number of tokens is low. Therefore, for example, if the second server 19002 provides natural language text summarization process via an API at a lower cost than the large-scale language model of the large-scale language model server 19001, the second server 19002 can be requested to perform the text summarization process via the API, and the text summary of the conversation history can be stored in a setting instruction statement for the large-scale language model server 19001 and transmitted.
[0102] Furthermore, if it is only text summarization processing, it can be performed on the terminal side, and text summarization can be performed by the control unit 1110 executing a document summarization program loaded in the memory 1109 of the AI response output device 10010. In this case, the effect of saving tokens is high. Furthermore, even if the conversation history becomes long, if an upper limit on the number of characters after summary is specified in the text summarization processing, the upper limit on the text length of the conversation history is determined, so that an upper limit on the number of tokens can be set and tokens can be saved.
[0103] In addition, since the text information of the character's initial settings, such as the character's role, name, conversational characteristics, or personality, does not increase as much as the conversation history, it is efficient and preferable to maintain the text information of the character's initial setting instructions and reduce the number of tokens of the text information in the conversation history.
[0104] The processing described in Example 1 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0105] Example 2 is another example of reducing the number of tokens in the conversation history text stored in the API setting instruction and transmitted. For example, the number of tokens is reduced by deleting the oldest conversation history among the conversation histories with characters recorded in the storage unit 1170. Specifying an upper limit on the number of characters in the conversation history determines the upper limit on the length of the conversation history text, thereby setting an upper limit on the number of tokens and saving tokens. Alternatively, a method may be used in which a predetermined period of the conversation history is specified and conversation history exceeding that period is deleted. This also saves tokens. Note that, even in Example 2, the text information for the character's initial setting, such as the character's role, name, conversation characteristics, or personality, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the text information in the character's initial setting instruction and reduce the number of tokens in the text information for the conversation history.
[0106] The processing described in Example 2 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0107] Example 3 is a method for reducing the number of tokens by reducing the frequency of sending setting instruction statements using an API. Specifically, even after the device is powered on, after the displayed character is switched, or after the video settings and synthetic voice settings for the displayed character are completed, setting instruction statements are not sent in advance. Instead, setting instruction statements are sent to the large-scale language model server 19001 only when the control unit 1110 determines that natural language text information included in the user's speech picked up by the microphone 1139 is text information that should be processed using a large-scale language model of artificial intelligence. This reduces the frequency of sending setting instruction statements to the large-scale language model server 19001 and reduces the number of tokens.
[0108] Specifically, for example, after the device is powered on or after an operation input to switch the displayed character is made, character 19051 (named "Koto") is displayed on display unit 10011 as shown in FIG. 2H by display processing on display unit 10011 controlled by a character operation program executed by control unit 1110. At this time, for example, if a synthetic voice for the character 19051 to appear is stored and prepared in storage unit 1170 or the like, a synthetic voice for the character's appearance such as "Good morning, I'm Koto," "Hello, I'm Koto," or "Good evening, I'm Koto" may be output from the speaker, which is audio output unit 1140. At this time, the image of character 19051 has already been set as the image of the character displayed on display unit 10011, and the synthetic voice output from the speaker, which is audio output unit 1140, is set to the synthetic voice corresponding to character 19051.
[0109] Here, the inference process of the large-scale language model of the artificial intelligence in the large-scale language model server 19001 already described also takes time as the instruction becomes longer. In particular, if the setting instruction includes text information related to past conversation history, the number of tokens in the instruction increases, resulting in a particularly long inference process time. The setting instruction itself and its response are not output to the user 230. Based on the response to the user instruction following the setting instruction, a synthesized voice as the character's "utterance" is output from the speaker, which is the voice output unit 1140. Therefore, it may seem preferable at first glance to transmit the setting instruction from the AI response output device 10010 to the large-scale language model server 19001 in advance and complete the inference process of the large-scale language model for the setting instruction in advance, because this would result in a faster response in the output of synthesized voice of the character's "utterance" after the user 230 speaks to the character 19051.
[0110] However, if the setting instruction is transmitted to the large-scale language model server 19001 before the user 230 speaks, completing the inference processing of the large-scale language model for the setting instruction in advance, for example, the user 230 may turn off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or the user 230 may switch the displayed character from character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. In this case, the number of tokens processed in the inference processing of the large-scale language model by transmitting the setting instruction to the large-scale language model server 19001 in advance becomes the number of processed tokens for which the usage fee is wasted. This hinders the provision of a character conversation device using the AI response output device 10010 and a character conversation service using a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 at a lower cost to users.
[0111] Therefore, after the AI response output device 10010 is powered on or an operation input is made to switch the displayed character, the AI response output device 10010 sets the image of character 19051 as the image of the character to be displayed on the display unit 10011 under the control of the character operation program executed by the control unit 1110, and even after setting the synthetic voice output from the speaker, which is the audio output unit 1140, to the synthetic voice corresponding to character 19051, it is desirable to continue not to send the setting instruction statement to the large-scale language model server 19001 until the user 230 recognizes that he or she is speaking to character 19051.
[0112] Here, the point in time at which it is recognized that the user 230 is speaking to the character 19051 may be, for example, the point in time at which the trigger keyword described in Figure 2B is detected, or the point in time at which the text of the words spoken by the user 230 is extracted. In this way, the number of processing tokens that waste usage fees can be reduced, and a character conversation service by a character conversation device using the AI response output device 10010 or a character conversation system using the AI response output device 10010 and the large-scale language model server 19001 can be provided to users at a lower cost.
[0113] Furthermore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051 has passed, it is desirable to continue not transmitting setting instructions to the large-scale language model server 19001, for example, if the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Specifically, examples of preset keywords include keywords such as "try jumping" and "try dancing," which are keywords by which the user 230 requests the character 19051 to react, such as by animating the character 19051 to move or emitting synthetic voice. In this case, the character operation program executed by the control unit 1110 reads out the motion data, animation video, and / or synthetic voice data corresponding to the character 19051 stored in the storage unit, and uses these data to generate video to be displayed on the display unit 10011 and output synthetic voice from the speaker, which is the audio output unit 1140.
[0114] Such processing does not necessarily require inference processing of the large-scale language model of the large-scale language model server 19001. After the processing, if the user 230 turns off the power of the AI response output device 10010 by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or if the user 230 switches the displayed character from the character 19051 to another character by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, for example, if the setting instruction statement is sent to the large-scale language model server 19001 first and processed by the inference processing of the large-scale language model, the number of tokens for that processing will be the number of processing tokens that unnecessarily consumes the usage fee.
[0115] Therefore, even after the point in time when it is recognized that the user 230 is speaking to the character 19051, it is desirable to continue not sending the setting instruction statement to the large-scale language model server 19001 until it is determined whether or not the text information extracted from the voice of the user 230 picked up by the microphone 1139 corresponds to a preset keyword that does not require inference processing of a large-scale language model. Only when it is determined that inference processing of a large-scale language model is required is it desirable to send the setting instruction statement to the large-scale language model server 19001 and proceed with the inference processing of the large-scale language model.
[0116] The processing described in Example 3 may be performed by the character operation program executed by the control unit 1110 controlling each unit.
[0117] According to the methods for reducing (saving) the number of processed tokens in a large-scale language model using the examples in Figure 2J described above, a character conversation device using the AI response output device 10010, or a character conversation service using a character conversation system using the AI response output device 10010 and the large-scale language model server 19001, can be provided to users at a lower cost.
[0118] Next, an example of the display of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2K. The example of FIG. 2K shows an example in which a response from a large-scale language model to an instruction statement from a user, as described in each of FIGS. 2A to 2J, is displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this example shows text 10063, which is the response from the large-scale language model, displayed on the display unit 10011 together with an image of a character 19051. The text 10063, which is the response from the large-scale language model, may be displayed superimposed on top of the image of the character 19051, as shown in FIG. 2K. Alternatively, the text 10063, which is the response from the large-scale language model, may be displayed together with the image of the character 19051 without being superimposed on the image of the character 19051.
[0119] The display in Figure 2K is an example, but for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to the minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by audio.
[0120] Therefore, in this case, the control unit 1110 may perform control to start a display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051, as shown in Fig. 2K. In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more suitably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, which is a response from the large-scale language model, is displayed together with the image of the character 19051 may be manually switched ON / OFF by the user 230 operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011.
[0121] Next, an example of a fixed response phrase database (fixed response phrase DB) in the character conversation device (artificial intelligence response output device 10010) capable of displaying multiple characters, as described in FIGS. 2H and 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition numbers and condition contents are the same as those in FIG. 1C. In the example of FIG. 2L, individual fixed response phrases are set for each of the multiple characters in response to these conditions. For example, fixed response phrases for each condition are stored for each of the three characters, Character 1: Koto, Character 2: Tom, and Character 3: Necco, as described in FIGS. 2H and 2I. The output control of fixed response phrases is the same as that in FIG. 1C, and therefore will not be repeated.
[0122] In the example of FIG. 2L , the control unit 1110 selects a corresponding fixed response phrase from a fixed response phrase database (fixed response phrase DB) based on the character displayed in the character conversation device (artificial intelligence response output device 10010) and the current conditions, and uses the selected fixed response phrase to control the output of the response uttered by the character. For example, in the example of the fixed response phrase database (fixed response phrase DB) in FIG. 2L , even under the same conditions, the fixed response phrases are changed to expressions or content that correspond to the individuality of the character. This allows the character conversation device (artificial intelligence response output device 10010) to provide the user with conversations that correspond to the individuality of the displayed character. This allows the user to feel as if each character has a more consistent personality. This makes it possible to realize a character conversation device (artificial intelligence response output device 10010) that gives multiple characters a more realistic presence.
[0123] The fixed response phrase database (fixed response DB) of FIG. 2L described above may be stored in the storage unit 1170, and may be used by the control unit 1110 of the AI response output device 10010. However, the fixed response phrase database (fixed response DB) shown in FIG. 2L may also be provided on the large-scale language model server 19001 side. In this case, the control unit of the large-scale language model server 19001 may generate a response using the fixed response phrase database (fixed response DB). The control unit of the large-scale language model server 19001 may transmit a response generated using the fixed response phrase database (fixed response DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in each server. In this way, even if the AI response output device 10010 does not have a fixed response phrase database (fixed response DB), it is possible to generate a response using the fixed response phrase database (fixed response DB).
[0124] The character conversation device and character conversation system according to the second embodiment described above can reduce the sense of discomfort felt by the user from a conversation with a character displayed on the AI response output device 10010. Furthermore, the character conversation device and character conversation system according to the second embodiment can provide a character conversation service to users at a lower cost.
[0125] In the above description of Example 2, an example has been described in which the large-scale language model included in the large-scale language model server 19001 is used as the large-scale language model. However, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B , and the large-scale language model included in the local LLM processing unit 10028 may be used instead of the large-scale language model included in the large-scale language model server 19001. In this case, in the above description of Example 2, the large-scale language model included in the large-scale language model server 19001 may be read as the large-scale language model included in the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).
[0126] In this case, too, it is possible to further reduce the sense of discomfort felt by the user from conversations with characters displayed on the AI response output device 10010. Note that when a large-scale language model held by the local LLM processing unit 10028 is used instead of the large-scale language model held by the large-scale language model server 19001, there is less need to consider usage fees according to the number of processed tokens, but by reducing the number of processed tokens even for the large-scale language model held by the local LLM processing unit 10028, it is possible to reduce the consumption of resources such as power required for inference. In this case, it is possible to provide the user with a character conversation service that consumes less power.
[0127] In the above description of Example 2, an example has been described in which a conversation history with a character is recorded and stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010). Alternatively, a conversation history with a character may be recorded and stored in a second server 19002 or another cloud server connected to the Internet 19000. In this case, when a user and a character start a new conversation, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or another cloud server, acquires (downloads) a past conversation history between the character and the user, stores it in the storage unit 1170 or memory 1109 of the character conversation device (artificial intelligence response output device 10010), and uses it to create an instruction for a large-scale language model. The specific method of using a past conversation history to create an instruction for a large-scale language model is as described in each figure of Example 2, so a repeated description will be omitted.
[0128] Furthermore, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character's conversation history up to a predetermined point in time, such as each time a conversation between the user and the character takes place or when the conversation between the user and the character ends, to the second server 19002 or another cloud server. That is, the character conversation device (artificial intelligence response output device 10010) may upload the conversation history with the character to the second server 19002 or another cloud server at a predetermined timing, and when the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or another cloud server and use it to generate an instruction sentence for the large-scale language model. In this way, the character conversation device (artificial intelligence response output device 10010) that the user used the previous day and the character conversation device (artificial intelligence response output device 10010) that the user is about to use are different individual devices but can display the same character, and when the user has multiple conversations with the same character between the different individual devices at different times, it is possible to realize a conversation in which the character's memories are pseudo-carried over from the previous conversation, which is more convenient for the user.
[0129] The process described above in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads a conversation history with a character to the second server 19002 or another cloud server to pseudo-continue the character's memory is also effective when handling the database 19200 containing the conversation histories of multiple characters described in Figures 2H and 2I. In other words, if the database 19200 described in Figure 2I is configured to be uploaded and downloaded to the second server 19002 or another cloud server, not only for one character but for multiple characters, and between different individual devices, when a user has multiple conversations with each of the multiple characters at different times, it is possible to realize a conversation in which the memory of each character is pseudo-continued from the previous conversation, which is more convenient for the user.
[0130] <Embodiment 3> Next, embodiment 3 of the present invention is an improvement of the character conversation device (artificial intelligence response output device 10010) and character conversation system described in the drawings of embodiment 2. In this embodiment, differences from embodiment 2 will be described, and repeated description of the same configurations as those of these embodiments will be omitted.
[0131] As in Example 2, the character in Example 3 can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can be helpful to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and character conversation system in this example may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0132] An example of a character conversation device and a character conversation system according to a third embodiment of the present invention will be described with reference to Fig. 3A. The character conversation system according to the third embodiment is provided with a large-scale language model server 20001 instead of the large-scale language model server 19001 shown in Fig. 2A, and is connected to the Internet 19000.
[0133] Here, the large-scale language model server 20001 is a server equipped with large-scale language model artificial intelligence, and is a multimodal large-scale language model artificial intelligence that can process not only the natural language text information that the large-scale language model server 19001 could process, but also types of information other than natural language text information.
[0134] Furthermore, the AI response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (AI response output device 10010) of the second embodiment, as an example.
[0135] In the third embodiment as well, the AI response output device 10010, which is a character conversation device, can communicate with the large-scale language model of the large-scale language model server 20001 via the Internet 19000 using an API.
[0136] The character conversation system of the third embodiment includes a mobile information processing terminal 20010 used by a user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.
[0137] 3B , an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011, which is a touch operation input panel, a control unit 20012, an external power supply input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, a posture sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.
[0138] The display panel 20011 is equipped with a touch operation input sensor and can accept touch operation input by the finger of the user 230. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel. The display panel 20011 may also be referred to as a display unit.
[0139] The communication unit 20020 may be configured with a Wi-Fi communication interface, a Bluetooth communication interface, a mobile communication interface such as 4G or 5G, or the like. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and memory, and the control unit controls the display panel 20011 and the communication unit 20020, etc. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with a communication device 19011 connected to the Internet 19000. This allows the mobile information processing terminal 20010 to communicate with various servers connected to the Internet 19000.
[0140] The power supply 20014 converts AC current input from the outside via the external power supply input interface 20013 into DC current and supplies the required DC current to each unit of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Furthermore, the secondary battery 20015 supplies power to each unit requiring power via the external power supply input interface 20013 when power is not supplied from the outside.
[0141] The video signal input unit 20023 connects to an external video output device and inputs video data. The video signal input unit 20023 may be configured with various digital video input interfaces. For example, it may be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be configured with various USB interfaces.
[0142] The audio signal input unit 20024 is connected to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, a coaxial digital terminal interface, or the like. The audio signal input unit 20024 may also be various USB interfaces, etc. In the case of an HDMI-standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface in which a terminal and a cable are integrated.
[0143] The audio output unit 20021 can output audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 can also output audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured with a speaker. The audio output unit 20021 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output audio as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard.
[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may record a person's voice, such as a user's voice, and the control unit 20012, which will be described later, may perform voice recognition processing on the generated audio signal to acquire text information from the audio signal.
[0145] The imaging unit 20025 is a camera having an image sensor. A camera may be provided on the front side of the mobile information processing terminal 20010, on the display panel 20011 side, or on the back side of the display panel 20011 side. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.
[0146] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be configured with a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid-state drive (SSD). For example, various types of information, such as video data, image data, and audio data, may be recorded in the storage unit 20016 before product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices, external servers, etc., via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may also be output to external devices, external servers, etc., via the communication unit 20020.
[0147] The video control unit 20017 performs various controls related to the video signal input to the display panel 20011. The video control unit 20017 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, an FPGA, or a video processor. The video control unit 20017 may also be referred to as a video processing unit or an image processing unit. The video control unit 20017 performs video switching control, such as determining which video signal to input to the display panel 20011, between the video signal to be stored in the memory 20026 and the video signal (video data) input to the video signal input unit 20023. The video control unit 20017 may also perform image processing control on the video signal input from the video signal input unit 20023 and the video signal to be stored in the memory 20026. Examples of image processing include scaling processing, which enlarges, reduces, or deforms an image; brightness adjustment processing, which changes the brightness; contrast adjustment processing, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into light components and changes the weighting of each component.
[0148] The attitude sensor 20018 is a sensor configured by a gravity sensor or an acceleration sensor, or a combination of these, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each unit connected thereto.
[0149] The nonvolatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the nonvolatile memory 20027 includes, for example, data for various operations to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for objects to be operated by user operations, etc. The memory 20026 stores video data to be displayed on the display panel 20011, data for controlling the device, etc. The control unit 20012 may read various software from the storage unit 20016 and expand and store it in the memory 20026.
[0150] The control unit 20012 controls the operation of each unit connected to it. The control unit 20012 may also work in conjunction with a program stored in the memory 20026 to perform arithmetic processing based on information acquired from each unit within the mobile information processing terminal 20010.
[0151] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to Figure 3C. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. In the third embodiment, the character conversation device (artificial intelligence response output device 10010) also loads a character operation program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character operation program, thereby realizing the various processes described below.
[0152] In Example 2, the actions performed by the user 230 with respect to the character conversation device (artificial intelligence response output device 10010) were mainly calls made by the user 230 using his / her voice. The character conversation device (artificial intelligence response output device 10010) of Example 2 performed a series of operations starting with a process of collecting the user 230's voice with a microphone. In contrast, the character conversation device (artificial intelligence response output device 10010) of Example 3 is also capable of performing a series of operations starting with a process of collecting the user 230's voice with a microphone, as described in Example 2. In addition, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action with respect to the character conversation device (artificial intelligence response output device 10010) by user operation via the operation input unit 1107 of FIG. 1B . Here, examples of the operation input unit 1107 of FIG. 1B include a mouse, a keyboard, and a touch panel.
[0153] Furthermore, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by performing a touch operation by the user that can be detected by the touch operation input sensor of the display unit 10011 of Figure 1B.
[0154] In addition, the user 230 can operate the mobile information processing terminal 20010 and communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, thereby inputting the user's 230 operation input into the character conversation device (artificial intelligence response output device 10010).
[0155] Alternatively, an information storage image such as a two-dimensional code storing information that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on display panel 20011 of mobile information processing terminal 20010, and the image of the display may be captured by imaging unit 1180 of FIG. 1B possessed by the character conversation device (artificial intelligence response output device 10010). Control unit 1110 of character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as a two-dimensional code captured by imaging unit 1180, and obtain the information. Alternatively, an image that the user wishes to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on display panel 20011 of mobile information processing terminal 20010, and the image of the display may be captured by imaging unit 1180 of FIG. 1B possessed by the character conversation device (artificial intelligence response output device 10010). The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180 and obtain the results of the image recognition processing.
[0156] As described above, the character conversation device (artificial intelligence response output device 10010) of Example 3 has a greater variety of actions that the user 230 can take toward the character conversation device (artificial intelligence response output device 10010) than the character conversation device (artificial intelligence response output device 10010) described in Example 2. As a result, the character conversation device (artificial intelligence response output device 10010) of Example 3 can acquire the results of actions taken by the user 230 other than the user's voice, and generate instructions (prompts) to be sent to the large-scale language model server 20001 based on the results. As a result, the instructions to be sent to the large-scale language model server 20001 can more preferably include information other than natural language text information extracted from the user's voice. Examples of information other than natural language text information extracted from the user's voice include images, videos, and audio.
[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this embodiment uses an API to send an instruction to the large-scale language model server 20001. In this embodiment, the instruction may be metadata containing information written in a notation using tags, such as the markup format of a markup language, a notation using predetermined symbols, such as Markdown, or an object notation of a predetermined script, such as JSON. In this embodiment, the instruction may be classified into two types: a setting instruction that stores instructions, such as initial settings, and a user instruction that reflects instructions from a user. Type identification information identifying whether the instruction is a setting instruction or a user instruction may be stored in a portion other than the main message of the instruction. In this case, the instruction includes natural language text information as the main message. Furthermore, in this embodiment, the main message of the instruction may include, in addition to the natural language text information, a non-natural language information source, such as an image, video, or audio, as information of a type other than the natural language text information. A specific method for including a non-natural language information source in the instruction will be described later.
[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources in addition to natural language text information. The large-scale language model server 20001 receives an instruction from a character conversation device (artificial intelligence response output device 10010). Based on the instruction, the multimodal large-scale language model performs inference and generates a response including natural language text information that is the result of the inference. Here, because the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.
[0159] The character conversation device (artificial intelligence response output device 10010) receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or audio stored as the main message in the response. The character operation program of the character conversation device (artificial intelligence response output device 10010) may use speech synthesis technology to generate a natural language voice as a reply to the user based on the natural language text information extracted from the response, and output the voice from the audio output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.
[0160] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display natural language characters serving as a response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on natural language text information extracted from the above-mentioned response. In this case, the characters may be displayed together with the character 19051, may be superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0161] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display an image on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the information of the image of the non-natural language information source extracted from the response. At this time, the image may be displayed together with the character 19051, may be superimposed on the image of the character 19051, or may be displayed in place of the image of the character 19051. These specific processes may be executed by the image control unit 1160.
[0162] Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video on the display screen of the character conversation device (artificial intelligence response output device 10010) to present to the user based on the video information of the non-natural language information source extracted from the above-mentioned response. In this case, the video may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed in place of the video of the character 19051. These specific processes may be executed by the video control unit 1160.
[0163] In addition, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output a voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140, which is a speaker.
[0164] According to the character conversation device (artificial intelligence response output device 10010) of FIG. 3C or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001 described above, it is not necessary to install a large-scale language model, which requires vast amounts of data and computational resources for learning, in the character conversation device (artificial intelligence response output device 10010). Furthermore, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via an API. In response to a user's action toward a character, a response based on a non-natural language information source can be provided in addition to a response based on natural language text, enabling a more appropriate conversation to be held.
[0165] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3D . This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3D shows examples of non-natural language information sources, such as natural language text and images, of the main message of an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001, and examples of non-natural language information sources, such as natural language text and images, of the main message of a server response that is the response. In this embodiment, the non-natural language information source can be an image, video, audio, or the like, but FIG. 3D shows an example of an image as the non-natural language information source.
[0166] 3D also shows an exchange of instructions and responses in chronological order, from a setting instruction, a first round of user instructions and their responses, to a second round of user instructions and their responses. The instructions and responses shown in FIG. 3D include non-natural language information source 20061 and non-natural language information source 20062, which were not shown in FIG. 2D of Example 2. In the example of FIG. 3D, both non-natural language information source 20061 and non-natural language information source 20062 are images.
[0167] 3D, for ease of explanation, an image of the non-natural language information source 20061 is shown pasted into the instruction. However, there are multiple methods for transmitting or specifying data from the non-natural language information source 20061 in an instruction sent from the character conversation device (artificial intelligence response output device 10010) to the large-scale language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these methods or switch between them. An example of each method will be described below.
[0168] The first method for transmitting or specifying non-natural language information source data in a directive is used, for example, when the non-natural language information source to be specified is a non-natural language information source located in a location such as a server connected to a network such as the Internet. A specific example of the first method is to use information such as tags and symbols in the directive to specify a non-natural language information source file located on a network such as the Internet by using location information (such as a URL) and a file name on the network such as the Internet.
[0169] For example, it is a tag that specifies an image in a markup language.<img src=“****”> You can also specify an image on a network such as the Internet by writing the location information and file name information of the image file in the **** part using the tag.<video src=“****”> You can also specify a video file on a network such as the Internet by writing the location information and file name information of the video file in the **** part using the tag.<audio src=“****”> Using the above, audio existing on a network such as the Internet may be specified by entering location information and file name information of the audio file in the **** portion. Furthermore, in the case of the JSON format notation, an image existing on a network such as the Internet may be specified by preparing a key such as img_src and entering location information and file name information of the image file as the value. For video files and audio files, respective keys and values may also be prepared. This specific example of a format is merely an example, and other unique formats may also be used. In either case, information specifying the location information and file name information of the non-natural language information source file may be stored in the directive.
[0170] When information specifying the location information and file name information of a non-natural language information source file is stored in a directive as in the first method, the directive itself does not need to store the data of the non-natural language information source file itself. Therefore, the amount of data in the directive can be reduced. In the first method, the large-scale language model server 20001 that receives a directive specifying non-natural language information source data simply uses the location information and file name information of the non-natural language information source file stored in the directive to obtain the non-natural language information source file located in a location such as a server connected to a network such as the Internet.
[0171] Here, we will explain how location information and file name information are input when the character conversation device (artificial intelligence response output device 10010) specifies non-natural language information source data in an instruction sentence using the first method. In FIG. 3C , we explained that in this embodiment, the types of actions that the user 230 can perform on the character conversation device (artificial intelligence response output device 10010) are increased beyond the voice of the user 230 compared to Example 2. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, etc., by user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 of FIG. 1B .
[0172] Furthermore, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a web browser program and display a GUI of the web browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the web browser program may be accepted by a user operation (e.g., a mouse, keyboard, or touch panel) via the operation input unit 1107 or a user touch operation detectable by a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, video, or audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction. In this case, the web browser program may acquire location information and file name information of the non-natural language information source data and pass them to the character operation program.
[0173] Furthermore, user 230 may input location information such as a URL for specifying non-natural language information source data to character conversation device (artificial intelligence response output device 10010) by operating mobile information processing terminal 20010 and communicating with the character conversation device (artificial intelligence response output device 10010) from mobile information processing terminal 20010. Alternatively, location information such as a URL for specifying non-natural language information source data, file name information, and the like may be input by displaying an information storage image such as a two-dimensional code on display panel 20011 of mobile information processing terminal 20010, performing image recognition processing on an image captured by imaging unit 1180 of character conversation device (artificial intelligence response output device 10010), and acquiring the results of the image recognition processing, as described in FIG.
[0174] Note that the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file is previously present in a location such as a server connected to a network such as the Internet. For example, if non-natural language information source data such as images, videos, or audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) is to be included in the instruction statement, the character conversation device (artificial intelligence response output device 10010) may upload the non-natural language information source data to the second server 19002 via the Internet 19000, and include in the instruction statement the location information (such as a URL) and file name of the non-natural language information source data on the uploaded second server 19002. In this case, the second server 19002 functions as a so-called intermediate server.
[0175] Similarly, when it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in the instruction, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit location information on the Internet (such as a URL) and a file name of the non-natural language information source data on the second server 19002 to the character conversation device (artificial intelligence response output device 10010), and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include in the instruction the location information on the Internet (such as a URL) and a file name of the non-natural language information source data that has been acquired and uploaded to the second server 19002.
[0176] Furthermore, a media server may be constructed within the character conversation device (artificial intelligence response output device 10010) that can be accessed from other servers via the Internet 19000 by the character operation program of the character conversation device (artificial intelligence response output device 10010) in cooperation with memory 1109 and storage unit 1170. In this case, when non-natural language information source data is specified in an instruction statement using the first method, the character conversation device (artificial intelligence response output device 10010) may store location information on the Internet (such as a so-called URL) indicating the location within the media server constructed within the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data in the instruction statement.
[0177] Next, a second method for specifying the transmission or designation of non-natural language information source data in a prompt is, for example, to simply store (attach) the non-natural language information source data itself in the prompt and transmit it. Non-natural language information source data, such as images, videos, and audio, generally has a larger data volume than text information in natural language. Therefore, in this case, the data volume of the prompt itself is larger than in the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) temporarily stores the non-natural language information source data to be stored (attached) in the prompt in memory 1109, and when transmitting the prompt, stores (attaches) the non-natural language information source data in memory 1109 via the communication unit 1132 and outputs it to the large-scale language model server 20001. The non-natural language information source data itself that the character operation program of the character conversation device (artificial intelligence response output device 10010) stores in memory 1109 may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read out from the storage unit 1170 and stored in memory 1109.
[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data using instruction sentences.
[0179] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources along with natural language text information. Therefore, in the first pass of the user instruction shown in the example of Figure 3D, it can obtain images of the swimming pool and poolside, which are non-natural language information source 20061, and text information in natural language, and output the text information in natural language as shown in the figure as a response to the first pass of the user instruction as an inference result.
[0180] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example of the response to the second round of user instructions in Figure 3D, the large-scale language model server 20001 can include in the response a non-natural language information source 20062 generated by inference from the multimodal large-scale language model and send it to the character conversation device (artificial intelligence response output device 10010). In Figure 3D, the non-natural language information source 20062 is an example of an image in which a circle image is added to the image of a swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 3D, and may be video or audio.
[0181] When a response from the large-scale language model server 20001 includes a non-natural language information source other than natural language text information, a method similar to the first method or the second method in which the character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence can be used.
[0182] Specifically, as a method similar to the first method described above, the large-scale language model server 20001 may store, in the response, information specifying the location information and file name information of the non-natural language information source file in an instruction statement. The non-natural language information source 20062, such as an image, video, or audio, itself may be stored in the large-scale language model server 20001, or the non-natural language information source 20062 may be transferred to and stored on the second server 19002 functioning as an intermediate server. In either case, the large-scale language model server 20001 may store, in the response, information specifying the location information and file name information of the non-natural language information source file in an instruction statement. The character conversation device (artificial intelligence response output device 10010) that receives the response may access the large-scale language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction statement to acquire the non-natural language information source 20062.
[0183] Specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction sentence and use it for various outputs to the user 230.
[0184] According to the operation of the character conversation device (artificial intelligence response output device 10010) and character conversation system of Example 3 described above using Fig. 3D, instructions and responses are sent and received between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230 to realize a conversation using non-natural language information such as images, videos, and sounds. This makes it possible to realize a more advanced and natural conversation as shown in each message in Fig. 3D.
[0185] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3E. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3E shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0186] Figure 3E shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in Figure 3D has ended. In the example of Figure 3E, processing using the conversation history as described in Figures 2F, 2G, and 2I of Example 2 is not performed. Therefore, Figure 3E, like Figure 2E of Example 2, is a response with content that does not remember at all the name of the large-scale language model itself, the role to be played, the characteristics of the conversation, the user's name, the conversation history, etc., that were included in the setting instruction statement.
[0187] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3F. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3F shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0188] Figure 3F shows an example of a case where the user 230 speaks to the character 19051 again to start a new conversation after the series of conversations shown in Figure 3D has ended. Here, Figure 3F shows an example in which the method of storing a message explaining the history of past conversations in the setting instruction sentence, which was described in Figure 2F of Example 2, is also applied to the character conversation device (artificial intelligence response output device 10010) of Example 3. Specifically, in Figure 3F, the message that is the content of the setting instruction sentence in Figure 3D is stored as a reset message, and a message explaining the history of past conversations is stored as a conversation history message following the reset message.
[0189] The large-scale language model server 20001 of Example 3 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, and therefore non-natural language information source data may have been transmitted or specified in past instruction sentences and responses. Therefore, in the example of Figure 3F, the conversation history message reflects not only the natural language text information in past instruction sentences and responses, but also the transmission or specification of non-natural language information source data in past instruction sentences and responses. The specific method for transmitting or specifying non-natural language information source data in the instruction sentence of Figure 3F is similar to the transmission or specification of non-natural language information source data described in Figure 3D, and therefore a repeated description will be omitted.
[0190] In the example of Fig. 3D, the method of transmitting or specifying non-natural language information source data includes storing (attaching) the non-natural language information source data itself in the instruction statement, and storing (attaching) the non-natural language information source data in the instruction statement without storing (attaching) the non-natural language information source data in the instruction statement. This also applies to the instruction statement of Fig. 3F.
[0191] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) of Example 3 of the present invention will be described using Figure 3G. This can also be said to be an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3G shows an example of the main message of an instruction sentence sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, which serves as the basis for a conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that serves as the response.
[0192] Figure 3G shows an example of a series of conversations in the series of conversations shown in Figure 3F, from the first user instruction and its response to the third user instruction and its response following the first setting instruction. Figure 3G shows the exchange of instructions and responses in chronological order. The contents of the setting instructions are the same as those shown in Figure 3F, so repeated description will be omitted.
[0193] As described above, even when using the large-scale language model server 20001 having a multimodal large-scale language model capable of processing non-natural language information sources together with natural language text information according to the third embodiment, even when the user 230 speaks to the character 19051 again to start a new conversation after a series of conversations has ended, if the generation and transmission processes of the setting instruction sentences in Figure 3F are performed, the response of the subsequent user instruction sentence will reflect the settings and conversation history of the character's role, name, conversational characteristics, personality, and / or conversational characteristics at the time of the previous conversation, as shown in Figure 3G. This is preferable because it allows the user to recognize that the settings and memories of the character's role, name, conversational characteristics, or personality at the time of the previous conversation are more closely matched.
[0194] Next, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3H . This can also be considered as an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, FIG. 3H uses the example described in FIG. 2H of the second embodiment for the settings of multiple characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Therefore, repeated explanations of the settings of multiple characters will be omitted.
[0195] Furthermore, database 20200 shown in Fig. 3H for managing character settings and character conversation history has the same format as database 19200 shown in Fig. 2I of Example 2, and only the differences from database 19200 shown in Fig. 2I will be explained in Fig. 3H. Also, the content of the character "Koto" in the database will be explained, and the content of other characters will be omitted.
[0196] As described above, the large-scale language model server 20001 of Example 3 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, and therefore both the instruction sentences from the character conversation device (artificial intelligence response output device 10010) and the responses from the large-scale language model server 20001 include not only natural language text information but also the transmission or specification of non-natural language information source data. Therefore, the database 20200 shown in Figure 3H records, in the conversation history data, not only the natural language text information included in these instruction sentences and responses, but also the transmission or specification of non-natural language information source data. The specific method for transmitting or specifying non-natural language information source data in the conversation history recording is the same as the transmission or specification of non-natural language information source data described in Figure 3D, and therefore a repeated description will be omitted.
[0197] In the example of Figure 3D, methods for transmitting or specifying non-natural language information source data include storing (attaching) the non-natural language information source data itself to the instruction statement and not storing (attaching) the non-natural language information source data to the instruction statement. This is also true for the conversation history of Figure 3H. However, in the conversation history of Figure 3H, if the method for specifying non-natural language information source data is to specify location information and file name information of a non-natural language information source file on a server on a network such as the Internet (such as the second server 19002 functioning as an intermediate server or another cloud server), there is a possibility that the non-natural language information source file on the server will be deleted if the conversation history period is long. In this case, even if the location information and file name information are used, the non-natural language information source file cannot be retrieved at a later date, and information from the conversation record may be lost.
[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts and records instruction sentences and response messages into a conversation history, it can use the location information and file name information to obtain the non-natural language information source file itself specified in the instruction sentence and response from a server on the network and store it in the storage unit 1170. Furthermore, the character operation program of the character conversation device (artificial intelligence response output device 10010) can rewrite the location information and file name of the non-natural language information source file to Internet location information (such as a URL) indicating the location of the media server of the media server built within the character conversation device (artificial intelligence response output device 10010), and then record the non-natural language information source in the conversation record. In this way, unless the character conversation device (artificial intelligence response output device 10010) itself deletes the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the conversation record information, which is more suitable for preserving the conversation record.
[0199] By using the database of Figure 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch between multiple character candidates to display on the display unit 10011, the user will feel less uncomfortable when conversing with each character, and will be able to share memories with each of the multiple characters, resulting in a more enjoyable character conversation experience. This effect can be achieved as shown in Figure 2I of Example 2, and can also be achieved when the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources in addition to natural language text information.
[0200] In addition, in the character conversation device (artificial intelligence response output device 10010) or character conversation system of Example 3, the large-scale language model server 20001 uses a multimodal large-scale language model artificial intelligence that can process not only natural language text information but also non-natural language information other than natural language text information.
[0201] Here, an API is used for communication between the character conversation device (artificial intelligence response output device 10010) and the large-scale language model server 20001. In a multimodal large-scale language model, an API usage fee may be charged according to the number of processed natural language text information units, called tokens, which are units of words that separate sentences, as well as the amount of data from non-natural language information sources.
[0202] Therefore, in order to provide users with a character conversation service by the character conversation system according to this embodiment at a lower cost, the following modified example may be used.
[0203] In a first variation, the database conversation history record of FIG. 3H also records information about the transmission or specification of non-natural language information source data. However, the character and the user exchange conversations about the natural language information source data using text information in natural language, and the content of those conversations is recorded as text information in natural language. Even if the transmission or specification of the natural language information source data is omitted from the database conversation history record of FIG. 3H, the conversation itself about the natural language information source data will still be recorded as text information in natural language to some extent. Therefore, if a certain amount of information reduction is acceptable, the transmission or specification of the natural language information source data may be omitted from the database conversation history record of FIG. 3H. In this case, the transmission or specification of the natural language information source data is also omitted from the conversation history message of the setting instruction sentence of FIG. 3F. This reduces the amount of data from non-natural language information sources communicated using the API.
[0204] Next, as a second variation, in the database conversation history recording of FIG. 3H , instead of transmitting non-natural language information source data or recording specified information, natural language text information describing the content of the non-natural language information source data is recorded. The natural language text information describing the content of the non-natural language information source data may be acquired, for example, by starting a conversation between the large-scale language model of the large-scale language model server 20001 and a character conversation device (artificial intelligence response output device 10010) separately from a conversation as a character, and having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit. Alternatively, the natural language text information may be acquired by having the large-scale language model server 20001 explain the content of the non-natural language information source data by specifying a predetermined character limit through a conversation with another large-scale language model of another server that is more inexpensive to use than the large-scale language model of the large-scale language model server 20001. Furthermore, if alternative text data is prepared at the time of acquiring the non-natural language information source data, the alternative text data may be used as the natural language text information describing the content of the non-natural language information source data. A specific example of alternative text data for non-natural language information source data is the text information written in the **** part of markup language tags such as <img src="" alt="****">, <video src="" alt="****">, and <audio src="" alt="****">.
[0205] Furthermore, in the case of the JSON format notation, in an object that is stored in association with the location information and file name information of the non-natural language information source data, which are keys and values indicating the location information of the non-natural language information source data, a key corresponding to the alternative text can be further linked and stored with a value that is the alternative text data itself.
[0206] In this case, too, the transmission of the natural language information source data or the recording of the specified information can be omitted in the recording of the conversation history in the database of Fig. 3H, and the transmission of the natural language information source data or the specified information can also be omitted from the conversation history message of the setting instruction sentence of Fig. 3F, thereby reducing the amount of data from non-natural language information sources communicated using the API.
[0207] Next, as a third variation, at the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data is not stored in the user instructions, but is replaced with natural language text information explaining the content of the non-natural language information source data. For example, in the first round of user instructions in FIG. 3D , the transmitted or specified information of non-natural language information source data 20061 may be replaced with the user instructions, and a description such as "This image is of a swimming pool, poolside seats, and parasols. There is water in the swimming pool. There are drinks on the table next to the seats" may be stored as natural language text information. In this case, the description may be obtained by having another large-scale language model on another server, which is more inexpensive to use than the large-scale language model on the large-scale language model server 20001, explain the content of the non-natural language information source data with a specified character limit. Alternatively, the description may be obtained from a server of various other services that can obtain summaries and descriptions of non-natural language information source data such as images, videos, and audio. Furthermore, if alternative text data is prepared from the time of acquiring the non-natural language information source data, the alternative text data may be natural language text information that explains the content of the non-natural language information source data.
[0208] Next, an example of a display of the character conversation device (artificial intelligence response output device 10010) according to the third embodiment of the present invention will be described with reference to FIG. 3I. The example of FIG. 3I illustrates an example of displaying a response from a large-scale language model to a user instruction, as described in each of FIGS. 3A to 3H, on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, this example illustrates displaying text 10063 of natural language information source data, image 10064 of non-natural language information source data, and / or video 10065 of non-natural language information source data, which are responses from the large-scale language model, together with a video of the character 19051, on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed superimposed on top of the video of the character 19051, as shown in FIG. 3I.
[0209] Furthermore, text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on the video of character 19051. The display in FIG. 3I is an example. However, for example, if the user 230 adjusts the volume of the audio output of the audio output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to minimum or sets the audio output to OFF by operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011, the user 230 will not be able to hear the response from the large-scale language model by audio. Therefore, in this case, the control unit 1110 may control to start a display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in FIG. 3I.
[0210] In this way, even when the user 230 wishes to refrain from audio output, the user 230 can more preferably use the character conversation device (artificial intelligence response output device 10010). Note that the display mode in which text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the image of the character 19051, may be manually switched ON / OFF by the user 230 operating the operation input unit 1107 or the touch operation input sensor of the display unit 10011. According to the display example of FIG. 3I, the character conversation device (artificial intelligence response output device 10010) that supports multimodal display can more preferably output responses from the large-scale language model.
[0211] According to the character conversation device and character conversation system of Example 3 described above, in addition to the effects of the character conversation device and character conversation system of Example 2, it is possible to provide users with a more advanced conversation experience that includes information in non-natural languages in addition to information in natural languages by using a multimodal large-scale language model. Furthermore, according to the character conversation device and character conversation system of Example 3, it is possible to provide users with a character conversation service at a lower cost.
[0212] In the above description of Example 3, an example has been described in which the large-scale language model included in the large-scale language model server 20001 is used as the large-scale language model. However, the character conversation device (artificial intelligence response output device 10010) may include the local LLM processing unit 10028 shown in FIG. 1B and use the multimodal large-scale language model included in the local LLM processing unit 10028. In this case, the multimodal large-scale language model included in the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model included in the large-scale language model server 20001.
[0213] In this case, in the above description of Example 3, the multimodal large-scale language model of the large-scale language model server 20001 can be replaced with the multimodal large-scale language model of the local LLM processing unit 10028 of the character conversation device (the AI response output device 10010). In this case, the multimodal large-scale language model can be used to provide users with a more sophisticated conversation experience that includes not only natural language information but also non-natural language information. Note that using the multimodal large-scale language model of the local LLM processing unit 10028 instead of the multimodal large-scale language model of the large-scale language model server 20001 reduces the need to consider usage fees based on the number of processed tokens and the amount of data in the non-natural language information source. However, even with the multimodal large-scale language model of the local LLM processing unit 10028, reducing the number of processed tokens and the amount of data in the non-natural language information source can reduce the consumption of resources, such as power required for inference. In this case, a character conversation service with less power consumption can be provided to users.
[0214] The configuration described in Example 2 for uploading and downloading the conversation history with a character or database data including the conversation history with a character to the second server 19002 or another cloud server can also be used in the example using a multimodal large-scale language model described in Example 3. In this case, too, when a user has multiple conversations with each of the multiple characters at different times between different individual devices for one character or multiple characters, it is possible to realize a conversation in which the memory of each character is pseudo-carried over from the previous conversation, which is more convenient for the user.
[0215] <Embodiment 4> Next, embodiment 4 of the present invention is an improvement of the AI response output device 10010, the character conversation device, or these systems described in the drawings of embodiment 2 or embodiment 3. In this embodiment, differences from embodiment 2 or embodiment 3 will be described, and repeated description of the same configuration as these embodiments will be omitted.
[0216] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0217] An example of operation using a database in a character conversation device (artificial intelligence response output device 10010) according to Example 4 of the present invention will be described using Figure 4A. The database according to Example 4 shown in Figure 4A is an extension of the database described with reference to Figure 2I or Figure 3I. Specifically, the database shown in Figure 4A assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores in the database initial setting instructions and conversation histories corresponding to each user and character.
[0218] 4A, for user 1 with a user ID of 1, the initial setting command sentences and conversation histories are stored for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3. In addition, for user 2 with a user ID of 2 and user 3 with a user ID of 3, the initial setting command sentences and conversation histories are stored for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3.
[0219] These initial setting instructions and conversation history data are stored as separate data in different areas for each user-character combination. For the sake of explanation, in Figure 4A, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (AI response output device 10010) uses the initial setting instructions and conversation history stored in different areas for each user-character combination based on the user currently using (logging in to) the character conversation device (AI response output device 10010) or its system, thereby making it possible to more appropriately maintain consistency in the character's personality and continuity of memory for each different user.
[0220] Specifically, consider a situation in which User 1 has previously had a conversation with character Tom using a character conversation device (AI response output device 10010), and User 2 is unaware of that conversation, and then User 2 subsequently has a conversation with character Tom. In this case, if AI response output device 10010 uses an initial setting instruction statement or a conversation history database that does not identify users, the response output from AI response output device 10010 will be based on a conversation history that User 2 does not remember, and there is a possibility that the conversation between User 2 and the character of AI response output device 10010 will become inconsistent.
[0221] In contrast, even in a similar situation, if the database shown in Figure 4A is used, the control unit 1110 of the character conversation device (AI response output device 10010) identifies users by ID, stores the initial setting instructions and conversation history in a different area for each user, and uses the initial setting instructions and conversation history stored in a different area for each user to generate an AI response. As a result, the initial setting instructions and conversation history used to generate an AI response for each user are based on the operation or conversation history of that user and are managed separately from the operation or conversation history of other users. This makes it possible to more appropriately maintain consistency in the conversation history between each user and each character in the AI response output device 10010.
[0222] The database of initial setting directives and / or conversation history described in FIG. 4A may be stored in the storage unit 1170 of the AI response output device 10010 and used by the control unit 1110. Furthermore, without being limited to this, the database of initial setting directives and / or conversation history may be stored in a server on the network. For example, if the AI response output device 10010 uses the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the large-scale language model server 20001 to generate an AI response, the database of initial setting directives and / or conversation history described in FIG. 4A may be stored in these servers themselves. In this way, it is possible to omit the process of transmitting the initial setting directives and conversation history again from the AI response output device 10010 to these servers by including them in directives, thereby reducing the number of transmission tokens required for using the large-scale language model.
[0223] When storing the database of initial setting instructions and / or conversation history described in FIG. 4A , the user ID, character ID, and user instructions for subsequent conversations can be transmitted from the AI response output device 10010 to these servers. The large-scale language models on these servers use the user ID and character ID acquired from the AI response output device 10010 to acquire the corresponding initial setting instructions and conversation history from the database of initial setting instructions and / or conversation history of FIG. 4A . The large-scale language models on these servers perform inference using the initial setting instructions and conversation history, and the user instructions for subsequent conversations transmitted from the AI response output device 10010, to generate an AI response and transmit it to the AI response output device 10010. In this way, the effect of more appropriately maintaining the consistency of a character's personality and the continuity of memory for each different user can be obtained while saving on the number of transmitted tokens when using a large-scale language model.
[0224] Next, an example of operation using a database in the character conversation device (artificial intelligence response output device 10010) according to Example 4 of the present invention will be described with reference to Figure 4B. The database according to Example 4 shown in Figure 4B is an extension of the database described with reference to Figure 1C or Figure 2L. Specifically, the database shown in Figure 4B assumes a case in which multiple different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores data on standard response phrases corresponding to each user and character in the database.
[0225] In the example of Fig. 4B, for user 1, whose user ID is 1, fixed response phrase data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3. In addition, for user 2, whose user ID is 2, and user 3, whose user ID is 3, fixed response phrase data is stored for each of the characters Koto, whose character ID is 1, Tom, whose character ID is 2, and Necco, whose character ID is 3.
[0226] These fixed response phrase data are stored as separate data in different areas for each user-character combination. For ease of explanation, in FIG. 4B , the data stored in each area is represented as fixed response phrase data 101, 102, 103, 201, 202, 203, 301, 302, and 303. For example, fixed response phrase data 101 is stored as a database, such as a table, corresponding to fixed response phrases corresponding to condition numbers 1 to 7 for character 1: Koto shown in FIG. 2L. Data 201 in FIG. 4B is stored as a database, such as a table, corresponding to fixed response phrases corresponding to condition numbers 1 to 7 for character 2: Tom shown in FIG. 2L.
[0227] Data 301 in Figure 4B is stored as a database such as a table corresponding to the fixed response phrases corresponding to condition numbers 1 to 7 of character 3: Necco shown in Figure 2L. Data 102, 202, and 302 in Figure 4B store fixed response phrases in a similar format that have been modified for user 2. Data 103, 203, and 303 in Figure 4B store fixed response phrases in a similar format that have been modified for user 3. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses fixed response phrase data stored in different areas for each combination of user and character, based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.
[0228] In this way, even for the same character, it is possible to respond with different template responses for each user. That is, even for the same character, it may be more appropriate to vary the content of the template responses depending on the relationship between the character and the user. For example, depending on the relationship between the character's set age and the age of the user registered in the AI response output device 10010 or the system, the user may be older, the same age, or younger than the character. In this case, different content for the template responses of the character to older users, the template responses of the character to users of the same age, and the template responses of the character to younger users may result in more appropriate or natural conversations between the user and the character. That is, by performing operations using the database of FIG. 4B and varying the content of the template responses depending on the relationship between the character and the user, it is possible to produce more appropriate or natural conversations.
[0229] The fixed response phrase database (fixed response phrase DB) of FIG. 4B described above may be stored in the storage unit 1170 and used by the control unit 1110 of the AI response output device 10010. However, the fixed response phrase database (fixed response phrase DB) shown in FIG. 4B may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side. In this case, the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response phrase database (fixed response phrase DB). The control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may transmit a response generated using the fixed response phrase database (fixed response phrase DB) to the AI response output device 10010, instead of a response generated using a large-scale language model stored in the respective servers. In this way, even if the artificial intelligence response output device 10010 is not equipped with a standard response phrase database (standard response phrase DB), it is possible to generate a response using the standard response phrase database (standard response phrase DB).
[0230] According to the character conversation device and character conversation system of Example 4 described above, it is possible to produce more suitable or more natural conversations depending on the relationship between the character and the user, the conversation history, etc.
[0231] <Example 5> Next, Example 5 of the present invention is an improvement on the AI response output device 10010 or the AI response output system described in the drawings of Examples 1, 2, and 3. Specifically, this is an example in which the response generation process of the AI response output device 10010 is switched from response generation process using a large-scale language model on a network to response generation process using a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, or response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated description of configurations similar to these examples will be omitted.
[0232] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0233] An example of response generation process switching processing in the AI response output device 10010 according to the fifth embodiment of the present invention will be described using FIG. 5A. The table in FIG. 5A shows examples 1 to 9 of response generation process switching processing in the AI response output device 10010. In the table in FIG. 5A, the column "Switching Overview" shows an overview of the switching processing for each example. The column "State before switching of LLM on network (API-connected LLM)" shows the state before response generation processing by a large-scale language model on a network (large-scale language model connected using an API), such as the large-scale language model provided in the large-scale language model server 19001 in FIG. 1 and the multimodal large-scale language model provided in the large-scale language model server 20001, is switched to another response generation processing. The column "Switching Occurrence Condition" shows the conditions under which switching processing of the response generation processing occurs. The column "Switching destination from LLM on the network (API-connected LLM)" indicates the switching destination to which the response generation process of the AI response output device 10010 is switched from a large-scale language model on the network (a large-scale language model connected using an API), such as the large-scale language model provided in the large-scale language model server 19001 and the multimodal large-scale language model provided in the large-scale language model server 20001. When the condition indicated in "Switching occurrence condition" occurs in the state of "State before switching of LLM on the network (API-connected LLM)" shown in Figure 5A, the control unit 1110 of the AI response output device 10010 may perform control to switch to the large-scale language model, database, or correspondence indicated in "Switching destination from LLM on the network (API-connected LLM)."
[0234] Below, each example shown in the table of FIG. 5A will be described. Example 1, as shown in the "Switching Overview," is an example in which switching is performed depending on the network connection status of the AI response output device 10010. In Example 1, the "State before switching of the LLM (API-connected LLM) on the network" indicates that the network connection status of the AI response output device 10010 is connectable. Here, in Example 1, the "Switching Occurrence Condition" indicates "when network connection becomes unavailable." That is, this means that the connection via the network between the AI response output device 10010 and a large-scale language model on the network (a large-scale language model connected using an API) becomes unavailable. Specifically, this connection inability may be due to a communication inability on the connection path from the AI response output device 10010 to the Internet 19000. Alternatively, this connection inability may be due to a communication inability on the Internet 19000. Alternatively, the connection failure may be due to a situation in which the large-scale language model on the network (large-scale language model connected using an API) itself cannot connect to the Internet 19000. Furthermore, in Example 1, "local LLM" is indicated as the "switching destination from LLM on the network (API-connected LLM)." This specifically means that switching to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010 is performed. That is, in Example 1, even if connection to the large-scale language model on the network (large-scale language model connected using an API) becomes impossible for some reason and response generation processing by the large-scale language model on the network (large-scale language model connected using an API) is unavailable, response generation processing is switched to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. This allows response generation processing to continue using the large-scale language model, despite differences in performance between large-scale language models.
[0235] Next, Example 2 of FIG. 5A will be described. In Example 2, the "switching destination from the networked LLM (API-connected LLM)" in Example 1 is changed from "local LLM" to "prepared response DB (database)." The response generation process using the "prepared response DB (database)" is similar to the process described in FIG. 1C, FIG. 2L, or FIG. 4B, and therefore a repeated description will be omitted. That is, in Example 2, if a connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason and the response generation process using the large-scale language model on the network (a large-scale language model connected using an API) cannot be used, by switching to response generation process using a preparatory response database, it becomes possible to generate a response through simpler processing and output the response to the user.
[0236] Next, Example 3 of FIG. 5A will be described. In Example 3, the "switching destination from the LLM on the network (API-connected LLM)" in Example 1 is changed from "local LLM" to "no-response support." The "no-response support" means that even if a user input requesting a response from a large-scale language model is received from the user via the touch panel, microphone 1139, or operation input unit 1107, no response is generated to this input, or even if a user input requesting a response from a large-scale language model is received, no response is output to this. In other words, Example 3 makes it easier to respond to a situation where connection to a large-scale language model on the network (a large-scale language model connected using an API) becomes impossible for some reason, making it impossible to use the large-scale language model on the network (a large-scale language model connected using an API) for response generation processing.
[0237] Next, Example 4 of FIG. 5A will be described. As shown in the "Switching Overview," Example 4 is an example in which switching is performed due to a response delay of an LLM on the network. In Example 4, the "state before switching of an LLM on the network (API-connected LLM)" indicates a state in which a response is obtained from an LLM on the network within a predetermined time. Here, in Example 4, the "condition for switching" indicates a case in which a response from an LLM on the network is not obtained within the predetermined time, but exceeds the predetermined time. Furthermore, in Example 4, the "destination for switching from an LLM on the network (API-connected LLM)" indicates a "local LLM." The "local LLM" of the switching destination is the same as in Example 1, so repeated explanation will be omitted. That is, in Example 4, even if for some reason a response from an LLM (large-scale language model connected using an API) on the network exceeds a predetermined time and response generation processing by the LLM (large-scale language model connected using an API) on the network cannot be used smoothly, response generation processing is switched to response generation processing by the local LLM processing unit 10028 of the AI response output device 10010. This makes it possible to continue response generation processing using the large-scale language model, despite differences in performance as a large-scale language model.
[0238] Next, Example 5 of FIG. 5A will be described. Example 2 is an example in which the "switching destination from the networked LLM (API-connected LLM)" in Example 4 is changed from "local LLM" to "prepared response DB (database)." The response generation process using the "prepared response DB (database)" is similar to the process described in FIG. 1C , FIG. 2L , or FIG. 4B , and therefore a repeated description will be omitted. That is, in Example 5, if a response from an LLM (large-scale language model connected using an API) on the network exceeds a predetermined time for some reason and the response generation process using the LLM (large-scale language model connected using an API) on the network cannot be used smoothly, by switching to response generation process using a preparatory response database, it becomes possible to generate a response through simpler processing and output the response to the user.
[0239] Next, Examples 6 to 9 of FIG. 5A will be described. As shown in "Switching Overview," Examples 6 to 9 are examples in which switching occurs when the API usage or usage fee limit is reached. As described in Example 2, providers of large-scale language models often collect the costs used to train the large-scale language model from users of the terminal as a usage fee for the terminal's API. In such cases, with natural language models, the API usage fee is often charged based on the number of processing times of units of words that separate sentences, called tokens. Here, various methods of charging and limiting API usage fees are conceivable. One possible example is to define the upper limit of the amount of large-scale language model usage services that a user can receive under normal circumstances using the number of tokens processed.
[0240] In this case, the user may be able to receive services using large-scale language models at a specified API usage fee until the usage volume (or corresponding usage fee) is reached, and once the upper limit of the usage volume (or corresponding usage fee) is reached, certain restrictions may be imposed, such as the user being unable to receive services using large-scale language models in their normal state (performance or frequency).
[0241] Examples 6 to 9 in FIG. 5A are examples of response generation process switching control by the control unit 1110 of the AI response output device 10010 when such restrictions occur in a service using a large-scale language model. Specifically, in Example 6, the "state before switching the LLM (API-connected LLM) on the network" is a state in which the API usage volume and API usage fee have not reached a predetermined upper limit. This means that the usage volume of the LLM (API-connected LLM) on the network has not reached a predetermined upper limit. In this case, the user can use the LLM (API-connected LLM) on the network in a normal state.
[0242] Here, in Example 6, the "switching condition" is when the API usage volume or API usage fee reaches a predetermined upper limit. This means that the usage volume of an LLM on the network (API-connected LLM) reaches a predetermined upper limit. Also, in Example 6, the "switching destination from an LLM on the network (API-connected LLM)" is a second LLM on the network that is different from the LLM (which may be referred to as the first LLM) used in the normal state. An example of a second LLM on the network is an LLM that charges a lower fee than the first LLM used in the normal state. Because it is a lower-cost service, the performance of the second LLM may be lower than that of the first LLM. Even in this case, there is still a significant advantage if a large-scale language model can be used inexpensively even after the usage volume / fee limit of the first LLM is reached.
[0243] Next, Example 7 in FIG. 5A will be described. In Example 7, the "switching destination from the networked LLM (API-connected LLM)" in Example 6 is changed from a second LLM on a network different from the LLM (which may be referred to as the first LLM) used in the normal state to a "local LLM." In Example 7, even if the API usage or API usage fee reaches a predetermined upper limit, i.e., even if the usage of the networked LLM (API-connected LLM) reaches a predetermined upper limit, it is possible to continue response generation processing using a large-scale language model by switching to response generation processing using a local LLM that is not subject to restrictions such as the usage amount of the networked LLM, the API usage amount, or the API usage fee.
[0244] Next, Example 8 of FIG. 5A will be described. In Example 8, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "prepared response DB (database)." The response generation process using the "prepared response DB (database)" is similar to the process described in FIG. 1C , FIG. 2L , or FIG. 4B , and therefore will not be described again. In Example 8, even if the API usage volume or API usage fee reaches a predetermined upper limit, i.e., even if the usage volume of the networked LLM (API-connected LLM) reaches a predetermined upper limit, the process switches to a response generation process using a preparatory response database that is not limited by the usage volume of the networked LLM, the API usage volume, or the API usage fee. This enables a response to be generated and output to the user through simpler processing.
[0245] Next, Example 9 in FIG. 5A will be described. In Example 9, the "switching destination from the networked LLM (API-connected LLM)" in Example 7 is changed from "local LLM" to "no-response handling." The "no-response handling" refers to a handling in which no response is generated to the user or no response is output to the user. In Example 9, it becomes possible to more easily handle a situation in which response generation processing using a large-scale language model on a network (a large-scale language model connected using an API) becomes unavailable due to the API usage or API usage fee reaching a predetermined upper limit, i.e., the usage of the networked LLM (API-connected LLM) reaching a predetermined upper limit.
[0246] According to the switching control of the response generation process of the AI response output device 10010 shown in Examples 1 to 9 of Figure 5A as described above, even in situations where the response generation process using an LLM (large-scale language model connected using an API) on the network cannot be used as usual, more appropriate switching or response can be made according to each situation.
[0247] Note that the switching control of Examples 1 to 9 in FIG. 5A may be performed by combining a plurality of examples. For example, the switching control of Examples 1 to 3 may be combined with any of the controls of Examples 4 to 9, respectively. Similarly, the control of Example 4 or Example 5 may be combined with any of the controls of Examples 1 to 3 or Examples 6 to 9, respectively. Similarly, the control of Examples 6 to 9 may be combined with any of the controls of Examples 1 to 5, respectively.
[0248] Next, using Figures 5B to 5D, an example of displaying an AI assistant or character when the artificial intelligence response output device 10010 of Example 5 is configured as an AI assistant device or a character conversation device will be described.
[0249] First, Figure 5B is an example of the display of an AI assistant or character in the artificial intelligence response output device 10010 when performing the switching control of Example 3 in Figure 5A. In the example of Figure 5B, the display state of the AI assistant or character is changed depending on whether the network connection status of the artificial intelligence response output device 10010 is network connectable or network unconnectable. The network connectable and network unconnectable states of the artificial intelligence response output device 10010 are as described in Figure 5A, so a repeated description will not be provided.
[0250] In the example of Figure 5B, the artificial intelligence response output device 10010 (1) displays the AI assistant or character in a normal awake state when a network connection is possible, but (2) displays the AI assistant or character in a "sleeping" state when a network connection is not possible. In the switching control of example 3 of Figure 5A, when the artificial intelligence response output device 10010 cannot connect to the network, it does not generate a response or output a response even if a command is input from the user. In this case, if the AI assistant or character displayed by the artificial intelligence response output device 10010 is in a normal awake state, the user will feel uncomfortable. However, if the AI assistant or character displayed by the artificial intelligence response output device 10010 is displayed in a sleeping state, the user will understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the discomfort felt by the user.
[0251] In the case of Figure 5B (2), it is desirable that the user understand that "the reason the AI assistant or character is not responding is because it is sleeping" before making a user input requesting a response using a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 1001. Therefore, it is desirable that the start timing of the state in Figure 5B (2) where the AI assistant or character is displayed in a "sleeping" state when network connection is not possible, be immediately after the control unit 1110 of the AI response output device 10010 determines that network connection is not possible, before the user makes a user input requesting a response using a large-scale language model.
[0252] Next, as another display example, the display example of FIG. 5C will be described. The display example of FIG. 5C is an example in which the display state of the AI assistant or character is changed depending on the state of the "switching destination from the LLM on the network (API-connected LLM)" in the table during the switching control of FIG. 5A. Specifically, FIG. 5C shows (1) a display example of the AI assistant or character in a state in which the AI response output device 10010 can connect to a large-scale language model on the network (a large-scale language model connected using an API) and is able to use response generation processing using the large-scale language model on the network (referred to as the normal state in this figure), (2) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to response generation processing using an LLM or a fixed response phrase database with lower performance than the large-scale language model on the network (a large-scale language model connected using an API), and (3) a display example of the AI assistant or character in a state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A.
[0253] In the example of FIG. 5C , for example, (1) when the artificial intelligence response output device 10010 is in a "normal state," the artificial intelligence response output device 10010 displays the AI assistant or character in a state where there are no particular problems. Note that the "normal state" in FIG. 5C may be considered a state other than states (2) and (3). Also, for example, (2) when the artificial intelligence response output device 10010 has switched to a response generation process using an LLM or a response template database, which has lower performance than a large-scale language model on the network (a large-scale language model connected using an API), the artificial intelligence response output device 10010 displays the AI assistant or character in a "sleepy" state. Note that "displaying the AI assistant or character in a "sleepy" state" may also be expressed as "a display indicating that the AI assistant or character is feeling drowsy."
[0254] The response generation process (2) has lower performance than the response generation process using a large-scale language model on the network (a large-scale language model connected using an API) in the normal state (1). Therefore, by displaying the AI assistant or character in a "sleepy" state, it is possible to implicitly convey to the user that the response performance of the AI assistant or character is low. This makes it possible to further reduce the sense of discomfort felt by the user due to a low-performance response. Note that the switching conditions under which the artificial intelligence response output device 10010 switches to a response generation process using an LLM or a fixed response phrase database, which has lower performance than the large-scale language model on the network (a large-scale language model connected using an API), are as described in FIG. 5A, and therefore a repeated explanation will be omitted.
[0255] 5C (2), it is desirable to implicitly inform the user that the response performance of the AI assistant or character is low before the user makes a user input requesting a response from a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 10010. Therefore, it is desirable that the start timing of the state in which the AI assistant or character is displayed in a "sleepy" state in FIG. 5C (2) be immediately after the artificial intelligence response output device 10010 switches to a response generation process using an LLM or a fixed response phrase database that has lower performance than the large-scale language model on the network (the large-scale language model connected using an API), before the user makes a user input requesting a response from a large-scale language model.
[0256] Further, for example, in (3) the state in which the AI response output device 10010 has switched to the no-response mode described in FIG. 5A , the AI response output device 10010 displays the AI assistant or character in a "sleeping" state. As also described in FIG. 5B , by displaying the AI assistant or character displayed by the AI response output device 10010 in a "sleeping" state, the user can understand that "the AI assistant or character is not responding because it is sleeping," which can further reduce the sense of discomfort felt by the user. Note that the conditions under which the AI response output device 10010 switches to the no-response mode described in FIG. 5A are the same as those described in Example 3 or Example 9 of FIG. 5A , and therefore a repeated description will be omitted. Note that in the case of FIG. 5C (3), it is desirable for the user to understand that "the AI assistant or character is not responding because it is sleeping" before making a user input requesting a response using a large-scale language model via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010. Therefore, it is desirable that the timing for starting the state (3) in Figure 5C in which the AI assistant or character is displayed in a "sleeping" state be immediately after the artificial intelligence response output device 10010 switches to the no-response response described in Figure 5A, before the user input requesting a response using a large-scale language model.
[0257] 5C , the AI response output device 10010 displays a technical explanation of the state of the AI response output device 10010 related to the response generation process to the user, without directly explaining the state to the user, implicitly reflecting the change in the state of the AI assistant or character. This can further reduce the sense of discomfort felt by the user compared to when a technical explanation of the state of the AI response output device 10010 related to the response generation process is directly given to the user. Furthermore, this can further reduce the sense of discomfort felt by the user compared to when the display state of the AI assistant or character remains the same as its normal state despite a change in the state of the AI response output device 10010 related to the response generation process.
[0258] However, some users may wish to know a more precise explanation of the technical state of each state. Therefore, a display example for such users will be described with reference to FIG. 5D . Among the rows of the table shown in FIG. 5D , the rows for explaining the device state and display state are identical to those in FIG. 5C , and therefore, repeated explanations will be omitted. Furthermore, the display example of the AI assistant or character shown in the row for the display example of the AI assistant or character is almost identical to that in FIG. 5C , except that a question mark (?) is displayed in the display example. The question mark (?) is a mark that the user operates when requesting an explanation from the AI response output device 10010, and may also be referred to as a help mark.
[0259] In the example of FIG. 5D , when a user selects the question mark (?) by user operation via the operation input unit 1107 or the touch panel of the display unit 10011 in FIG. 1B , the display of the AI assistant or character on the AI response output device 10010 changes to the display example shown in the row of the display example after user operation. Specifically, regardless of whether the device state is (1), (2), or (3), an explanation of the technical state of each state is displayed. For example, in the example of FIG. 5D , if the device state is (1) normal, a display may be displayed explaining that the device is in a normal state with no particular technical limitations, such as "normal state." Furthermore, if the device state is (2) using a low-performance LLM or a standard response phrase database, a display may be displayed explaining that the device is in a low-performance mode, such as "low-performance mode." This display may be considered to explain the reason why the AI assistant or character is displayed in a "sleepy" state.
[0260] In this case, a more technically detailed explanation may be provided. Specifically, a display such as "Low-performance LLM usage mode" or "Canned response mode" may be displayed. Furthermore, if the device status is (3) unresponsive, a display such as "Network connection unavailable" may be displayed, providing a technical explanation of the cause of the switch to unresponsive mode. If the cause of the switch to unresponsive mode is that a response from an LLM (large-scale language model connected using an API) on the network exceeds a predetermined time, a display such as "Response from the LLM is delayed" may be displayed. Furthermore, if the cause of the switch to unresponsive mode is that the usage amount, API usage amount, or API usage fee of the LLM on the network has reached its upper limit, a display such as "LLM usage limit reached," "API usage limit reached," or "API usage fee has reached a predetermined amount" may be displayed. These displays may be considered to explain the cause of the AI assistant or character display being displayed in the "asleep" state.
[0261] According to the display example of Figure 5D described above, even if there are technical constraints in the response generation process in the AI response output device 10010, first, by implicitly indicating the device state by changing the display state of the AI assistant or character without providing a direct explanation to the user, it is possible to further reduce the sense of discomfort felt by the user. This display is more suitable for users who do not need technical explanations. Furthermore, by displaying an operation mark to explain the technical state, a display is provided to users who operate the mark that technically explains the state of the response generation process in the AI response output device 10010 (normal state or state with technical constraints). This makes it possible to provide a more suitable display for users who want to know the technical state accurately.
[0262] In the examples of Figures 5B, 5C, and 5D, a "sleeping" state is shown as an example of the display state of the AI assistant or character when the AI assistant or character is in a "no response" state, but this is just one example, and the embodiment is not limited to this. Instead of the "sleeping" state, another display state that implies a situation where the AI assistant or character is unable to respond, such as "taking a break," may be used. In the examples of Figures 5C and 5D, a "sleepy" state is shown as an example of the display state of the AI assistant or character when a low-performance LLM or a standard response phrase database is being used, but this is just one example, and the embodiment is not limited to this. Alternatively, another display state that implies that the AI assistant or character has low response performance, such as "hungry," may be used.
[0263] According to the AI response output device and AI response output system of Example 5 described above, it becomes possible to more appropriately switch the response generation process used by the AI response output device depending on the connection state between the large-scale language model on the network and the AI response output device, the response delay state from the large-scale language model on the network, the usage amount of the large-scale language model on the network, etc. Furthermore, when the AI response output device of Example 5 is configured as an AI assistant device or a character conversation device, it becomes possible to perform a display that is less strange to the user.
[0264] Example 6 Next, Example 6 of the present invention is an improvement on the AI response output device 10010 or the AI response output system described in the drawings of Examples 1 to 5. Specifically, this is an example in which the response output is generated by more suitably combining the response generation process of the AI response output device 10010 with a response generation process using a large-scale language model on a network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, and a response generation process using a fixed response phrase database. In this example, differences from these examples will be described, and repeated description of configurations similar to these examples will be omitted.
[0265] As in the above-described embodiments, the AI response output device 10010 may be referred to as an AI response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an AI interface device. A system including the AI response output device 10010 and the large-scale language model server may be referred to as an AI response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an AI interface system.
[0266] An example of a response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 shows an example of a flowchart of the response generation process in the AI response output device 10010 according to the sixth embodiment of the present invention. Specifically, a time axis in which time progresses from top to bottom, a processing flow, and an example response output are shown. The response shown in the example response output may be output via display on the display unit 10011 of the AI response output device 10010 or audio output by the audio output unit 1140.
[0267] In the example of Figure 6, first, at time t0, a user input requesting a response based on a large-scale language model is made by the user via the touch panel, microphone 1139, or operation input unit 1107 of the AI response output device 10010, and the control unit 1110 of the AI response output device 10010 acquires the user input (step 600). Next, at time t1, the control unit 1110 begins preparation for response output using the fixed response phrase database stored in the storage unit 1170, and begins response output using the fixed response phrase database (step 601). In the example of Figure 6, response output using the fixed response phrase database begins at time t2, and as shown in the figure, the fixed response is being output but has not yet been completed. "Good morning" in the figure indicates the output of part of the sentence that continues "Good morning..."
[0268] At time t3, before the response output using the fixed response phrase database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and transmits the generated instruction sentence to a large-scale language model on the network or a local large-scale language model (such as the local LLM processing unit 10028) provided in the AI response output device 10010, thereby starting a request for a response from the large-scale language model (step 602). Furthermore, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 starts acquiring a response from the large-scale language model (step 603).
[0269] An example of a response output in which the response output using the fixed response phrase database is completed at time t5 is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is the ____ day of the month." is completed at time t5 using the fixed phrases stored in the fixed response phrase database and date information stored in memory. Here, at time t4, before the response output using the fixed response phrase database is completed, the control unit 1110 has already started acquiring a response from the large-scale language model. Therefore, at time t6, which follows time t5 when the response output display using the fixed response phrase database is completed, the control unit 1110 starts outputting a response from the large-scale language model following the response output using the fixed response phrase database (step 604). Then, at time t7, a response from the large-scale language model is output following the response output using the fixed response phrase database. When the response output from the large-scale language model is completed, the response output according to the processing flow shown in FIG. 6 is completed (step 605).
[0270] Next, the effect of the processing flow shown in FIG. 6 of the present invention will be described. Processing a large-scale language model requires a large amount of computational resources. Even if inference, which generally requires fewer computational resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several to several tens of seconds from the time the control unit initiates a response request to the large-scale language model until it is able to obtain a response from the large-scale language model. This period corresponds to the time t3 to t4 shown in FIG. 6. Furthermore, from time t0, when a user input is made, until time t4, the control unit 1110 is unable to obtain a response output from the large-scale language model, and therefore is unable to output a response from the large-scale language model to the user.
[0271] 6 , there is no start of preparation for response output using the fixed response phrase database, and no start of response output using the fixed response phrase database, the user may continue to wait for a period of several seconds to more than ten seconds from time t0 when the user input is made until time t4 without receiving a response from the AI response output device 10010. For example, when the AI response output device 10010 is configured as an AI assistant device or a character conversation device, this waiting time may cause the user to feel uncomfortable.
[0272] In contrast, in the processing flow according to the sixth embodiment of the present invention shown in Fig. 6, the control unit 1110 starts the process of outputting a response using a fixed response phrase database, which requires fewer computational resources than the process of a large-scale language model, before starting to acquire a response from the large-scale language model. This prevents the user from having to wait from time t0 to time t4 without receiving a response from the AI response output device 10010. From the user's perspective, whether the response output is using a fixed response phrase database or a response output from a large-scale language model, it is the same as receiving a response from the AI response output device 10010.
[0273] 6, by providing step 601 before step 603, the response of the AI response output device 10010 to the user can be artificially accelerated. This can further reduce the sense of discomfort felt by the user due to long waiting times. Furthermore, by outputting a response from a large-scale language model following a response using the template response database in step 604, the user can perceive these outputs as if they were a series of more natural outputs.
[0274] According to the AI response output device and AI response output system of Example 6 described above, the waiting time for a response from the AI response output device can be shortened, and the sense of discomfort felt by the user can be further reduced.
[0275] <Embodiment 7> Embodiment 7 of the present invention is an improvement of the AI response output device 10010 or the AI response output system described in embodiment 1 etc. In this embodiment, differences from embodiment 1 will be described, and repeated explanations of configurations similar to those of these embodiments will be omitted.
[0276] The AI response output system according to Example 7 is an example in which a conversion process is performed to convert a user's question as needed when an instruction sentence is sent from the AI response output device 10010 to a large-scale language model. More specifically, when creating an instruction sentence (prompt) for a large-scale language model AI from a user's question sentence, a conversion process is performed to convert the user's question as needed, so that a more appropriate answer can be obtained for the user's question sentence.
[0277] An example of a response output system according to a seventh embodiment of the present invention will be described with reference to Fig. 7. As shown in Fig. 7, the configuration of the response output system according to the seventh embodiment is basically the same as that according to the first embodiment. In the seventh embodiment, an AI response output device 10010 is also communicatively connected to a large-scale language model server 19001 or a large-scale language model server 20001 equipped with a large-scale language model, and a second server 19002 different from these servers, via the Internet 19000. However, this embodiment differs from the first embodiment in that the second server 19002 includes a check database (check DB) 19010.
[0278] The checking DB 19010 will be described in detail later. Furthermore, in Example 7, a configuration will be described in which the large-scale language model server 19001 is equipped with a large-scale language model (LLM). In the following description, the large-scale language model equipped in the large-scale language model server 19001 may be simply referred to as the large-scale language model or LLM. Incidentally, the large-scale language model may be a large-scale language model on a network, or a local large-scale language model (such as the local LLM processing unit 10028) equipped in the AI response output device 10010.
[0279] When a user asks a question to a large-scale language model, which is an artificial intelligence, and attempts to obtain a response (answer), the process is, for example, as follows: First, the user inputs a question to the large-scale language model to the artificial intelligence response output device 10010. As an example, the user operates the operation input unit 1107, which is an input unit, to input a question sentence, which is text information in a natural language. The question input to the large-scale language model may be operated using the operation input unit 1107 or may be input by voice. The artificial intelligence response output device 10010 generates a response instruction sentence (prompt) directed to the large-scale language model based on the question sentence, and transmits it to the large-scale language model server 19001 equipped with the large-scale language model. When the large-scale language model generates a response to the response instruction sentence (e.g., a response sentence to the user's question), the generated response is transmitted from the large-scale language model server 19001 to the artificial intelligence response output device 10010.
[0280] Upon receiving this response, the AI response output device 10010 processes the response of the large-scale language model (LLM) for the user. As an example of processing the response of the large-scale language model, the AI response output device 10010 displays an answer sentence (also called a response sentence) generated by the large-scale language model on the display unit 10011. In this manner, the user can obtain an answer to the question from the large-scale language model.
[0281] However, users may not always be able to obtain the answers they expect from large-scale language models. For example, depending on how a question is written, the content of the question may be interpreted differently from the user's intention, and the user may not be able to obtain the answer they want from the large-scale language model.
[0282] Therefore, in the response output system according to Example 7, when a user inputs a question to the AI response output device 10010, the AI response output device 10010 executes a conversion process on the question as necessary, and generates a response instruction sentence based on the question after the conversion process (hereinafter, sometimes referred to as a processed question). This allows the question to be interpreted more appropriately by the large-scale language model, making it easier for the user to obtain the answer they want from the large-scale language model.
[0283] The method for inputting a question by a user to the AI response output device 10010 may be a method of converting a question inputted by voice using the operation input unit 1107 as an input unit into text using the microphone 1139. The process of converting the question, the process of generating an instruction sentence, and the process of outputting an answer from the large-scale language model to the display unit 10011 are executed by, for example, the control unit 1110. The control unit 1110 in this embodiment generates a response instruction sentence for the large-scale language model based on the question sentence, and acquires the response sentence generated by the large-scale language model in response to the response instruction sentence. Specifically, the control unit 1110 executes a process of converting the question sentence and generates a response instruction sentence based on the question sentence after the conversion process.
[0284] The response output processing flow in the response output system according to Example 7, mainly the processing flow when obtaining a response from a large-scale language model to a user's question, will be described in detail below. Figures 8, 9, and 11 are diagrams showing an example of the processing flow of the response output system according to Example 7. Figure 10 is a diagram showing an example of a check database stored in the second server according to Example 7. Figure 12 is a diagram showing an example of a database matching priority table. Figures 13 and 14 are diagrams showing an example of a modified example of the processing flow of the response output system according to Example 7. Note that in the diagrams showing the processing flow of the response output system, the same steps are assigned the same symbols, and duplicate explanations of each step may be omitted.
[0285] First, a basic processing flow of the response output system according to the seventh embodiment will be described with reference to Fig. 8 . As shown in Fig. 8 , a user first operates the operation input unit 1107 to input a question for a large-scale language model. In step S810, the control unit 1110 receives the user input (question). In step S820, the control unit 1110 executes a check process to determine whether conversion processing of the question input by the user is necessary. Also, in step S820, the control unit 1110 executes conversion processing of the question as necessary based on the results of this check process. These check processes and conversion processes will be described in detail below.
[0286] Next, the process proceeds to step S830, where an instruction (prompt) is generated for the large-scale language model (LLM) provided in the large-scale language model server 19001 based on the question entered by the user or the question that has undergone the conversion process (hereinafter referred to as the converted question). In this example, a response instruction is generated that instructs the creation of an answer to the user's question. This response instruction is then sent to the large-scale language model (step S840). Note that the above response instruction, along with the conversion instruction, summarization instruction, and addition instruction described below, are collectively referred to as "instructions."
[0287] Next, when the large-scale language model (LLM) receives the instruction sentence, for example, a response instruction sentence, sent by the AI response output device 10010 (step S850), it generates a response based on this instruction sentence (step S860). In the example shown in Figure 8, the instruction sentence is a response instruction sentence, so the large-scale language model creates a response sentence to the user's question as the response generation in step S860. Then, in step S870, the response sentence created by the large-scale language model is sent to the AI response output device 10010.
[0288] When the control unit 1110 of the AI response output device 10010 receives a response (e.g., an answer sentence) created by the large-scale language model (step S880), the control unit 1110 executes processing on the response of the large-scale language model in step S890. As an example of processing the response of the large-scale language model, the control unit 1110 causes the display unit 10011 to display the answer sentence to the user's question.
[0289] As described above, in the response output system according to the seventh embodiment, when performing the response output process, the question input by the user is checked and, if necessary, the question is converted. This makes it easier to obtain the answer intended by the user from the large-scale language model in response to the user's question.
[0290] Next, the question checking and conversion processing in step S820 will be described in more detail. For example, as shown in Fig. 9, when control unit 1110 receives a user input in step S810, control unit 1110 first checks the question input by the user and determines whether conversion processing of the question is necessary (step S8210).
[0291] If it is determined as a result of this check process that question conversion processing is not required (step S8210: No), the process proceeds to step S830, where an instruction sentence is generated based on the question sentence. On the other hand, if it is determined that question conversion processing is required (step S8210: Yes), the process proceeds to step S8220, where question conversion processing is executed. In this example, the question conversion processing involves requesting the user to rewrite the question sentence (step S8220).
[0292] To explain the check process of step S820 in more detail, the control unit 1110 executes a process of checking whether the question corresponds to a predetermined check item as the check process. As an example, the control unit 1110 executes a process of checking whether the question contains a predetermined specific expression. In other words, one of the check items is whether the question contains the specific expression. The content of the check item is not particularly limited, and may be any item that can determine whether conversion processing of the question is necessary.
[0293] In this example, the specific expressions refer to words, phrases, grammar, etc. that are pre-registered in a check database (check DB) 19010 of the second server 19002. That is, a list of specific expressions is stored in the check DB 19010 of the second server 19002. As an example, the specific expressions stored in the check DB 19010 include specific words, phrases such as "technical terms," "abbreviations," and "homonyms / homographs," as well as specific grammar, such as "ambiguous sentence expressions" and "dialects," as shown in FIG. 10 . Note that, in this example, the meanings or explanations of the specific words and phrases registered in the check DB 19010 are also stored.
[0294] As shown in FIG. 10 , the specific expressions stored in the checking DB 19010 are classified into multiple groups, including "ambiguous sentence expressions," "abbreviations," "technical terms," "homonyms / homographs," and "dialects." Furthermore, each group is set up into multiple hierarchical levels. That is, each hierarchical level group is further classified into multiple groups in the next hierarchical level. For example, the "ambiguous sentence expressions" group is classified into "ambiguous sentence expression category 1" and "ambiguous sentence expression category 2," and these "ambiguous sentence expression category 1" and "ambiguous sentence expression category 2" groups are further classified into multiple groups. Similarly, the "abbreviations" group is classified into "abbreviation category 1" and "abbreviation category 2," which are further classified into "abbreviation category 1" and "abbreviation category 2."
[0295] In this example, the groups of specific expressions are set to three levels. However, the number of levels of the groups is not particularly limited and may be determined as appropriate. The number of groups set to the same level is also not particularly limited and may be determined as appropriate.
[0296] The control unit 1110 then executes a question check process using the check DB 19010. For example, as shown in Fig. 11 , as part of the question check process in step S820, the control unit 1110 first executes a database matching process for the question in step S8211. In this matching process, the words and grammar in the question entered by the user are matched with specific expressions registered in the check DB 19010.
[0297] Next, in step S8212, it is determined whether or not the words and grammar in the question match any of the specific expressions registered in the check DB 19010. In other words, the control unit 1110 performs processing to confirm whether or not the question contains words and grammar that fall into the multiple groups illustrated in Fig. 10, that is, whether or not the question contains any of the specific expressions registered in the check DB 19010.
[0298] If the words or grammar in the question are not found in the check DB 19010 (step S8212: No), it is determined that the question does not correspond to the above check item, and the process proceeds to step S830, where an instruction statement is generated. On the other hand, if the words or grammar in the question are found in the check DB 19010 (step S8212: Yes), it is determined that the question corresponds to the above check item, and the process proceeds to step S8220, where a process is performed to request the user to rewrite the question as a question conversion process. The control unit 1110, for example, causes the display unit 10011 to display information requesting the user to rewrite the question.
[0299] At this time, it is preferable to also display the reason for requesting a revision of the question. For example, it is preferable to also display that the question contains a specific expression and the words / grammar corresponding to the specific expression contained in the question. Note that the method for requesting the user to rewrite the question is not particularly limited, and may be, for example, by voice.
[0300] Furthermore, the question rewritten by the user in response to the rewriting request as part of this conversion process corresponds to the "converted question." In the example shown in FIG. 9 , when the user rewrites the question in response to the request in step S8220 and the control unit 1110 receives the rewritten question (converted question) (step S810), a check process is performed on the converted question to determine whether further conversion processing of the converted question is required (step S8210). If it is determined that further conversion processing of the converted question is required, the control unit 1110 again requests the user to rewrite the converted question (step S8220). If it is determined that further conversion processing of the converted question is not required (step S8210: No), the process proceeds to step S830, where a directive is generated based on the converted question. The subsequent processing flow is similar to that of the example in FIG. 8 , and therefore will not be described again.
[0301] Meanwhile, the AI response output device 10010 is preset with priorities for matching a question sentence with specific expressions registered in the check DB 19010. For example, as shown in Fig. 12, the AI response output device 10010 stores in advance a DB matching priority table that sets the priority for matching a question sentence with the check DB 19010. In this example, matching priorities 1 to 5 are set for five groups in the check DB 19010: "Ambiguous sentence expressions," "Abbreviations," "Terminology," "Homophones / Homographs," and "Dialects."
[0302] In this case, in steps S8211 and S8212 shown in Fig. 11, the group set to priority 1 is processed (matched) first, and the group set to priority 5 is processed last. In the example shown in Fig. 12, in the case of default settings, the "ambiguous sentence expression" in the question sentence is matched first, and the "dialect" in the question sentence is matched last.
[0303] The priority of the groups to be compared can be set arbitrarily, for example, according to the user's habits and preferences in writing questions. Furthermore, the types and number of groups to be compared can also be set arbitrarily. The groups set as groups to be compared are referred to as "set groups."
[0304] 12, "User 1 personal settings" and "User 2 personal settings" are examples in which five groups, "Ambiguous sentence expressions," "Abbreviations," "Terminology," "Homophones / Homographs," and "Dialects," are set as "setting groups," and the priorities of these five setting groups are arbitrarily set by the user. Also, "User 3 personal settings" is an example in which three groups, "Dialects," "Ambiguous sentence expressions," and "Homophones / Homographs," are set as "setting groups," and the priorities of the three setting groups are arbitrarily set by the user.
[0305] By appropriately setting the priority of the groups to be compared in this way, the conversion process of the question can be executed efficiently. For example, by setting a higher priority for a group of specific expressions that are likely to be used in questions, it becomes easier to perform the comparison of the question with the check DB 19010 and the conversion process of the question in parallel. This reduces the time required to obtain an answer from a large-scale language model for the question.
[0306] An example of the processing flow of the response output system of Example 7 has been described above, but the processing flow for obtaining a response from a large-scale language model to a user's question is not limited to the above example. In the above example, a case has been described in which the processing for requesting the user to rewrite the question is performed as the processing for converting the question, but the processing for converting the question is not limited to this processing.
[0307] Next, another example of the question sentence conversion process will be described with reference to Figures 13 to 15. The processing flow in Figure 13 shows an example in which the question sentence conversion process is executed by the control unit 1110 of the AI response output device 10010. The control unit 1110 may also be referred to as a processor.
[0308] 13 , if the words and grammar in the question are found in the check DB 19010 in step S8212 (step S8212: Yes), the process proceeds to step S8213, where the control unit 1110 converts the question as a question conversion process. For example, the control unit 1110 replaces specific expressions used in the question with other expressions (words and grammar) and converts the question into a sentence that does not use specific expressions. Note that if the words and grammar in the question are not found in the check DB 19010 in step S8212 (step S8212: No), the process proceeds to step S830, where a command is generated based on the question entered by the user, as in the above example.
[0309] Furthermore, if the words and grammar in the question are found in the check DB 19010 in step S8212 (step S8212: Yes), the control unit 1110 converts the question in step S8213, and then in step S8214, the control unit 1110 performs a process of inquiring the user about whether the converted question matches the intention of the original question. For example, confirmation items for confirming whether the converted question matches the intention of the original question are displayed on the display unit 10011, and the user is requested to respond to these confirmation items.
[0310] Next, in step S8215, it is confirmed whether the converted question matches the intention of the original question (user's intention). For example, it is determined whether the converted question matches the user's intention based on the user's response to a confirmation item displayed on the display unit 10011. Specifically, a confirmation item such as "Does the converted question match the intention of the original question?" is displayed, along with a "Yes / No" option as the user's response to that question. If "Yes" is selected, it is determined that the converted question matches the user's intention, and if "No" is selected, it is determined that the converted question does not match the user's intention.
[0311] If the user replies that the converted question matches the user's intention (step S8215: Yes), the process proceeds to step S830, where the control unit 1110 generates an instruction sentence based on the converted question (converted question sentence) and transmits the generated instruction sentence to the large-scale language model (step S840). On the other hand, if the user replies that the converted question sentence does not match the user's intention (step S8215: No), the process proceeds to step S8220, where further question sentence conversion processing is performed. In the example shown in Fig. 13, as in the example shown in Fig. 9, etc., the question sentence conversion processing in step S8220 involves requesting the user to rewrite the question sentence.
[0312] Note that a question rewritten by the user in response to this request also corresponds to a "converted question" and is treated in the same manner as in the above example. In the example shown in FIG. 13, the processing flow after the large-scale language model receives the instruction sentence in step S850 is the same as the example shown in FIG. 8, etc., and is therefore not shown. Furthermore, step S820 is a system process, and may be a process other than the above. Furthermore, the processes of steps S8214 to S8220 do not necessarily have to be performed. For example, once the question sentence is converted in step S8213, the process may proceed to step S830, where a response instruction sentence is created based on the converted question sentence.
[0313] The process flows in Figures 14 and 15 illustrate an example in which the large-scale language model (LLM) of the large-scale language model server 19001 performs the conversion process for the question sentence. In the example shown in Figure 14, if the control unit 1110 determines in step S8210 that the conversion process for the question sentence is not necessary (step S8210: No), a response instruction sentence is generated based on the question sentence in step S830, as in the above example. Then, the generated instruction sentence (response instruction sentence) is sent to the large-scale language model (step S840). Upon receiving this response instruction sentence (step S850), the large-scale language model proceeds to step S860 and generates a response to the response instruction sentence. Specifically, an answer to the question is generated.
[0314] Thereafter, the answer to the question is transmitted to the control unit 1110 as the large-scale language model's response to the instruction statement (response instruction statement) (step S870). Upon receiving the large-scale language model's response to the response instruction statement (step S880), the control unit 1110 then performs processing such as presenting the response of the large-scale language model, in other words, the answer of the large-scale language model to the question entered by the user, to the user in step S890. The user may check the answer of the large-scale language model, and if satisfied with the answer, may end the query to the large-scale language model or enter a new question, or if not satisfied with the answer, may enter an additional question.
[0315] On the other hand, if it is determined that conversion processing of the question sentence is necessary (step S8210: Yes), the process proceeds to step S831, where a conversion instruction statement is generated to instruct the large-scale language model to perform conversion processing of the question sentence, and the generated instruction statement (conversion instruction statement) is sent to the large-scale language model (step S840).
[0316] When the large-scale language model receives this conversion instruction (step S850), it then proceeds to step S861 and generates a response to the conversion instruction. Specifically, as a conversion process for a question created by a user, the question is rewritten. For example, if the question contains a specific expression as described above, the question is converted into a sentence that does not use the specific expression. Of course, the content of the conversion process for a question is not particularly limited. Thereafter, the question that has undergone the conversion process (converted question) is transmitted to the control unit 1110 as a response of the large-scale language model to the instruction (conversion instruction).
[0317] When the control unit 1110 receives the converted question from the large-scale language model (LLM) in response to the conversion instruction (step S880), it then asks the user in step S900 whether the question reconstructed by the large-scale language model matches the original question. In other words, it asks the user to confirm whether the converted question matches the intention of the original question.
[0318] If the user replies that the converted question sentence matches the original question sentence (step S900: Yes), the process proceeds to step S830, where an instruction sentence (response instruction sentence) is generated based on the converted question sentence converted by the large-scale language model, and the generated response instruction sentence is sent to the large-scale language model (step S840).
[0319] On the other hand, if the user replies that the converted question does not match the user's intention (step S900: No), the process proceeds to step S910, where the user is requested to rewrite the question.
[0320] It can be said that steps S900 and S910 in the processing flow of Fig. 14 correspond to steps S8214 to S8220 in the processing flow of Fig. 13. In addition, in the processing flow of Fig. 14, the processing from step S831 through step S861 and step S900 to step S910 corresponds to the question sentence conversion processing.
[0321] In the above-described processing flows of the seventh embodiment, when a question is input by a user, a check process is performed on the input question, and based on the results of the check process, a conversion process is performed on the question as needed. However, the check process on the question does not necessarily have to be performed. In other words, when a question is input by a user, a conversion process on the question may always be performed.
[0322] Specifically, in the example shown in Fig. 14, when a question is input by the user in step S810, the question is checked in step S8210, and if conversion processing is required (step S8210: Yes), a conversion instruction statement is generated (step S831). In contrast, in the processing flow of Fig. 15, when the control unit 1110 receives a question input by the user (step S810), the control unit 1110 generates a conversion instruction statement in step S831 without checking the question. Note that in the processing flow of Fig. 15 as well, the processing from step S831 through step S861 and step S900 to step S910 corresponds to the question conversion processing.
[0323] The processing flow shown in Fig. 15 will now be further described with reference to Figs. 16A and 16B. For example, assume that the user inputs a question such as "Why can't I run as fast as my brother?" as shown in Fig. 16A (step S810). In this case, in step S830, the control unit 1110 generates a conversion instruction statement in the form of a question statement plus a system instruction statement, for example. The system instruction statement instructs the large-scale language model on the response content (processing content), and in the example of Fig. 16A, it instructs the user to rephrase the question statement.
[0324] Example 1 is an example in which the question and system instructions in the conversion directive are written in natural language, with "Why can't I run as fast as my brother?" written in the question field of the conversion directive and "Please rephrase it in other words" written in the system directive field. Example 2 is an example in which the question in the conversion directive is written in natural language and the system instructions are written in code, with "Why can't I run as fast as my brother?" written in the question field of the conversion directive and "paraphrase" written in the system directive field.
[0325] Then, the large-scale language model converts the question to generate a response based on the conversion instruction (step S861). As described above, if the user's question is "Why can't I run as fast as my brother?", the large-scale language model generates a converted question as an answer by paraphrasing the question, such as "Why can't I run as slow as my brother?", as shown in FIG. 16B.
[0326] Furthermore, as a process (output to the user) for the response of this large-scale language model, the control unit 1110 displays, for example, the original question "Why can't I run as fast as my brother?" and the converted question "Why can't I run as slow as my brother?" on the display unit 10011, as shown in FIG. 16B, and displays a message requesting the user to respond as to whether these converted questions match the original question (step S900).
[0327] By having the large-scale language model perform the conversion process of the question in this way, it is possible to check whether the large-scale language model properly understands the user's question before creating an answer, making it even easier to obtain an appropriate answer to the user's question from the large-scale language model.
[0328] The process flow shown in Fig. 17 is a modified example of the process flow shown in Fig. 16, and is an example in which a process for creating a summary of a question is performed as a question conversion process. Specifically, in the process flow shown in Fig. 17, when a question is input by a user in step S810, the process proceeds to step S831 without performing any check process on the question, and a summarization instruction statement is created to instruct the creation of a summary of the question. The generated instruction statement (summarization instruction statement) is then sent to the large-scale language model (step S840).
[0329] When the large-scale language model receives this summarization instruction (step S850), it proceeds to step S861 and generates a response to the summarization instruction. Specifically, as a conversion process for the question created by the user, a summary of the question is created. In other words, the question entered by the user is converted into a sentence that concisely summarizes the main points. Then, as a response of the large-scale language model to the instruction (summarization instruction), the converted question (converted question) is sent to the control unit 1110.
[0330] The subsequent processing flow by the control unit 1110 has already been described, and therefore will not be described here. Furthermore, in the example shown in FIG. 17 , the question entered by the user is not checked; however, the question may be checked in the same manner as in the example shown in FIG. 14 . In the example of FIG. 17 , for example, the question may be checked by determining whether the number of words in the question is equal to or greater than a predetermined number, and if the number of words in the question is equal to or greater than the predetermined number, a summary instruction sentence may be created. In other words, the check items may include whether the number of words in the question is equal to or greater than a predetermined number. Of course, the content of the check process in this case is not particularly limited and may be determined as appropriate.
[0331] The processing flow shown in FIG. 17 will now be further described with reference to FIGS. 18A and 18B. Assume that in step S810, the user inputs a question, as shown in FIG. 18A, such as, "The government has officially announced a policy to reduce subsidy amounts, focusing on large-scale enterprises starting next fiscal year. Please explain the impact of this policy on society, including the keyword "corporate independence." In this case, in step S830, as an example, a summary instruction statement is generated in the form of a question statement plus a system instruction statement. Note that in this example, the system instruction statement instructs the user to summarize the question statement.
[0332] Example 1 in Figure 18A is an example in which the question and system instruction in the summary instruction are written in natural language, with the question field of the summary instruction reading, "The government has officially announced a policy to narrow subsidies to large companies starting next fiscal year and reduce the amount of subsidies. Please explain the impact of this policy on society, including the keyword "corporate independence." The system instruction field reads, "Please summarize." Example 2 in Figure 18A is an example in which the question in the summary instruction is written in natural language and the system instruction is written in code, with the question field of the summary instruction reading, "The government has officially announced a policy to narrow subsidies to large companies starting next fiscal year and reduce the amount of subsidies. Please explain the impact of this policy on society, including the keyword "corporate independence." The system instruction field reads, "summarize."
[0333] Then, the large-scale language model converts the question to generate a response based on the summary instruction (step S861). As described above, if the user's question is "The government has officially announced a policy to reduce subsidies to large companies starting next year. Please explain the impact of this policy on society, including the keyword "corporate independence," the large-scale language model generates a converted question as an answer that summarizes the question, for example, as shown in FIG. 18B : "Subsidies will be provided only to large companies starting next year, and the amount will be reduced from the current amount. Please explain the impact this will have on corporate independence."
[0334] In addition, as a process (output to the user) for the response of this large-scale language model, the control unit 1110 displays, for example, as shown in FIG. 18B, the original question or request, "Subsidies will be provided only to large companies starting next year, and the amount of payment will be reduced from now. Please explain the impact this will have on companies becoming self-sufficient," and the converted question, "Subsidies will be provided only to large companies starting next year, and the amount of payment will be reduced from now. Please explain the impact this will have on companies becoming self-sufficient," on the display unit 10011, and displays a message requesting the user to answer whether these converted questions match the original question (step S900).
[0335] The above-described function for checking and converting questions does not need to be enabled at all times; for example, the user may be able to select whether to enable this function when using the system or when shutting down the system. The function for checking and converting questions is a function for correcting questions and requests, and is hereinafter also referred to as a question correction function. Figure 19A is a diagram showing an example of the processing flow when the system is started, and Figure 19B is a diagram showing an example of the processing flow when the system is shut down.
[0336] 19A, for example, an operator operates the operation input unit 1107 or the like, and the control unit 1110 receives a system startup request (step S1910). Upon receiving the system startup request, the control unit 1110 performs user authentication (step S1920). Next, in step S1930, the control unit 1110 checks whether the operator operating the response output system is a registered user. Note that the method for user authentication and checking whether the user is a registered user is not particularly limited, and any existing method may be used.
[0337] If the operator is a registered user (step S1930: Yes), the system then checks whether the usage period is within the period of the fixed-term contract for the system, function, or service (step S1940). If the system, function, or service can be used without a time limit, the system returns Yes in step S1940 and proceeds to step S1950. If the usage period is within the period of the fixed-term contract (step S1940: Yes), the system proceeds to step S1950, where the previously set user information for each user is read and, if necessary, the user information is reset or changed. The ON / OFF information for the question correction function is included in this user information. In other words, changing the user information switches the ON / OFF of the question correction function. The system then proceeds to step S1960, where the main process, such as the process of answering questions described in FIG. 7, is initiated.
[0338] If the expiration date has not passed in step S1940, that is, if the expiration date has passed (step S1940: No), the process proceeds to step S1970, where the user is notified of the expiration date. If the operator is not a registered user in step S1930 (step S1930: No), the process proceeds to step S1980, where it is confirmed whether the operator is a rental user.
[0339] If the operator is not a rental user, for example, if the operator is a new user (step S1980: No), the process proceeds to step S2000, where user registration processing is executed. On the other hand, if the operator is a rental user (step S1980: Yes), the rental user is asked in step S1990 whether or not to register as a user. If the rental user replies that they will register as a user (step S1990: Yes), the process proceeds to step S2000, where user registration processing is executed.
[0340] When this user registration process is performed, the user is then requested to set the question correction function ON / OFF as part of the user information (step S2010). If the user selects "ON" for the question correction function ON / OFF setting (step S2010: Yes), the question correction function is set to "ON." If the user selects "OFF" for the ON / OFF setting (step S2010: No), the question correction function is set to "OFF," and the user information is set. Then, the process proceeds to step S1960, where the main process begins.
[0341] If the rental user replies in step S1990 that he or she will not register as a user (step S1990: No), the process proceeds to step S2010 without performing the user registration process, and the rental user is only requested to select ON / OFF setting of the question correction function.Incidentally, the ON / OFF setting information of the question correction function for the rental user is temporarily stored only while the system is in use.
[0342] 19B, for example, a user, including a rental user, operates operation input unit 1107, and control unit 1110 receives a system termination request (step S2040). When control unit 1110 receives the system termination request, it then checks whether the user who used the system this time is a registered user (step S2050).
[0343] If the user using the system is not a registered user, for example, if the user is a rental user as described above (step S2050: No), the process proceeds to step S2060. In step S2060, user information including the ON / OFF setting of the question / request correction function for the user who currently used the system is stored. Then, system shutdown processing is performed (step S2070). Also, if the user using the system is a registered user (step S2050: Yes), the user information including the ON / OFF setting of the question correction function is already stored, so the system shutdown processing is performed in step S2070 without re-storing the information.
[0344] In this way, the ON / OFF setting of the question correction function can be performed, for example, when setting user information in the initial settings, but the setting may also be changed at any time while the system is being used. For example, by selecting a predetermined menu item displayed on the display unit 10011 while using the system, a user setting menu screen for setting user information may be displayed, and the ON / OFF setting of the question / request correction function may be changed on this user setting menu screen. Figure 20 is a diagram showing an example of the user setting menu screen.
[0345] In the example shown in Figure 20, the question / request correction function can be turned on or off, and an upper limit on the number of repetitions when the question / request correction function is turned on can be set. This allows the question / request correction function to be executed within the range of the number of repetitions set as the upper limit. As the number of repetitions of the question / request correction function increases, the accuracy of the answer to the question sentence improves, but it takes a long time to obtain the answer. Depending on the user or the content of the question, there may be cases where a quick answer is required rather than high accuracy. By setting an upper limit on the number of repetitions, it becomes easier to obtain an answer that meets the user's request, even in such cases.
[0346] Next, the processing flow of the response output system of Example 7 will be further described with reference to the processing flow of Fig. 21. Fig. 21 is a modified example of the processing flow shown in Fig. 11, and is an example of a check process in which it is determined whether a specific expression, "technical terminology," is included in the question sentence.
[0347] 21 , upon receiving a user input (step S810), the control unit 1110 performs a database matching process for the question in step S8211. In this matching process, the question entered by the user is matched with specific expressions registered in the check DB 19010. In this example, the matching process involves matching the question with technical terms registered in the check DB 19010.
[0348] Next, in step S8212, the control unit 1110 determines whether the question contains a technical term registered in the check DB 19010. If the question does not contain a technical term as a specific expression (step S8212: No), the control unit 1110 proceeds to step S830, generates a response instruction sentence, and sends the generated response instruction sentence to the large-scale language model (step S840).
[0349] On the other hand, if the question contains technical terms (step S8212: Yes), the process proceeds to step S8230, where it is determined whether the user's question inquires about the meaning of the technical term as a specific expression. If the user's question inquires about the meaning of the technical term (step S8230: Yes), the control unit 1110, for example, references the check DB 19010 and outputs the meaning or explanation of the technical term to the user. Specifically, the control unit 1110 displays the explanation of the technical term on the display unit 10011. In this embodiment, as described above, the check DB 19010 also stores the meanings and explanations of specific words and phrases such as technical terms. Therefore, the control unit 1110 can output the explanation of the term by referencing the check DB 19010. Of course, the control unit 1110 may also output the explanation of the term by referencing another database connected via the Internet, for example.
[0350] In this way, a user's question about the meaning of technical terms can be answered without using a large-scale language model. Also, if a user's question is about the meaning of technical terms, the answer the user wants cannot be presented unless the large-scale language model correctly understands the technical terms. However, by referring to the check DB 19010, the user can more easily obtain the answer they want.
[0351] If the user's question does not inquire about the meaning of technical terms in step S8230 (step S8230: No), the process proceeds to step S8240, where the question is converted. Specifically, technical terms as specific expressions contained in the question are converted into other terms (general-purpose terms) that are commonly used. Then, the process proceeds to step S830, where a response instruction sentence is created based on the converted question sentence using the general-purpose terms, and the created response instruction sentence is sent to the large-scale language model (step S840). By converting technical terms into general-purpose terms in this way, the user can more easily obtain the answer they want, even if the large-scale language model does not correctly understand the technical terms.
[0352] In the processing flow shown in Fig. 21, steps S8211, S8212, S8230, and S8240 correspond to the question sentence checking process and conversion process. In addition, the processing after the large-scale language model receives the response instruction sentence in step S850 is the same as the processing flow in Fig. 11, etc., and therefore will not be illustrated or described in Fig. 21.
[0353] The processing flow shown in Fig. 21 will be further described with reference to Fig. 22A and Fig. 22B. Fig. 22A and Fig. 22B are diagrams illustrating processing when a user's question includes technical terms. Fig. 22A is a diagram illustrating processing when a user's question inquires about the meaning of technical terms, and Fig. 22B is a diagram illustrating processing when a user's question does not inquire about the meaning of technical terms.
[0354] Assume that in step S810, the user inputs a question such as, "Please tell me the meaning of the word 'fail-safe'," as shown in FIG. 22A. The control unit 1110 then compares the question with the check DB 19010 (step S8211) and finds the technical term "fail-safe" in the check DB 19010 (step S8212: Yes). In this case, the control unit 1110 determines whether the user's question inquires about the meaning of the technical term "fail-safe" (step S8230). In this example, the control unit 1110 determines from the content of the question that the user's question inquires about the meaning of the technical term "fail-safe" (step S8230: Yes), and outputs a glossary of the term "fail-safe" to the user by referencing a check database or the like (step S825). The control unit 1110 displays, as an output to the user, a glossary of the term "fail-safe," for example, "Fail-safe means always controlling a device or system to the safe side when a failure occurs." on the display unit 10011.
[0355] Also, suppose that in step S810, the user inputs a question such as "Please tell me an example of a fail-safe that is familiar to you," as shown in FIG. 22B. Then, suppose that the control unit 1110 compares the question with the check DB 19010 (step S8211) and finds the technical term "fail-safe" in the check DB 19010 (step S8212: Yes). In this case, the control unit 1110 also determines whether the user's question inquires about the meaning of the technical term "fail-safe" (step S8230). In this example, the control unit 1110 determines, based on the content of the question, that the user's question is not inquiring about the meaning of the technical term "fail-safe" (step S8230: No).
[0356] Therefore, the control unit 1110 then creates a question (converted question) in which the technical terms in the question are replaced with general-purpose terms (step S8240). As an example, the control unit 1110 creates the converted question from the question input by the user: "Please tell me a familiar example of a mechanism for maintaining safety when a failure occurs in a device or system." Furthermore, the control unit 1110 generates a response instruction sentence based on the created converted question sentence (step S830) and sends the generated response instruction sentence to the large-scale language model (step S840).
[0357] Thereafter, when the large-scale language model receives the response instruction sentence (step S850), it creates an answer sentence to the converted question sentence as a response based on the response instruction sentence (step S860). As an example, the large-scale language model creates the answer sentence "A familiar example is that a heating appliance automatically stops operating when it detects shaking or tipping over" as a response to the converted question sentence. The created answer sentence is transmitted to the control unit 1110 (step S870). When the control unit 1110 receives the transmitted answer sentence (step S880), it outputs the answer sentence to the user's question to the user as a process for the response of the large-scale language model (step S890).
[0358] In this case, the answer generated by the large-scale language model to the user's question may contain technical terms that the user is not familiar with. In such cases, an explanation of the technical terms may be added to the answer as needed.
[0359] Fig. 23 shows a modified example after the control unit 1110 receives a response by a large-scale language model in step S880. Note that in Fig. 23, the processing flow before transmitting the response by a large-scale language model (step S870) is the same as the processing flow in Fig. 11 etc., and therefore will not be illustrated or described in Fig. 23.
[0360] As shown in FIG. 23 , when the control unit 1110 receives an answer sentence as a response from the large-scale language model (step S880), it determines whether the answer sentence requires an explanation of technical terms (step S920). The determination method here is not particularly limited, but as an example, a method similar to the question sentence check process described above can be employed. Specifically, the control unit 1110 compares the answer sentence with the check DB 19010 and determines whether an explanation of technical terms is required based on the determination result of whether the answer sentence contains technical terms registered in the check DB 19010 as specific expressions. That is, if the answer sentence contains technical terms as specific expressions, it determines that an explanation of technical terms is required (step S920: Yes). If the answer sentence does not contain technical terms as specific expressions, it determines that an explanation of technical terms is not required (step S920: No).
[0361] If it is determined that an explanation of the technical term is not necessary (step S920: No), the control unit 1110 outputs to the user an answer to the user's question transmitted from the large-scale language model as a process for the response of the large-scale language model (step S891). The case where it is determined that an explanation of the technical term is not necessary may be, for example, when an explanation of the technical term included in the answer has already been presented to the user in the past, or when the technical term is used as a general-purpose word in a certain expression due to the usage of the expression.
[0362] On the other hand, if it is determined that an explanation of the technical term is necessary (step S920: Yes), the control unit 1110 proceeds to step S832, where it creates an instruction statement (addition instruction statement) that instructs the large-scale language model to add an explanation of the technical term to the answer sentence, and sends the created addition instruction statement to the large-scale language model (step S841).
[0363] When the large-scale language model receives the additional instruction sent from the control unit 1110 (step S851), it creates a terminology explanation in response to the additional instruction (step S861) and then sends the created terminology explanation to the control unit 1110 (step S871).
[0364] When the control unit 1110 receives a response, i.e., a glossary, sent from the large-scale language model (step S881), it outputs the glossary along with the answer sentence that has already been created to the user as processing for the response from the large-scale language model (step S892).
[0365] The processing flow in Fig. 23 will be further described with reference to the explanatory diagram in Fig. 24. As shown in Fig. 24, for example, suppose that a user inputs a question such as "What kind of surgery is Tommy John surgery?", and the large-scale language model creates an answer to this question, saying, "Tommy John surgery is a surgical procedure in which elbow ligaments damaged mainly by pitching or the like are removed and healthy tendons extracted from the patient's palmaris longus muscle are transplanted into the affected area to reconstruct the ligaments," and transmits the created answer to the control unit 1110 (step S870).
[0366] When the control unit 1110 receives this answer sentence (step S880), it compares the answer sentence with the check DB 19010, for example. If the "palmis longus muscle" included in the answer sentence is registered as a technical term in the check DB 19010, the control unit 1110 determines that an explanation of the technical term "palmis longus muscle" is necessary (step S920), generates an additional instruction sentence such as "Please explain the term palmaris longus muscle" (step S832), and sends the generated additional instruction sentence to the large-scale language model (step S841).
[0367] When the large-scale language model receives the additional instruction sentence (step S851), it creates a term explanation in response to the additional instruction sentence, such as "The palmaris longus is a muscle in the human upper limb that flexes the wrist joint and tensions the palmar aponeurosis" (step S861). The created term explanation is then transmitted (step S871). At this time, a response sentence may also be transmitted, if necessary.
[0368] When the control unit 1110 receives the term explanation (step S881), it processes the response (answer to the question / with technical term explanation) of the large-scale language model by displaying on the display unit 10011 the user's question, "What kind of surgery is Tommy John surgery?", the large-scale language model's answer, "Tommy John surgery is a surgical procedure in which elbow ligaments damaged mainly by pitching or other actions are removed and healthy tendons extracted from the patient's palmaris longus muscle are transplanted into the affected area to reconstruct the ligaments," and the term explanation from the large-scale language model, "The palmaris longus muscle is a muscle in the human upper limb that flexes the wrist and tensions the palmar aponeurosis." (step S892).
[0369] In addition, if no technical term registered in the checking DB 19010 is found in the answer sentence in step S920 (step S920: No), the control unit 1110 processes the response of the large-scale language model (answer to the question / no technical term explanation) by displaying the user's question sentence and the answer sentence of the large-scale language model on the display unit 10011 (step S891).
[0370] This response output process flow can reduce the user's effort when asking a question to a large-scale language model. For example, if an answer sentence from a large-scale language model contains technical terms that the user does not know, the user may need to ask the large-scale language model a further question to inquire about the meaning of the technical terms. However, by appropriately adding an explanation of the technical terms to the answer sentence, the user can avoid the effort of asking a further question.
[0371] However, even if an answer contains technical terms, it is not necessary to add an explanation of the technical terms. For example, as shown in FIG. 25 , whether or not to provide explanations of technical terms depending on the situation may be preset. In this example, even if an answer contains technical terms, if the technical terms in the answer are also included in the question, the explanation of the technical terms is set to "unnecessary." In other words, even if an answer contains technical terms, if the technical terms in the answer are also included in the user's question, it is determined that the user knows the technical terms, and no explanation of the term is added. Therefore, in the setting example shown in FIG. 25 , a term explanation is added to the answer only if the answer contains technical terms but the technical terms in the answer are not included in the user's question. Note that the setting of whether or not to provide explanations of the term may be changed by the user at will.
[0372] Next, the response output processing flow of the response output system of Example 7 will be further described with reference to the processing flow of Fig. 26. Fig. 26 is a modified example of the processing flow shown in Fig. 11 etc., and is an example of a check process in which it is determined whether or not the question sentence contains an "abbreviation" as a specific expression.
[0373] 26 , upon receiving a user input (step S810), the control unit 1110 performs a database matching process for the question in step S8211. In this matching process, the question entered by the user is matched with specific expressions registered in the check DB 19010. In this example, the matching process involves matching the question with "abbreviations" registered in the check DB 19010.
[0374] Next, in step S8212, it is determined whether the question contains an abbreviation registered in the check DB 19010. If the question does not contain an abbreviation as a specific expression (step S8212: No), a response instruction sentence is generated in step S830, and the generated response instruction sentence is sent to the large-scale language model (step S840).
[0375] On the other hand, if the question contains an abbreviation as a specific expression (step S8212: Yes), the control unit 1110 proceeds to step S8250, where it performs a conversion process for the question. In this example, the control unit 1110 performs a process of querying the user for details about the abbreviation included in the question. In other words, the control unit 1110 performs a process of requesting the user to input an explanation of the abbreviation included in the question or a phrase for the abbreviation in full (for example, the abbreviation ETC is written as Electronic Tool Collection without abbreviation). Then, upon receiving user input in response to this query, i.e., detailed information about the abbreviation (step S8260), the control unit 1110 generates a response instruction sentence based on the user's question and the detailed information about the abbreviation (step S830) and sends the created response instruction sentence to the large-scale language model (step S840).
[0376] This process flow makes it easier to obtain a more appropriate answer from a large-scale language model, even when a user's question contains an abbreviation. This is particularly effective when the question contains an abbreviation with multiple meanings, or when the abbreviation is familiar to the user but not widely recognized.
[0377] In the processing flow of Fig. 26, steps S8211, S8212, and S8250 correspond to the question sentence checking process and conversion process. Furthermore, the processing flow after the large-scale language model receives the response instruction sentence in step S850 is the same as the processing flow of Fig. 11, etc., and therefore will not be illustrated or described in Fig. 26.
[0378] Next, the response output processing flow of the response output system of Example 7 will be further described with reference to the processing flow of Fig. 27. Fig. 27 is a modified example of the processing flow shown in Fig. 11 etc., and is an example of a check process in which it is determined whether or not the question sentence contains a "dialect" as a specific expression.
[0379] 27, upon receiving a user input (step S810), the control unit 1110 performs a database matching process for the question in step S8211. In this matching process, the question entered by the user is matched with a specific expression registered in the check DB 19010. In this example, the matching process involves matching the question with the "dialects" registered in the check DB 19010.
[0380] Next, in step S8212, it is determined whether the question sentence includes a dialect as a specific expression. If the question sentence does not include a dialect as a specific expression (step S8212: No), the control unit 1110 generates a response instruction sentence in step S830 and sends the generated response instruction sentence to the large-scale language model (step S840).
[0381] On the other hand, if the question contains a dialect (step S8212: Yes), the process proceeds to step S8270, where the control unit 1110 executes a conversion process for the question. In this example, the control unit 1110 identifies a dialect region likely to be used by the user based on pre-registered user information, a conversation between the user and the large-scale language model prior to the question, and performs a process for converting the dialect of the question into standard Japanese based on the identified dialect region. Then, the control unit 1110 generates a response instruction sentence based on the converted question (converted question) (step S830) and transmits the generated response instruction sentence to the large-scale language model (step S840). In other words, regardless of whether the user has used a dialect in the question, the control unit 1110 generates a response instruction sentence based on the question written in standard Japanese and transmits it to the large-scale language model.
[0382] As in the case of the processing flow shown in FIG. 11 etc., when the large-scale language model receives a response instruction sentence (step S850), it generates a response (answer sentence) according to the received response instruction sentence (step S860) and transmits the generated response (answer sentence) to the control unit 1110 (step S870).
[0383] When the control unit 111 receives the response sent from the large-scale language model (step S880), it next determines whether or not to use dialect in the user output (step S930).
[0384] In this example, if the question sentence entered by the user uses a dialect, it is determined that the dialect will be used in the answer sentence to be output to the user (step S930: Yes). In other words, if the response instruction sentence is based on the converted question sentence, it is determined that the dialect will be used in the user output. In this case, the process proceeds to step S940, where the control unit 1110 executes a conversion process for the answer sentence. Specifically, a process is performed to convert the standard language of the answer sentence into a dialect corresponding to the identified dialect region. Thereafter, the control unit 1110 outputs an answer sentence using the dialect (converted answer sentence) to the user as a process for responding to the response of the large-scale language model (step S892).
[0385] On the other hand, if the question sentence input by the user is in standard language at step S930, it is determined that standard language will be used in the answer sentence to be output to the user (step S930: No). In other words, if the response instruction sentence is based on the original question sentence input by the user, it is determined that standard language will be used in the user output. In this case, the process proceeds to step S893, where the control unit 1110 outputs an answer sentence using standard language to the user as processing for the response of the large-scale language model (step S893).
[0386] With this processing flow, even if a user's question contains a dialect, the meaning of the unique words in the dialect that differ from standard Japanese is converted so that the large-scale language model can understand the question, making it easier to obtain a more appropriate answer to the question from the large-scale language model.In addition, by converting the response of the large-scale language model into a dialect corresponding to the dialect region entered by the user as the question and presenting it to the user, the user can feel as if they are interacting with a familiar person when interacting with the large-scale language model.
[0387] The necessity of question conversion processing and the content of processing for the response of the large-scale language model are set in advance as shown in an example in Fig. 28. In the example shown in Fig. 28, the necessity of question conversion processing (identifying the dialect region and converting the dialect into standard Japanese) is set to "necessary" if the question contains a dialect, and set to "not necessary" if the question does not contain a dialect. The necessity setting may be changed by the user as appropriate.
[0388] Furthermore, for the processing of the response of the large-scale language model, processing details are set for the cases where the question contains a dialect and the case where the question does not contain a dialect. In this example, as the processing details for the case where "the question contains a dialect," one of "1. Answer using dialect (dialect of the same region as the question)," "2. Answer in standard Japanese (without using dialect)," "3. Answer by listing both 1 and 2," and "4. User selection of 1-3" is preset. Similarly, as the processing details for the case where "the question does not contain a dialect," one of "1. Answer in standard Japanese (without using dialect)," "2. Answer using dialect (dialect of a region closely related to the user)," "3. Answer by listing both 1 and 2," and "4. User selection of 1-3" is preset.
[0389] When "4. User selection of 1-3" is set, the user selects any one of the above items 1 to 3 at any timing, such as when entering a question. This setting may be changed on the user setting menu screen by selecting a predetermined menu item displayed on the display unit 10011 while using the system, for example.
[0390] Next, the processing flow of Fig. 27 will be further described with reference to the explanatory diagram of Fig. 29. As shown in Fig. 29, for example, when a question such as "Tell me why this summer has become so hot" is input by voice by the user (step S810), the control unit 1110 checks the question, which has been converted into text, against the check DB 19010 (step S8211), and determines whether the question contains a dialect (step S8212).
[0391] If it is determined that the question contains a dialect (step S8212: Yes), the control unit 1110 identifies a dialect region and converts the dialect in the question to standard Japanese based on the identified dialect region (step S8270). In the example of FIG. 29, the presence or absence of a dialect in the question is determined to be "yes," and the dialect region is identified as "Hokkaido." The dialect expression "namara" in the question is then converted to standard Japanese "temo" (very much), and the dialect expression "re" is converted to standard Japanese imperative form. As a result, the question is converted to "Please tell me why this summer has become so hot."
[0392] Thereafter, the control unit 1110 generates a response instruction sentence based on the converted question sentence (converted question sentence) as described above (step S830), and sends the generated instruction sentence to the large-scale language model (step S840).
[0393] Furthermore, when the large-scale language model receives a response instruction sentence (step S850), it creates a response sentence in standard Japanese in response to the received response instruction sentence, such as "Global warming is progressing and the average temperature is continuing to rise. In Japan alone, the reason for this summer's extreme heat is that the meandering westerly winds have caused consecutive days of hot air covering the area," (step S860).The created response sentence is then sent to the control unit 1110 (step S870).
[0394] When the control unit 111 receives the answer sentence transmitted from the large-scale language model (step S880), it determines whether or not to use a dialect in the user output (step S930). If it determines that a dialect will be used in the answer sentence to be output to the user (step S930: Yes), the control unit 1110 converts the answer sentence to create a converted answer sentence (an answer sentence using a dialect) such as, for example, "Global warming is progressing and the average temperature is continuing to rise. Therefore, in Japan only, the reason this summer has been so hot is because the meandering westerly winds have caused the country to be covered in hot air for several days" (step S940), and outputs the converted answer sentence to the user (step S892).
[0395] On the other hand, if it is determined in step S930 that a dialect will not be used in the answer to be output to the user (step S930: No), the process for processing the response of the large-scale language model is to output a standard answer to the user using standard Japanese: "Global warming is progressing on a global scale, and average temperatures are continuing to rise. In Japan alone, the reason for this summer's extreme heat is that the meandering westerly winds have meant that the country has been covered in hot air for several days." (step S893).
[0396] Next, the response output processing flow of the response output system of Example 7 will be further described with reference to the processing flow of Fig. 30. Fig. 30 shows a modified example of the processing flow shown in Fig. 27, specifically, a modified example of the question sentence conversion processing when a dialect is included in the question sentence.
[0397] The example shown in Fig. 30 is a case where a user inputs a question by voice using a microphone 1139 or the like. As shown in Fig. 30, when the control unit 1110 receives a user's question (user input) input by voice (step S811), it then performs a database matching process on the question text (step S8211). In this matching process, as in the example of Fig. 27, the question text is matched with the "dialects" registered in the check DB 19010. The control unit 1110 then determines whether the question text includes a dialect as a specific expression (step S8212).
[0398] Here, if the question sentence does not include a dialect as a specific expression (step S8212: No), as in the example of Figure 27, the control unit 1110 generates a response instruction sentence in step S830 and sends the generated response instruction sentence to the large-scale language model (step S840).
[0399] On the other hand, if the question contains a dialect (step S8212: Yes), the process proceeds to step S8280, where control unit 1110 performs a process of converting the question by requesting the user to input the voice-input question into text by operating operation input unit 1107. For example, information requesting the user to re-input the question by operating operation input unit 1107 is displayed on display unit 1011.
[0400] In response to this request, the user inputs a question in text, and when the control unit 1110 receives the input question text (step S950), the control unit 1110 generates a response instruction text based on the input question text (converted question text) (step S830), and sends the generated response instruction text to the large-scale language model (step S840).
[0401] In speech input, even if a question contains a dialect, it is expected that the question will be entered in standard Japanese when the user enters the same question in text. In other words, by changing the way the user enters the question, it is expected that a question using a dialect will be converted into a question using standard Japanese. Therefore, the processing flow of this example also makes it easier to obtain an appropriate answer using a large-scale language model for a user's question that contains a dialect.
[0402] In the processing flow of Fig. 30, steps S8211, S8212, and S8280 correspond to the question sentence checking process and conversion process. Also, the processing flow after the large-scale language model receives the response instruction sentence in step S850 is the same as the processing flow of Fig. 27, and therefore will not be illustrated or described in Fig. 30.
[0403] As described above, the region where a user speaks a dialect is identified, for example, from pre-registered user information. For this reason, it is preferable that the user information include information suitable for identifying the dialect region. FIG. 31 is a diagram illustrating an example of a user information registration screen. In the example screen illustrated in FIG. 31 , the following user information is registered: "the user's birthplace," "information about areas where the user has lived," and "the birthplaces of people with whom the user frequently converses." The reason for registering this information is that dialects may vary slightly from city, town, or village to city, even within the same region, and may change over time. Furthermore, a user's language may be influenced by the places where the user has spent time and the words of people with whom the user frequently interacts. Registering such information, such as birthplace, as user information makes it easier to identify the dialect region where the user speaks based on the user information.
[0404] Next, the response output processing flow of the response output system of Example 7 will be further described with reference to the processing flow of Figure 32. Figure 32 is a modified example of the processing flow shown in Figure 11 etc., and is an example of a check process in which it is determined whether a specific expression, "homonyms / homographs," is included in a question sentence. Note that the expression "homonyms / homographs" here means "homonyms and / or homographs." Furthermore, "including homonyms / homographs" means "including specific words and phrases that include homonyms / homographs."
[0405] 32 , upon receiving a user input (step S810), the control unit 1110 performs a database matching process for the question in step S8211. In this matching process, the question input by the user is matched with specific expressions registered in the check DB 19010. In this example, the matching process involves matching the question with "homonyms / homophones" registered in the check DB 19010.
[0406] Next, in step S8212, it is determined whether the question contains a homonym / homograph as a specific expression registered in the check DB 19010. In other words, in step S8212, it is determined whether the question contains a homonym / homograph with an unclear meaning. If the question does not contain a homonym / homograph with an unclear meaning (step S8212: No), in step S830, a response instruction sentence is generated based on the original question entered by the user, and the generated response instruction sentence is sent to the large-scale language model (step S840).
[0407] On the other hand, if the question contains a homonym / homograph as a specific expression and its meaning is unclear (step S8212: Yes), the process proceeds to step S8290, where the control unit 1110 performs a conversion process for the question. In other words, the control unit 1110 performs a process for clarifying the homonym / homograph whose meaning is unclear. In this example, the control unit 1110 performs a process for requesting or suggesting the user to change the question input method. That is, the control unit 1110 performs a process for requesting or suggesting the user to ask the question again using a second input method that is different from the first input method.
[0408] When the control unit 1110 receives a user re-input in response to this request, i.e., a re-question (converted question sentence) with a changed input method (step S950), the control unit 1110 returns to step S8212 and determines whether the converted question sentence contains homonyms / homographs with unclear meanings. If the control unit 1110 determines that the converted question sentence does not contain homonyms / homographs with unclear meanings (step S8212: No), the control unit 1110 proceeds to step S830, generates a response instruction sentence based on the converted question sentence, and sends the generated response instruction sentence to the large-scale language model (step S840).
[0409] In the processing flow of Fig. 32, steps S8211, S8212, and S8290 correspond to the question sentence checking process and conversion process. Also, the processing flow after the large-scale language model receives the response instruction sentence in step S850 is the same as the processing flow of Fig. 11, etc., and therefore will not be illustrated or described in Fig. 32.
[0410] Furthermore, whether or not homonym / homograph clarification processing is required, i.e., whether or not a re-question with a changed input method is required, and the content of a proposal or request to change the input method to the user, are set in advance, as shown in an example in Fig. 33. In the example shown in Fig. 33, the homonym / homograph clarification processing is set to "required" when "the question contains a homonym and it is unclear what meaning the word or phrase is used in," and "not required" when "the question contains a homonym but the meaning of the word or phrase is clear, or it does not contain a homonym / homograph." Note that this setting of whether or not homonym / homograph clarification is required may be changed by the user as appropriate.
[0411] Furthermore, suggestions to the user for changing the input method are only made when "homonyms are included in the question and it is unclear what meaning the phrase is being used in," and when homonym / homograph disambiguation processing is set to "required." In this example, the suggestion "when a voice-input question contains a homonym" is to "switch the question input method from voice input to text input," and the suggestion "when a text-input question contains a homonym" is to "switch the question input method from text input to voice input."
[0412] Next, the processing flow shown in Fig. 32 will be further described with reference to Fig. 34A to Fig. 34C. Each of Fig. 34A to Fig. 34C is a diagram for explaining a specific example of the processing flow shown in Fig. 32.
[0413] For example, as shown in FIG. 34A , assume that a question such as "Where is the largest bank in my area?" is received from a user (step S810). Furthermore, assume that in step S8212, the word "bank" in the question is registered as a "homonym" and a "homograph" in the check DB 19010, and it is determined that the meaning of the question is unclear (step S8212: Yes). As an example, the meaning of "bank" in this question could be "bank" or "slope" such as a bank. In this case, as shown in FIG. 33 , the homonym / homograph clarification process satisfies the condition "necessary." Therefore, the control unit 1110 performs a process to suggest to the user that they change the question input method (step S8290). This allows the meaning of the homonym / homograph contained in the question to be clarified.
[0414] As an example of a suggestion process when voice input is used as the question input method, the control unit 1110 displays the message "The question contains words that can have multiple meanings. Entering a question by text may solve the problem." An example of a user's response to this suggestion is to switch the question input method from voice input to text input. This response is considered to be particularly effective when the question contains homonyms.
[0415] Furthermore, as an example of a suggestion process when text input is used as the question input method, the control unit 1110 displays the message, "The question contains words that can have multiple meanings. Asking by voice may solve your problem." An example of a user's response to this suggestion is to switch the question input method from text input to voice input. This response is considered to be particularly effective when the question contains homographs.
[0416] Furthermore, as an example of a suggestion process common to when voice input is used as a question input method and when text input is used, the control unit 1110 displays, "The question contains words that can have multiple meanings. Please revise your question to make the meaning clearer." An example of a user's response to this suggestion is to revise the question and change it to, for example, "What is the most well-funded bank in my area?"
[0417] For example, the words "Flower" and "Flower" are pronounced the same but have different meanings, and depending on the context of the co...
Claims
1. An input unit for a user to input a question sentence, a control unit that generates a response instruction sentence for a large language model based on the question sentence and obtains a response sentence generated by the large language model for the response instruction sentence, and an output unit that performs an output based on the response sentence obtained by the control unit. The control unit executes a conversion process of the question sentence and generates the response instruction sentence based on the question sentence after the conversion process. A response output system.
2. In the response output system according to claim 1, the control unit performs a check process to check whether the question sentence corresponds to a predetermined check item, and when the question sentence corresponds to the check item, executes the conversion process of the question sentence. A response output system.
3. The response output system according to claim 2, wherein the check item includes that the question sentence includes a preset specific expression, and the control unit performs the conversion process of the question sentence when the specific expression is included in the question sentence in the check process. A response output system.
4. The response output system according to claim 3, wherein the specific expression is classified into a plurality of groups, and the control unit checks whether a specific expression classified into a preset group among the plurality of groups is included in the question sentence as the check process, and when the specific expression classified into the preset group is included in the question sentence, performs the conversion process of the question sentence. A response output system.
5. The response output system according to claim 4, wherein a priority is preset for each of the preset groups, and the control unit checks whether the specific expression classified into each preset group is included in the question sentence in the order according to the priority. A response output system.
6. The response output system according to claim 1, wherein as the conversion process of the question sentence, the control unit generates a conversion instruction sentence for requesting the conversion of the question sentence, and performs a process of obtaining, as the question sentence after the conversion process, the question sentence converted by the large language model according to the conversion instruction sentence. A response output system.
7. The response output system according to claim 1, wherein the control unit, as the conversion process of the question text, generates a summary instruction text that requests a summary of the question text, and obtains, as the question text after the conversion process, a summary of the question text created by the large language model according to the summary instruction text.
8. The response output system according to claim 1, wherein the control unit, as the conversion process of the question text, performs a process of requesting the user to re-create the question text.
9. The response output system according to claim 3, comprising: an audio input unit that receives the question text as audio input; and an operation input unit that receives the question text as character input. When the question text input by the user from the audio input unit includes a dialect as the specific expression, the control unit, as the conversion process of the question text, performs a process of requesting the user to input the question text from the operation input unit.
10. The response output system according to claim 1, wherein after executing the conversion process of the question text, the control unit asks the user whether the question text after the conversion process matches the user's intention. When an answer is received from the user that the question text after the conversion process matches the user's intention, the control unit generates the response instruction text based on the question text after the conversion process.
11. The response output system according to claim 19, wherein when an answer is received from the user that the question text after the conversion process does not match the user's intention for the inquiry, the control unit further performs a process of requesting the user to re-create the question text.
12. The response output system according to claim 1, wherein the control unit, as the conversion process of the question text, performs a first translation process of translating the question text input by the user in a first language into a second language different from the first language, and a second translation process of translating the question text after the first translation process from the second language into the first language.
13. The response output system according to claim 12, wherein the control unit makes an inquiry as to whether the question sentence after the second translation process matches the intention of the user, and when there is an answer from the user that the question sentence after the conversion process matches the intention of the user, generates the response instruction sentence based on the question sentence after the second translation process. Response output system.
14. The response output system according to claim 13, wherein the control unit, when there is an answer to the inquiry that the question sentence after the second translation process does not match the intention of the user, further performs a process of instructing the user to request a re-creation of the question sentence. Response output system.
Citation Information
Patent Citations
Artificial intelligence-based human-machine interaction method and device
JP2019528512A
Systems, methods, and media for retrieving an entity from a data table using semantic search
US20230259507A1