Response output device

The response output device enhances user interaction by processing user inputs through a control unit to generate and output responses using a large language model, addressing the lack of user-centric configurations in existing AI technologies.

JP2025109546APending Publication Date: 2025-07-25MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024003501
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing response output technologies using artificial intelligence, such as language models, lack sufficient consideration for user-centric configurations.

Method used

A response output device comprising an input unit, a control unit, and an output unit, which processes user inputs to generate and output responses through a large language model, including a conversion process to enhance user interaction.

Benefits of technology

Provides a more suitable and user-friendly response output technology by leveraging a control unit to generate and output responses effectively, addressing the limitations of existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109546000001_ABST
    Figure 2025109546000001_ABST
Patent Text Reader

Abstract

To provide a further suitable artificial intelligence response output technology, and to contribute to "9 Industry, Innovation and Infrastructure" and "11 Sustainable Cities and Communities" of the Sustainable Development Goals (SDGs).SOLUTION: A response output system includes: an input unit to which a question sentence is input by a user; a control unit that generates a response instruction sentence to a large-scale language model based on the question sentence, and acquires a response sentence generated by the large-scale language model in response to the response instruction sentence; and an output unit that produces an output based on the response sentence acquired by the control unit. The control unit performs conversion processing on the question sentence and generates the response instruction sentence on the basis of the question sentence after the conversion processing.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a response output device.

Background Art

[0002] Regarding response output technology using artificial intelligence such as a language model, for example, it is disclosed in Patent Document 1.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the disclosure of Patent Document 1, considerations regarding configurations for more suitably providing response output technology using artificial intelligence to users were not sufficient.

[0005] An object of the present invention is to provide a more suitable response output technology.

Means for Solving the Problems

[0006] In order to solve the above problems, for example, the configuration described in the claims is adopted. This application includes a plurality of means for solving the above problems. If an example is given, a response output device including an input unit into which a question sentence is input by a user, a control unit that generates a response instruction sentence for a large language model based on the question sentence and acquires a response sentence generated by the large language model for the response instruction sentence, and an output unit that performs an output based on the response sentence acquired by the control unit, wherein the control unit executes a conversion process of the question sentence and generates the response instruction sentence based on the question sentence after the conversion process may be configured.

Effects of the Invention

[0007] According to the present invention, a more suitable response output technology can be provided. Other problems, configurations, and effects will be clarified in the following description of the embodiments.

Brief Description of the Drawings

[0008]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 2F

Figure 2G

Figure 2H

Figure 2I

Figure 2J

Figure 2K

Figure 2L

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 3F

Figure 3G

Figure 3H

Figure 3I

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17

Figure 18A

Figure 18B

Figure 19A

Figure 19B

Figure 20

Figure 21

Figure 22A

Figure 22B

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34A

Figure 34B

Figure 34C

Figure 35A

Figure 35B

Figure 36A

Figure 36B

Figure 37

Figure 38

Mode for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. It should be noted that the present invention is not limited to the description of the embodiments, and various changes and modifications can be made by those skilled in the art within the scope of the technical idea disclosed in this specification. Also, in all the drawings for explaining the present invention, those having the same function are denoted by the same reference numerals, and the repeated description thereof may be omitted.

[0010] Note that when the artificial intelligence response output device according to each embodiment of the present invention has a display screen, it may be referred to as a display device. When the artificial intelligence response output device has a voice output function, it may be referred to as a voice output device. The artificial intelligence response output device may simply be referred to as an information processing device. A system including the artificial intelligence response output device and a large language model server that holds a large language model may be referred to as an artificial intelligence response output system. Also, when the artificial intelligence response output device provides a response service of a large language model that is artificial intelligence to a user and is helpful to the user, the artificial intelligence response output device or the display output of the artificial intelligence response output device can be an artificial intelligence (AI) assistant for the user. Therefore, in this case, the artificial intelligence response output device may be referred to as an AI assistant device or an AI assistant display device. Similarly, in this case, a system including the artificial intelligence response output device and a large language model server that holds a large language model may be referred to as an AI assistant system or an AI assistant display system. Also, in this case, since the artificial intelligence response output device serves as an interface between the user and the artificial intelligence, it may be referred to as an artificial intelligence interface device. In this case, a system including the artificial intelligence response output device and a large language model server that holds a large language model may be referred to as an artificial intelligence interface system.

[0011] <Example 1> As Example 1 of the present invention, an artificial intelligence response output device that outputs a response from a large language model artificial intelligence and its system will be described.

[0012] Using FIG. 1A, an example of the artificial intelligence response output device 10010 of the present invention will be described. Also, regarding the case where the artificial intelligence response output device 10010 is linked to the large language model server 19001 through communication or the like, an example of a system including the artificial intelligence response output device 10010 and / or the multimodal large language model server 20001 will be described.

[0013] In the example of FIG. 1A, the artificial intelligence response output device 10010 has a display unit 10011. In the example of FIG. 1A, the display unit 10011 may be a flat panel display, a screen that projects an image from the back, or an airborne floating image that forms an optical image in the air. When the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Also, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which pixels emit light spontaneously. Further, the display unit 10011 may be provided with a touch operation input sensor and configured as a touch panel.

[0014] In the example of FIG. 1A, the voice output unit 1140 included in the artificial intelligence response output device 10010 is composed of a speaker. Also, the artificial intelligence response output device 10010 is provided with a microphone 1139 and can pick up the user's voice. Through the voice input from the microphone 1139 and the user's operation input via the operation input unit described later, the artificial intelligence response output device 10010 can obtain user input that serves as the basis for an instruction sentence (prompt) to the large language model that is artificial intelligence.

[0015] The artificial intelligence response output device 10010 may be provided with a local large language model in the artificial intelligence response output device 10010 itself. In this case, the response of the large language model may be output as a display output on the display unit 10011 and / or a voice output from the voice output unit 1140.

[0016] Alternatively, the artificial intelligence response output device 10010 may not have a local large language model, communicate with an external large language model server 19001, and output the response received from the large language model server 19001 as a display output on the display unit 10011 and / or an audio output on the audio output unit 1140.

[0017] Or, the artificial intelligence response output device 10010 may also have a local large language model and be further configured to communicate with an external large language model server 19001 having a large language model or an external large language model server 20001 having a multimodal large language model. In this case, the response of the local large language model and the response received from the large language model of the large language model server 19001 or the multimodal large language model of the multimodal large language model server 20001 may be switched, and either one may be output as a display output on the display unit 10011 and / or an audio output on the audio output unit 1140. Or, a response generated based on both the response of the local large language model and the response received from the large language model of the large language model server 19001 or the multimodal large language model of the multimodal large language model server 20001 may be output as a display output on the display unit 10011 and / or an audio output on the audio output unit 1140.

[0018] The configuration when the artificial intelligence response output device 10010 communicates and cooperates with an external large language model server 19001 or large language model server 20001 is as follows. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 1A, the communication between the communication unit 1132 and the communication device 19011 shows a wireless example, but wired communication is also acceptable. In the communication path from the communication unit 1132 to the communication device 19011, there may be a wired part and a wireless part, or it may pass through a router or a repeater. Also, in the communication path from the communication unit 1132 to the Internet 19000, there may be a wired part and a wireless part, or it may pass through a router or a repeater. The artificial intelligence response output device 10010 can communicate with the large language model server 19001 via the communication device 19011 and the Internet 19000. Also, the artificial intelligence response output device 10010 can communicate with the large language model server 19001 or large language model server 20001, and a second server 19002 different from these servers via the communication device 19011 and the Internet 19000. The configuration including the artificial intelligence response output device 10010 and the large language model server 19001 or large language model server 20001 may be considered as one system.

[0019] In the following description, when simply referring to the "large language model" without special notice, it may be considered as a concept including the local large language model provided in the artificial intelligence response output device 10010, the large language model provided in the large language model server 19001, and the multimodal large language model provided in the large language model server 20001.

[0020] In the example of FIG. 1A, the display unit 10011 shows an example where each element is displayed in two display areas: an instruction display area 10051 for inputting an instruction sentence (prompt) from the user to a large language model which is an artificial intelligence, and an artificial intelligence response display area 10061 for displaying the response from the large language model. In the example of FIG. 1A, in the instruction display area 10051, examples of displaying an icon 10052 indicating the user, text 10053 such as natural language or software code as components of the instruction sentence, an image 10054 as a component of the instruction sentence, a video 10055 as a component of the instruction sentence, etc. are shown. In the example of FIG. 1A, in the artificial intelligence response display area 10061, examples of displaying an icon 10062 indicating the artificial intelligence or artificial intelligence assistant, text 10063 such as natural language or software code as components of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, a video 10065 as a component of the response from the artificial intelligence, etc. are shown. Note that the display example of the display unit 10011 of the artificial intelligence response output device 10010 shown in FIG. 1A is merely an example. Depending on the implementation example in which the artificial intelligence response output device 10010 is used, a display different from the example shown in FIG. 1A may be performed.

[0021] Here, the large language model will be described. The large language model is also denoted as LLM (Large Language Model). Specifically, various models such as GPT-1, GPT-2, GPT-3, InstructGPT, ChatGPT, etc. have been publicly released. These techniques may also be used in this embodiment. Note that these large language models are artificial intelligence models generated by performing large-scale pre-training on natural languages contained in a large number of documents and texts existing in the human world. The number of parameters of the artificial intelligence model exceeds hundreds of millions. Furthermore, in addition to this, there are also models that have undergone reinforcement learning based on feedback from humans. An example of the base model is a model called Transformer. As an example of the learning of these models, for example, Reference 1 etc. have been publicly released.

[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf

[0023] These large language models can perform tasks such as translation of natural language, grammar correction of natural language, summarization of natural language, etc. Among them, more advanced models can perform tasks such as question answering in natural language (also called dialogue or conversation), generation of proposals in natural language, generation of programming code, etc. Since the number of parameters of these artificial intelligence models is extremely large, a huge amount of data and computing resources are required for training. Therefore, it is very resource-inefficient to perform this level of artificial intelligence training only for specific applications. Thus, a foundation model that can be applied to various applications has been generated through large-scale pre-training. For example, the large language model server 19001 shown in Figure 1A is equipped with such a large language model and may be configured to be available on various terminals via an API (Application Programming Interface). Also, the artificial intelligence response output device 10010 shown in Figure 1A may be equipped with a local large language model and configured for use by the artificial intelligence response output device 10010 itself. The learning of any large language model itself can be separately generated through large-scale pre-training, and the generated large language model can be replicated and installed in the large language model server 19001, the artificial intelligence response output device 10010, etc. In this way, instead of performing pre-training for each application or terminal, by replicating the large language model, which is a foundation model generated through large-scale pre-training, and using it on individual servers or terminals, the resource consumption for learning can be shared, resulting in good resource efficiency.

[0024] Note that even for a large language model as a foundation model generated through large-scale pre-training, additional learning such as transfer learning may be configured to be performed in individual servers or devices according to the application and purpose.

[0025] In addition, large language models can pre-learn natural language and perform input / output processing for natural language. Furthermore, multi-modal large language model artificial intelligence that can process not only text information of natural language but also other types of information in addition to the text information of natural language is also applicable to the embodiments of the present invention. In FIG. 1A, a server having a multi-modal large language model is shown as the large language model server 20001. For example, as an example of multi-modal large language model artificial intelligence, specifically, GPT-4 (see Reference 2), Gato (see Reference 3), etc. have been published. These technologies may also be used in this embodiment. Note that these multi-modal large language models are artificial intelligence models generated by performing large-scale pre-learning on a large number of documents and texts existing in the human world, including natural language and other types of information (such as images, videos, voices, etc.) other than the text information of natural language. Furthermore, in addition to this, there are also models subjected to reinforcement learning based on feedback from humans. Hereinafter, other types of information other than the text information of natural language such as images, videos, and voices may be referred to as non-natural language information sources.

[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf

[0027] Next, with reference to FIG. 1B, a configuration example of the artificial intelligence response output device 10010 that receives an input from a user for artificial intelligence such as these large language models and outputs a response from the artificial intelligence such as the large language model for the input from the user will be described.

[0028] The artificial intelligence response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, an attitude sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, etc. The artificial intelligence response output device 10010 may have, for example, a large screen such as a so-called monitor or a TV.

[0029] The display unit 10011 may be a flat panel display, a screen that projects an image from the back, or one that displays an aerial floating image that forms an optical image in the air. When the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Also, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which pixels emit light spontaneously. When the display unit 10011 is a panel, it may be referred to as a display panel. A touch operation input sensor may be provided in the display unit 10011 so as to receive touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. Through the user's operation input via the touch panel, the artificial intelligence response output device 10010 can obtain user input that serves as the basis for an instruction sentence (prompt) to a large language model that is artificial intelligence.

[0030] The communication unit 1132 may be composed of a communication interface using the Wi-Fi method, a communication interface using the Bluetooth (registered trademark) method, a mobile communication interface such as 4G or 5G, and the like. Using these communication methods, the communication unit 1132 of the artificial intelligence response output device 10010 can communicate with the communication device 19011 connected to the Internet 19000. Note that in the communication path from the communication unit 1132 to the communication device 19011, there may be a wired part and a wireless part, or it may pass through a router or a repeater. In the case of wired communication, the communication unit 1132 may have an Ethernet connection interface as hardware and perform communication using the LAN communication method. Thereby, the artificial intelligence response output device 10010 can communicate with various servers connected to the Internet 19000.

[0031] The artificial intelligence response output device 10010 is provided with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011, the communication unit 1132, and the like.

[0032] The power supply 1106 converts the AC current input from the outside through the external power supply input interface 1111 into a DC current and supplies the necessary DC current to each part of the artificial intelligence response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. Further, when no power is supplied from the outside through the external power supply input interface 1111, the secondary battery 1112 supplies power to each part that requires power.

[0033] The operation input unit 1107 is, for example, an operation button, a signal reception unit such as a remote controller, or an infrared light receiving unit, and inputs signals for operations different from the touch operations by the user on the touch operation input sensor of the display unit 10011. Separate from the user who touches the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the artificial intelligence response output device 10010. Through the user's operation input via the operation input unit 1107, the artificial intelligence response output device 10010 can obtain user input that serves as the basis for an instruction sentence (prompt) to the large language model, which is artificial intelligence. Note that there may also be a modified example in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107.

[0034] The video signal input unit 1131 connects to an external video output device and inputs video data. Various digital video input interfaces are conceivable for the video signal input unit 1131. For example, it may be configured with a video input interface of the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface of the DVI (Digital Visual Interface) standard, or a video input interface of the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be various USB interfaces, etc.

[0035] The audio signal input unit 1133 connects to an external audio output device and inputs audio data. The audio signal input unit 1133 may be configured with an audio input interface of the HDMI standard, an optical digital terminal interface, or a coaxial digital terminal interface, etc. The audio signal input unit 1133 may also be various USB interfaces, etc. In the case of an interface of the HDMI standard, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an integrated interface with integrated terminals and cables.

[0036] The voice output unit 1140 can output voice based on the voice data input to the voice signal input unit 1133. The voice output unit 1140 can also output voice based on the voice data stored in the storage unit 1170. The voice output unit 1140 may be composed of a speaker. Also, the voice output unit 1140 may output built-in operation sounds or error warning sounds. Or, a configuration that outputs a voice signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard, may be used as the voice output unit 1140. Or, a configuration that outputs a voice signal as an analog signal to an external device such as headphones may be used as the voice output unit 1140.

[0037] The microphone 1039 is a microphone that picks up sounds around the artificial intelligence response output device 10010, converts them into signals, and generates voice signals. The microphone may record the voice of a person such as the user's voice, and the generated voice signal may be subjected to voice recognition processing by the control unit 1110 described later to obtain character information from the voice signal. Through the voice input from the microphone 1139, the artificial intelligence response output device 10010 can obtain user input that serves as the basis for an instruction sentence (prompt) to a large language model, which is an artificial intelligence.

[0038] The imaging unit 1180 is a camera having an image sensor. A camera may be provided on the front side of the display unit 10011 side of the artificial intelligence response output device 10010, or a camera may be provided on the back side of the display unit 10011 side. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.

[0039] The storage unit 1170 is a storage device that records various information such as various data like video data, image data, audio data, etc. The storage unit 1170 may be composed of a magnetic recording medium recording device such as a hard disk drive (HDD), or a semiconductor element memory such as a solid state drive (SSD). In the storage unit 1170, for example, various information such as various data like video data, image data, audio data, etc. may be recorded in advance at the time of product shipment. Also, the storage unit 1170 may record various information such as various data like video data, image data, audio data, etc. acquired from an external device or an external server via the communication unit 1132. The video data, image data, etc. recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, etc. recorded in the storage unit 1170 may be output to an external device or an external server via the communication unit 1132.

[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may be referred to as a video processing circuit and may be composed of hardware such as an ASIC, an FPGA, a video processor, etc. Note that the video control unit 1160 may also be referred to as a video processing unit or an image processing unit. The video control unit 1160 performs controls such as video switching control, for example, which video signal among the video signals stored in the memory 1109 and the video signals (video data) etc. input to the video signal input unit 1131 is to be input to the display unit 10011. Also, the video control unit 1160 may perform control to perform image processing on the video signal input from the video signal input unit 1131 or the video signal stored in the memory 1109. Examples of image processing include scaling processing such as enlarging, reducing, or deforming an image, brightness adjustment processing for changing brightness, contrast adjustment processing for changing the contrast curve of an image, and retinex processing for decomposing an image into light components and changing the weighting for each component.

[0041] The posture sensor 1113 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the posture of the artificial intelligence response output device 10010. Based on the posture detection result of the posture sensor 1113, the control unit 1110 may control the operations of the connected components.

[0042] The non-volatile memory 1108 stores various data used in the artificial intelligence response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, various operation data to be displayed on the display unit 10011 of the artificial intelligence response output device 10010, display icons, data of objects to be operated by the user's operations, layout information, and the like. The memory 1109 stores video data to be displayed on the display unit 10011, device control data, and the like. The control unit 1110 may read various software from the storage unit 1170, expand it, and store it in the memory 1109.

[0043] The local LLM processing unit 10028 is equipped with a memory capable of holding a large language model (LLM) and can execute inference of the large language model based on the control of the control unit 1110. It may be composed of, for example, a so-called GPU (Graphics Processing Unit) as hardware. The local LLM processing unit 10028 may perform not only inference but also learning. Note that the local LLM processing unit 10028 is not necessarily required when the execution of inference of the large language model in the local environment of the artificial intelligence response output device 10010 is unnecessary.

[0044] The control unit 1110 controls the operations of each connected unit. Further, the control unit 1110 may perform arithmetic processing based on the information acquired from each unit within the artificial intelligence response output device 10010 in cooperation with the program stored in the memory 1109. The control states by the control unit 1110 include, for example, a state of outputting a response from the large language model of the local LLM processing unit 10028, or a response from the large language model of the large language model server 19001 or the multimodal large language model of the multimodal large language model server 20001 acquired via the communication unit 1132, via the display unit 10011 or the audio output unit 1140 such as a speaker.

[0045] In addition, when there is an input from the user via the above-described touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and it is transmitted to the local large language model of the local LLM processing unit 10028 included in the artificial intelligence response output device 10010, the large language model included in the large language model server 19001, or the multimodal large language model included in the multimodal large language model server 20001, and the control for acquiring a response from these large language models may all be performed by the control unit 1110.

[0046] In addition, the storage unit 1170 may store a response fixed sentence database (which may also be referred to as a response fixed sentence DB) for outputting a fixed sentence as a response to the instruction sentence of the artificial intelligence response output device 10010. The control unit 1110 may perform control to generate a response output using the data stored in the response fixed sentence database. An example of the response fixed sentence database is shown in FIG. 1C. In the example of FIG. 1C, for each condition with a condition number, the fixed sentence response output by the artificial intelligence response output device 10010 is stored. For example, when "Good morning" is input from the user via the touch panel, the microphone 1139, or the operation input unit 1107 as described above under condition number 1, the response may be output using "Good morning" or "Today is the [month] [day], isn't it?" as the response fixed sentence. The parts such as "[month]" and "[day]" may be generated using the information stored in the memory 1109 or the like that the artificial intelligence response output device 10010 has.

[0047] Also, in the example of the response fixed sentence in the database shown in FIG. 1C, when a plurality of response fixed sentences separated by / are stored, the control unit 1110 may perform control to randomly select any one of the response fixed sentences using a random number or the like so that the response is output. By doing so, the situation where the response under the same condition becomes monotonous can be eliminated and improved. The examples of condition numbers 2, 3, and 4 are the same as the description of the example of condition number 1. The control unit 1110 may perform control so that an output is made using the response fixed sentence of each example shown in FIG. 1C for the condition content of each example shown in FIG. 1C.

[0048] Next, an example of condition number 5 shown in FIG. 1C will be described. When the control unit 1110 cannot understand the meaning of the user input obtained via the touch panel, the microphone 1139, or the operation input unit 1107 as natural language, or when there is an obvious grammar error in the user input, condition number 5 is an example in which the control unit 1110 outputs a response using "I couldn't quite catch that" or "I might not understand that" as a response fixed sentence. By responding in this way, the user can be prompted to input again, and the corrected user input can be awaited.

[0049] Next, an example of condition number 6 shown in FIG. 1C will be described. Condition number 6 is an example in the state where the control unit 1110 has detected that any part of each part constituting the artificial intelligence response output device 10010 shown in FIG. 1B is in an error (abnormal state), and there is a user input via the touch panel, the microphone 1139, or the operation input unit 1107. In this case, the control unit 1110 performs control to output a response using "It seems something is wrong" as a response fixed sentence. By responding in this way, it is possible to explain to the user that the artificial intelligence response output device 10010 is malfunctioning, and the user can be prompted to take error countermeasures and the like.

[0050] Instead of the responses of large language models such as the local large language model provided by the artificial intelligence response output device 10010, the large language model provided by the large language model server 19001, and the multimodal large language model provided by the large language model server 20001, the artificial intelligence response output device 10010 may output a response using the response fixed sentence database (response fixed sentence DB) described with reference to FIG. 1C. Alternatively, a response combining the responses of these large language models and the response using the response fixed sentence database (response fixed sentence DB) may be output.

[0051] Incidentally, the response template sentence database (response template sentence DB) of FIG. 1C described above is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 may use this. However, the response template sentence database (response template sentence DB) shown in FIG. 1C may be provided on the side of the large language model server 19001 or the side of the large language model server 20001. In this case, the control unit included in the large language model server 19001 or the control unit included in the large language model server 20001 may generate a response using the response template sentence database (response template sentence DB). Instead of the response generated by the large language model stored in each server, the control unit included in the large language model server 19001 or the control unit included in the large language model server 20001 may transmit the response generated using the response template sentence database (response template sentence DB) to the artificial intelligence response output device 10010. In this way, even when the artificial intelligence response output device 10010 is not provided with the response template sentence database (response template sentence DB), it is possible to generate a response using the response template sentence database (response template sentence DB).

[0052] Incidentally, in the above description, the artificial intelligence response output device 10010 was described as having a display panel of a display screen using fixed pixels. This concept may include a projection type video display device (projector) that provides a projection optical system after the display panel of the display screen using fixed pixels and projects an optical image of the video of the display panel of the display screen onto a screen or a wall.

[0053] Incidentally, in the examples of FIGS. 1A and 1B, an example in which the artificial intelligence response output device 10010 includes the display unit 10011 was described. However, the artificial intelligence response output device 10010 according to the embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if it does not include the display unit 10011, it may be configured to receive an input from a user to the artificial intelligence via the voice signal input unit 1133 or the microphone 1139 and output a response from the artificial intelligence such as a large language model to the input from the user via the voice output unit 1140.

[0054] According to the artificial intelligence response output device and the artificial intelligence response output system according to Embodiment 1 of the present invention described above, it is possible to receive an input from a user to an artificial intelligence such as a large language model, and output a response to the input from the user, which is generated by inference of an artificial intelligence such as a large language model possessed by a server device on a network or a local large language model possessed by the artificial intelligence response output device itself.

[0055] <Example 2> Next, as Example 2 of the present invention, an example will be described in which the artificial intelligence response output device 10010 described in Example 1 is connected to the Internet and operates by connecting to a server equipped with a large language model artificial intelligence via the Internet. In this embodiment, the differences from Example 1 will be described, and repeated descriptions of the same configurations as these embodiments will be omitted.

[0056] An example of the connection state between the artificial intelligence response output device 10010 and the large language model server 19001 according to Example 2 of the present invention will be described with reference to FIG. 2A. The artificial intelligence response output device 10010 according to Example 2 may be called a character conversation device. Also, a system including the artificial intelligence response output device 10010 and the large language model server 19001 according to Example 2 may be called a character conversation system. A video of the character 19051 is displayed on the display unit 10011 displayed by the artificial intelligence response output device 10010. The video of the character 19051 is a video generated by rendering a 3D model of a character in a virtual space.

[0057] In addition, the character in this embodiment can provide users with the service of a large language model, which is an artificial intelligence, and can assist users. Therefore, the character can be an artificial intelligence (AI) assistant for users. In this case, the character conversation device and the character conversation system in this embodiment may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, and an AI assistant response output system.

[0058] In the example of FIG. 2A, the voice output unit 1140 included in the artificial intelligence response output device 10010 is composed of a speaker. In addition, the artificial intelligence response output device 10010 is provided with a microphone 1139 and can pick up the user's voice. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example of FIG. 2A, the communication between the communication unit 1132 and the communication device 19011 shows a wireless example, but wired communication is also acceptable. In the communication path from the communication unit 1132 to the Internet 19000, there may be a wired part and a wireless part. The artificial intelligence response output device 10010 can communicate with a large language model server 19001 via the communication device 19011 and the Internet 19000. In addition, the artificial intelligence response output device 10010 can communicate with a second server 19002 different from the large language model server 19001 via the communication device 19011 and the Internet 19000. The configuration including the artificial intelligence response output device 10010 and the large language model server 19001 may be considered as one system.

[0059] Next, with reference to FIG. 2B, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Note that in FIG. 2B, illustration of communication paths such as the Internet 19000 shown in FIG. 2A is omitted. In FIG. 2B, the user 230 of the artificial intelligence response output device 10010 is also shown.

[0060] Here, a series of processes of the operation of the artificial intelligence response output device 10010 will be described. The artificial intelligence response output device 10010 expands the character operation program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character operation program, whereby various processes described below can be realized.

[0061] First, the artificial intelligence response output device 10010 is equipped with a microphone 1139. When the user 230 speaks to the character 19051, the voice of the user (words from the user) is picked up by the microphone 1139 and converted into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by the user 230 from the audio signal. The text is in natural language. Note that the extraction of the text of the words spoken by the user 230 may be continuously performed for all words, or may be started when words are uttered by the user within a predetermined period after a keyword serving as a trigger. For example, the keyword serving as a trigger may be a case where a character name is uttered after "Hello" from the user. For example, if the name of the character 19051 is "Koto", "Hello, Koto!" may be used as the keyword serving as a trigger.

[0062] The character operation program of the artificial intelligence response output device 10010 creates an instruction text (prompt) based on the text of the words spoken by the user 230, and sends the instruction text to the large language model server 19001 using an API. Here, the instruction text may be metadata storing information described by a notation using tags such as a markup format of a markup language, a notation using predetermined symbols such as a Markdown format, or an object notation of a predetermined script such as JSON. The instruction text stores text information in natural language as the main message. As types of instruction texts sent from the artificial intelligence response output device 10010 to the large language model server 19001, there are setting instruction texts storing instructions such as initial settings, and user instruction texts reflecting instructions from the user. Type identification information for identifying whether the instruction text is a setting instruction text or a user instruction text may be stored in a part other than the main message of the instruction text. When the character operation program of the artificial intelligence response output device 10010 creates an instruction text (prompt) based on the text of the words spoken by the user 230, it creates a user instruction text and sends it to the large language model server 19001.

[0063] Next, the large language model of the large language model server 19001 executes inferences based on the instruction text transmitted from the AI response output device 10010, and based on the results, generates a response including natural language text information. The large language model server 19001 uses the API to transmit the response to the AI response output device 10010. The response stores natural language text information as the main message. Here, the response may be metadata storing information described by a notation in the same format as the aforementioned instruction text (such as a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as the Markdown format, or an object notation of a predetermined script such as JSON). In the response, when using the same format as the aforementioned instruction text, type identification information may be stored in a part other than the main message to indicate that the initial setting instruction text and the user instruction text are different types of information. For example, information indicating that it is a response sentence from the large language model is stored.

[0064] Next, the AI response output device 10010 receives the response from the large language model server 19001, and extracts the natural language text information stored as the main message in the response. The character operation program of the AI response output device 10010 generates natural language voice as an answer to the user using voice synthesis technology based on the natural language text information extracted from the aforementioned response, and outputs it from the voice output unit 1140, which is a speaker, so that it sounds like the voice of the character 19051. This process may be expressed as the "utterance" of the character.

[0065] Specific examples of the response voice of the character 19051 to the words from the user 230 are shown in Conversation Examples 1 to 5 in FIG. 2C by the processing of the AI response output device 10010 and the large language model server 19001 described above. In this way, the user 230 can have a conversation as if the character 19051 were a real person.

[0066] According to the artificial intelligence response output device 10010 in FIG. 2B or the system including the artificial intelligence response output device 10010 described above, it is not necessary to install the large language model itself, which requires a huge amount of data for learning and computing resources, in the artificial intelligence response output device 10010 itself. Moreover, through the API, the advanced natural language processing capabilities of the large language model can be utilized, enabling a more suitable response to the user and a more suitable conversation when the user talks to the character.

[0067] Next, with reference to FIG. 2D, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, FIG. 2D shows an example of the natural language text of the main message of the instruction text sent from the artificial intelligence response output device 10010 to the large language model server 19001, which is the basis for the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the natural language text of the main message of the server response as the response.

[0068] Also, in FIG. 2D, it is shown that the exchange of instructions and responses is made in time series from the display setting instruction text, the first round of user instruction text and its response to the fourth round of user instruction text and its response.

[0069] As shown in FIG. 2D, through a setting instruction, the large language model of the artificial intelligence in the large language model server 19001 can be initially set to indicate the name of the large language model itself, the role to play, the characteristics of the conversation, etc. Also, the name of the user can be understood as an initial setting. As a result, the large language model generates responses after the first round while adhering to the role. Then, from the perspective of the user who hears the voice of the character 19051 based on the responses after the first round, it feels as if the character 19051 has the settings and personality of the person described in the setting instruction. In addition, the large language model server 19001 according to this embodiment is equipped with a memory for storing the content of the conversation until a series of conversations ends. After storing a series of user instructions and their responses, it is configured to generate responses. Thereby, a conversation as shown in FIG. 2D can be realized.

[0070] Next, with reference to FIG. 2E, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, FIG. 2E is an example of the natural language text of the main message of the instruction sent from the artificial intelligence response output device 10010 to the large language model server 19001, which is the basis for the conversation between the character 19051 and the user 230 displayed on the artificial intelligence response output device 10010, and the natural language text of the main message of the server response as the response.

[0071] FIG. 2E shows an example of the case where, after the series of conversations shown in FIG. 2D ends and the continuation of the series of conversations ends, the user 230 talks to the character 19051 again to start a new conversation. In FIG. 2E, it is shown that the exchange of instructions and responses is made in chronological order from the first round of user instructions and their responses to the third round of user instructions and their responses.

[0072] Here, the "end" of the "continuation of a series of conversations" means that when a predetermined condition is met, the large language model server 19001 deletes the conversation memory that it held while the series of conversations was ongoing, from the large language model server 19001. An example of a predetermined condition is, for example, when the artificial intelligence response output device 10010 instructs the large language model server 19001 to "end" the "continuation of a series of conversations" by an instruction statement. Another example of a predetermined condition is, for example, when the transmission of instruction statements from the artificial intelligence response output device 10010 to the large language model server 19001 regarding the series of conversations has stopped and a predetermined amount of time has passed (timeout). Also, in the connection between the artificial intelligence response output device 10010 and the large language model server 19001, when authentication processing has been performed and the above-mentioned instruction statements and responses are being exchanged, a case where the authentication processing is lost due to factors such as a communication disconnect or the power-off (OFF) of the artificial intelligence response output device 10010 is also included.

[0073] Note that when the "continuation of a series of conversations" "ends", the large language model server 19001 deletes the conversation memory that it held while the series of conversations was ongoing, from the large language model server 19001. Therefore, although the conversation shown in Figure 2E is after the series of conversations shown in Figure 2D, the server response to the user instruction statement completely does not remember the name as a character set in the large language model, the role to play, the characteristics of the conversation, the name of the user, etc. included in the setting instruction statement shown in Figure 2D, and is a response with the content in a state where nothing is remembered. Similarly, the conversation shown in Figure 2E is a response with the content in a state where there is no memory of the series of conversations shown in Figure 2D at all. That is, due to the "end" of the "continuation of a series of conversations" shown in Figure 2D, the conversation in Figure 2E starts from a state where the large language model of the artificial intelligence of the large language model server 19001 has been initialized.

[0074] This becomes a factor that makes it seem to user 230 as if character 19051 has lost its memories with the user or feels like a completely different person. From the perspective of user 230, the responses of this character feel very incongruous and result in an experience of loneliness and disappointment. In such operations, there was an issue that the identity of settings and memories such as the name, role, or conversation characteristics and personality of character 19051 displayed on the artificial intelligence response output device 10010 could not be ensured.

[0075] Next, with reference to FIG. 2F, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, FIG. 2F shows an example of the natural language text of the main message of the instruction text transmitted from the artificial intelligence response output device 10010 to the large language model server 19001, which is the basis for the conversation between character 19051 displayed on the artificial intelligence response output device 10010 and user 230, and the natural language text of the main message of the server response as the response.

[0076] Figure 2F shows an example of a case where, after the continuation of the series of conversations shown in Figure 2D ends, user 230 speaks to character 19051 again to start a new conversation. Different from the process of Figure 2E, in the process of Figure 2F, when starting a new conversation, the artificial intelligence response output device 10010 sends a setting instruction as the first instruction to the large language model server 19001. The setting instruction stores the same natural language text as the setting instruction of the initial setting in Figure 2D. This may be expressed as a re - setting text. Subsequently, the setting instruction stores natural language text explaining the history of past conversations. This may be expressed as conversation history text. The history of past conversations may be recorded by the artificial intelligence response output device 10010 in the storage unit 1170 as natural language text information associated with the date and time of the conversation during the continuation of the series of conversations described in Figure 2D. If there are conversations on different dates, they may be recorded associated with the date and time information for each conversation, and the conversation history may be accumulated. When generating the setting instruction of the first instruction for a later conversation as shown in Figure 2F, the natural language text information of the conversation recorded in the storage unit 1170 and the date and time information of the conversation may be read out and used for generating the setting instruction.

[0077] In addition, when using the natural language text information of the history of past conversations to generate the setting instruction, since it is data to be sent to the large language model, the format can be determined somewhat freely without problems. However, as shown in Figure 2F, prefixes and suffixes in natural language such as "I had the following conversation on month day." and "You had the following conversation on month day." are prepared and merged with the natural language text information of the recorded conversation to generate the text of the setting instruction. Also, the date and time information of the conversation read from the storage unit 1170 may be merged with parts such as the "month day" part above and used as part of the text of the setting instruction.

[0078] Even if, after a series of conversations, the continuation of the series of conversations ends and the user 230 addresses the character 19051 again to start a new conversation, if the generation process and transmission process of the setting instruction text of FIG. 2F described above are performed, the response to the subsequent user instruction text will reflect the role, name, conversation characteristics, personality, and / or conversation characteristics of the character at the time of the previous conversation, as well as the conversation history. As a result, from the user's perspective, it is recognized that the identity of the settings and memories such as the role, name, conversation characteristics, or personality of the character at the time of the previous conversation is more securely maintained, which is more preferable.

[0079] Next, with reference to FIG. 2G, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, FIG. 2G shows an example of the natural language text of the main message of the instruction text transmitted from the artificial intelligence response output device 10010 to the large language model server 19001, which is the basis of the conversation between the character 19051 and the user 230 displayed on the artificial intelligence response output device 10010, and the natural language text of the main message of the server response that is the response.

[0080] FIG. 2G shows an example of a series of conversations from the first round of user instruction text to the third round of user instruction text and their responses, following the first setting instruction text in the series of conversations shown in FIG. 2F. In FIG. 2G, the exchanges of instructions and responses are shown as being made in chronological order. Since the content of the setting instruction text is as shown in FIG. 2F, repeated description will be omitted.

[0081] As shown in the natural language text of the server response in the table of FIG. 2F, by using the setting instruction text shown in FIG. 2F, the server response by the large language model artificial intelligence of the large language model server 19001 reflects the settings and conversation history such as the role, name, conversation characteristics, or personality of the character at the time of the previous conversation. As a result, from the user's perspective, it is recognized that the identity of the settings and memory such as the role, name, conversation characteristics, or personality of the character at the time of the previous conversation can be more securely ensured, so it is more suitable. Note that since this means that the character can be recognized as the same from the user's perspective, it may be referred to as the pseudo-identity of the character as seen by the user.

[0082] Also, from the user's perspective, they can share memories with the character and obtain a more enjoyable character conversation experience.

[0083] Next, with reference to FIG. 2H, an example of the operation of the character conversation device (an example of the operation of the artificial intelligence response output device 10010) according to Example 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, FIG. 2H shows an example of an operation of switching and displaying a character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 from among a plurality of character candidates. The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 may switch the displayed character based on, for example, an operation input entered into the operation input unit 1107 or an operation detected by the touch operation input sensor of the display unit 10011.

[0084] In the example of FIG. 2H, in addition to the character 19051 (name: "Koto") used in the description of FIGS. 2A to 2G, the character 19052 (name: "Tom") and the character 19053 (name: "Necco") are shown. The character 19051 (name: "Koto") and the character 19052 (name: "Tom") are human characters, and the character 19053 (name: "Necco") is a cat character. To switch the display of the characters to be displayed on the display unit 10011, it is sufficient to switch and display on the display unit 10011 the video generated by rendering the characters in different virtual 3D spaces for each character. For the process of realizing the display of the rendering video of the 3D model of each character, for example, any of the first to third processing examples described in FIG. 15A may be performed. Also, for some characters, dynamic 2D images may be displayed.

[0085] Also, when the character operation program executed by the control unit 1110 switches the display of the characters to be displayed on the display unit 10011, it is preferable to also change the synthesized voice used for the "utterance" of each character. For this, in advance, the data of the synthesized voice of the voice color associated with each character is stored in the storage unit 1170, and the synthesized voice change process may also be performed when switching the display of the characters.

[0086] Note that in the example of FIG. 2H, it is configured so that the user 230 can talk to any of the characters. In the artificial intelligence response output device 10010 of FIG. 2H, different roles, names, conversation characteristics, or personalities are set for each of these characters. Also, the memory of each character based on the conversation history is managed as different for each character.

[0087] Therefore, the artificial intelligence response output device 10010 constructs the database shown in FIG. 2I in the storage unit 1170, and manages the character settings and the character conversation history by the database.

[0088] Next, with reference to FIG. 2I, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, FIG. 2I is an explanatory diagram of a database 19200 for managing character settings and character conversation histories for a plurality of characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.

[0089] A character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 constructs the database 19200 in, for example, the storage unit 1170. The character ID is an identification number for identifying each of a plurality of characters that can be displayed on the artificial intelligence response output device 10010, and may be a natural number or may use alphabets or the like. The name is data of the name of each of a plurality of characters that can be displayed on the artificial intelligence response output device 10010.

[0090] The initial setting instruction text is text information in natural language that describes settings such as the role, name, conversation characteristics, or personality of each of a plurality of characters that can be displayed on the artificial intelligence response output device 10010. Since the initial setting instruction text becomes the main data of the setting instruction text transmitted from the artificial intelligence response output device 10010 to the large language model server 19001 and is text information in natural language, it is desirable that the description content can be read by the artificial intelligence large language model of the large language model server 19001 as it is.

[0091] The conversation histories 1, 2, … that continue are records of conversations between each character and the user, and are recorded separately for each character. Since the conversation history will be included in the text information of natural language, which is the main data of the setting instruction text sent from the artificial intelligence response output device 10010 to the large language model server 19001, it is desirable that the description content can be directly read by the large language model of the artificial intelligence of the large language model server 19001.

[0092] When the character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 switches the character displayed on the display unit 10011 of the artificial intelligence response output device 10010, the initial setting instruction text and conversation history used for the text information of natural language, which is the main data of the setting instruction text sent from the artificial intelligence response output device 10010 to the large language model server 19001, are selected and switched to correspond to the character displayed on the display unit 10011 of the artificial intelligence response output device 10010 by using the database 19200 in FIG. 2I. In addition, every time a conversation between the user 230 and the character occurs, the character operation program records the conversation history in the area of the conversation history corresponding to the character displayed on the display unit 10011 in the database 19200 in FIG. 2I.

[0093] Even though the character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 uses the database 19200 in this way to establish a conversation between the user 230 and the character using the utterances of the character that utilize the responses of the same large language model of the same large language model server 19001, from the user's perspective, the uniqueness of the settings such as the personality of each character is maintained, and it feels like the conversation memories for each character continue separately for each character. From the user's perspective, in each character, it is recognized that the identity of the settings and memories such as the role, name, conversation characteristics, or personality of the character at the time of the previous conversation is more securely ensured, so it is more suitable. This can also be expressed as being able to ensure the pseudo-identity of the character as seen by the user for each character.

[0094] Therefore, even when the artificial intelligence response output device 10010 is configured to switch and display the character to be displayed on the display unit 10011 from among a plurality of character candidates, according to the operation using the database 19200 described above, from the user's perspective, there is less discomfort felt from the conversation with each character, and it is possible to share the memory with each of the plurality of characters, obtaining a more enjoyable character conversation experience.

[0095] In addition, if the user cannot edit the initial setting instructions for a plurality of characters, the settings such as the role, name, conversation characteristics, or personality of each character can be maintained in a state closer to the intention of the provider of the artificial intelligence response output device 10010 or the producer of the character content. On the other hand, the user may be able to edit the initial setting instructions of the character according to the input from the operation input unit 1107 or the like. In this case, the settings such as the role, name, conversation characteristics, or personality of the character can be set to the preferred settings, and it becomes possible to have a conversation with the character uniquely set by the user. In this case, the 3D model of the character, its rendering video, and the type of synthesized voice of the character may be replaced accordingly.

[0096] Next, with reference to FIG. 2J, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 2 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 19001. Specifically, a method for providing a character conversation service by a character conversation device using the artificial intelligence response output device 10010 or a character conversation system using the artificial intelligence response output device 10010 and the large language model server 19001 at a lower cost will be described.

[0097] As described with reference to FIG. 2B, it is very resource-inefficient to perform this level of artificial intelligence learning on a large language model limited to a specific application. Therefore, as a foundation model applicable to various applications, it is resource-efficient to generate a model through large-scale learning and use it on various terminals via an API (Application Programming Interface). Then, the provider of the large language model often recovers the cost used for learning the large language model from the terminal user as the usage fee of the terminal API. At that time, in the natural language model, the usage fee of the API is often billed in a form based on the number of processed units of words (called tokens) that divide a sentence.

[0098] Therefore, also in the artificial intelligence response output device 10010 according to Embodiment 2 of the present invention, by reducing the number of tokens in the text information of natural language transmitted between the artificial intelligence response output device 10010 and the large language model server 19001 using an API, a character conversation service by a character conversation device using the artificial intelligence response output device 10010 or a character conversation system using the artificial intelligence response output device 10010 and the large language model server 19001 can be provided to the user at a lower cost.

[0099] For example, by adopting the processing and configuration of Examples 1 to 3 as shown in the table of FIG. 2J, it is possible to technically reduce the number of tokens in the natural language text information transmitted between the artificial intelligence response output device 10010 and the large language model server 19001 using an API.

[0100] Example 1 is an example of reducing the number of tokens in the conversation history text stored and transmitted in the API setting instruction text by shortening the conversation history text using document summarization processing. For example, the natural language of the conversation history with the character recorded in the storage unit 1170 is summarized and recorded. The document summarization may be performed at the start of the next conversation, but it is better to perform it at the end of "a series of conversations" as there is more time available.

[0101] Also, for the document summarization processing, a request for summarization may be made to the large language model itself of the large language model server 19001. However, in this case, the effect of saving the number of tokens is low. Therefore, for example, when the second server 19002 provides natural language document summarization processing via an API at a lower cost than the large language model of the large language model server 19001, a request for document summarization processing may be made to the second server 19002 via the API, and the document summary of the conversation history may be stored in the setting instruction text to the large language model server 19001 and transmitted.

[0102] Also, if it is only the document summarization processing, it can also be performed on the terminal side. The control unit 1110 may execute a document summarization program deployed in the memory 1109 of the artificial intelligence response output device 10010 to perform document summarization. In this case, the effect of saving the number of tokens is high. Also, even if the conversation history becomes long, if the upper limit of the number of characters after summarization is specified in the document summarization processing, the upper limit of the length of the conversation history text is determined, so the upper limit value of the tokens can be determined and the tokens can be saved.

[0103] Note that the text information of the initial character settings, such as the role, name, conversation characteristics, or personality of the character, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the description of the text information of the initial character setting instruction text and reduce the number of tokens of the text information of the conversation history.

[0104] The process described in Example 1 may be performed by the character operation program executed by the control unit 1110 to control each part.

[0105] Example 2 is another example of reducing the number of tokens of the conversation history text stored in and transmitted by the API setting instruction text. For example, among the conversation history with the character recorded in the storage unit 1170, the older history is deleted to reduce the number of tokens. If the upper limit of the number of characters of the conversation history is specified, the upper limit of the length of the conversation history text is determined, so the upper limit value of the tokens can be determined and token savings can be achieved. Or, it may be a method of specifying a predetermined period of the conversation history and deleting the conversation history exceeding the period. Also in this case, token savings can be achieved. Note that also in Example 2, the text information of the initial character settings, such as the role, name, conversation characteristics, or personality of the character, does not increase as much as the conversation history. Therefore, it is efficient and preferable to maintain the description of the text information of the initial character setting instruction text and reduce the number of tokens of the text information of the conversation history.

[0106] The process described in Example 2 may be performed by the character operation program executed by the control unit 1110 to control each part.

[0107] Example 3 is a method of reducing the number of tokens by decreasing the transmission frequency of setting instruction texts using an API. Specifically, after the device power is turned on or after the display character is switched, and even after the video settings and synthesized voice settings of the displayed character are completed, without pre-transmitting the setting instruction text, when the control unit 1110 determines that the text information of the natural language included in the user's utterance picked up by the microphone 1139 is the text information to be used with the large language model of artificial intelligence, the setting instruction text is transmitted to the large language model server 19001 for the first time, thereby reducing the transmission frequency of the setting instruction text to the large language model server 19001 and reducing the number of tokens.

[0108] Specifically, for example, after the device is powered on (ON) or after the operation input for switching the display character, by the display process of the display unit 10011 under the control of the character operation program executed by the control unit 1110, the character 19051 (name: "Koto") is displayed on the display unit 10011 as shown in FIG. 2H. At this time, for example, if the synthesized voice for the appearance corresponding to the character 19051 is stored and prepared in the storage unit 1170 or the like, the synthesized voice for the appearance such as "Good morning. I'm Koto.", "Hello. I'm Koto.", "Good evening. I'm Koto." may be output from the speaker which is the voice output unit 1140. At this time, the video of the character 19051 has already been set as the video of the character displayed on the display unit 10011, and the synthesized voice output from the speaker which is the voice output unit 1140 is set as the synthesized voice corresponding to the character 19051.

[0109] Here, the inference process of the large language model of artificial intelligence in the large language model server 19001 described above also takes time as the instruction text becomes longer. In particular, when the setting instruction includes text information regarding the past conversation history, the number of tokens in the instruction text increases, and thus the inference process time becomes particularly long. The setting instruction itself and its response are not output to the user 230. From the response to the user instruction after the setting instruction, a synthesized voice as the "utterance" of the character is output from the speaker which is the voice output unit 1140. Then, it may seem preferable to transmit the setting instruction from the artificial intelligence response output device 10010 to the large language model server 19001 in advance and complete the inference process of the large language model for the setting instruction in advance, because the response of the synthesized voice output of the "utterance" of the character 19051 after the user 230 talks to the character 19051 becomes faster.

[0110] However, if, before the user 230 utters words, the setting instruction is transmitted to the large language model server 19001 and the inference process of the large language model for the setting instruction is completed in advance, for example, the user 230 turns off the power of the artificial intelligence response output device 10010 through an operation via the touch operation input sensor of the operation input unit 1107 or the display unit 10011, or for example, the user 230 switches the display character from the character 19051 to another character through an operation via the touch operation input sensor of the operation input unit 1107 or the display unit 10011. In this case, the number of tokens processed by transmitting the setting instruction to the large language model server 19001 in advance and processed by the inference process of the large language model becomes the number of processed tokens that waste the usage fee uselessly. This hinders the provision of the character conversation service by the character conversation device by the artificial intelligence response output device 10010 or the character conversation system by the artificial intelligence response output device 10010 and the large language model server 19001 to the user at a lower cost.

[0111] Therefore, after the artificial intelligence response output device 10010 is powered on or after an operation input for switching display characters, under the control of the character operation program executed by the control unit 1110, the video of character 19051 is set as the video of the character displayed on the display unit 10011, and the synthesized voice output from the speaker, which is the voice output unit 1140, is set as the synthesized voice corresponding to character 19051. However, it is desirable to continue the state of not sending the setting instruction text to the large language model server 19001 until the point in time when the user 230 is recognized as speaking to character 19051.

[0112] Here, the point in time when the user 230 is recognized as speaking to character 19051 may be, for example, until the point in time when the keyword serving as a trigger described in FIG. 2B is detected, or until the point in time when the text of the words spoken by the user 230 is extracted. By doing so, the number of processing tokens that waste usage fees can be reduced, and the character conversation service provided by the character conversation device using the artificial intelligence response output device 10010 or the character conversation system using the artificial intelligence response output device 10010 and the large language model server 19001 can be provided to the user at a lower cost.

[0113] Also, even after exceeding the point in time when the above-mentioned user 230 is recognized as speaking to the character 19051, for example, when the text information extracted from the voice of the user 230 recorded by the microphone 1139 corresponds to the text information of preset keywords that do not require inference processing of the large language model, it is desirable to continue the state of not sending the setting instruction text to the large language model server 19001. Specifically, as examples of preset keywords, there are cases such as "try jumping", "try dancing", etc., where the user 230 requests the character 19051 to perform reactions such as animations in which the character 19051 moves or emits synthesized voices. In this case, the character motion program executed by the control unit 1110 reads the motion data, animation video, and / or synthesized voice data corresponding to the character 19051 stored in the storage unit, and uses these data to perform the generation process of the video displayed on the display unit 10011 and the output process of the synthesized voice from the speaker which is the voice output unit 1140.

[0114] Such processing does not necessarily require the inference processing of the large language model of the large language model server 19001. After such processing, if the user 230 turns off the power of the artificial intelligence response output device 10010 through an operation via the operation input unit 1107 or the touch operation input sensor of the display unit 10011, or for example, if the user 230 switches the displayed character from the character 19051 to another character through an operation via the operation input unit 1107 or the touch operation input sensor of the display unit 10011, if the setting instruction text is sent to the large language model server 19001 first and processed by the inference processing of the large language model, the number of tokens of that processing will become the number of processing tokens that waste the usage fee uselessly.

[0115] Therefore, even after exceeding the point in time when it is recognized that the above-mentioned user 230 addresses the character 19051, for example, until it is determined whether the text information extracted from the voice of the user 230 recorded by the microphone 1139 corresponds to text information of preset keywords that do not require inference processing by the large language model, it is desirable to continue the state of not transmitting the setting instruction text to the large language model server 19001. When it is determined by this determination that inference processing by the large language model is required, it is desirable to transmit the setting instruction text to the large language model server 19001 for the first time and proceed with the inference processing of the large language model.

[0116] Note that the processing described in Example 3 may be performed by controlling each part with a character operation program executed by the control unit 1110.

[0117] According to the method for reducing (saving) the number of processing tokens of the large language model according to each example of FIG. 2J described above, the character conversation service by the character conversation device by the artificial intelligence response output device 10010 or the character conversation system by the artificial intelligence response output device 10010 and the large language model server 19001 can be provided to the user at a lower cost.

[0118] Next, an example of the display of the character conversation device (artificial intelligence response output device 10010) according to the second embodiment of the present invention will be described with reference to FIG. 2K. In the example of FIG. 2K, an example is shown in which the response from the large language model to the instruction text from the user described in each of FIGS. 2A to 2J is displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, it is an example in which the text 10063, which is the response from the large language model, is displayed on the display unit 10011 together with the video of the character 19051. The text 10063, which is the response from the large language model, may be displayed superimposed in front of the video of the character 19051 as shown in FIG. 2K. Also, the text 10063, which is the response from the large language model, may be displayed together with the video of the character 19051 without being superimposed on the video of the character 19051.

[0119] The display in FIG. 2K is an example. For example, when the user 230 adjusts the volume of the voice output of the voice output unit 1140 of the character conversation device (artificial intelligence response output device 10010) to the minimum or sets the voice output to OFF by operating the touch operation input sensor of the operation input unit 1107 or the display unit 10011, the user 230 cannot confirm the response from the large language model by voice.

[0120] Therefore, in this case, the control unit 1110 may control to start a display mode in which the text 10063, which is a response from the large language model, is displayed together with the video of the character 19051 as shown in FIG. 2K. In this way, even when the user wants to refrain from voice output, the user 230 can use the character conversation device (artificial intelligence response output device 10010) more suitably. Note that the ON / OFF of the display mode in which the text 10063, which is a response from the large language model, is displayed together with the video of the character 19051 may be configured to be manually switchable by an operation of the user 230 via the touch operation input sensor of the operation input unit 1107 or the display unit 10011.

[0121] Next, an example of a response template database (response template DB) in a character conversation device (artificial intelligence response output device 10010) that can display a plurality of characters, which was described with reference to FIGS. 2H and 2I, will be described with reference to FIG. 2L. In the example of FIG. 2L, the condition number and the condition content are the same as those in FIG. 1C. For these conditions, in the example of FIG. 2L, individual response templates are set for each of the plurality of characters. For example, response templates for each condition are stored for each of the three characters: character 1: Koto, character 2: Tom, and character 3: Necco, which were described with reference to FIGS. 2H and 2I. Since the output control of the response template is the same as that in FIG. 1C, repeated description will be omitted.

[0122] In the example of FIG. 2L, the control unit 1110 may select a corresponding response template sentence from the response template sentence database (response template sentence DB) based on the character displayed on the character conversation device (artificial intelligence response output device 10010) and the current conditions, and use it for output control as the response issued by the character. For example, in the example of the response template sentence database (response template sentence DB) in FIG. 2L, even under the same conditions, the response template sentences are respectively changed to expressions or contents corresponding to the personalities of the characters. Thereby, the character conversation device (artificial intelligence response output device 10010) can provide the user with a conversation corresponding to the personality of the displayed character. The user can feel that each character has a more consistent personality. Thereby, a character conversation device (artificial intelligence response output device 10010) that gives a more realistic feeling to a plurality of characters can be realized.

[0123] Note that the response template sentence database (response template sentence DB) of FIG. 2L described above is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 may use this. However, the response template sentence database (response template sentence DB) shown in FIG. 2L may be provided on the side of the large language model server 19001. In this case, the control unit of the large language model server 19001 may generate a response using the response template sentence database (response template sentence DB). The control unit of the large language model server 19001 may transmit the response generated using the response template sentence database (response template sentence DB) to the artificial intelligence response output device 10010 instead of the response generated by the large language model stored in each server. In this way, even when the artificial intelligence response output device 10010 is not provided with the response template sentence database (response template sentence DB), it is possible to generate a response using the response template sentence database (response template sentence DB).

[0124] According to the character conversation device and the character conversation system according to Example 2 described above, it is possible to reduce the sense of discomfort felt by the user from the conversation with the character displayed on the artificial intelligence response output device 10010. Also, according to the character conversation device and the character conversation system according to Example 2, it is possible to provide the character conversation service to the user at a lower cost.

[0125] In the above description of Example 2, an example using the large language model of the large language model server 19001 as the large language model was described. In contrast, the character conversation device (artificial intelligence response output device 10010) may be configured to include the local LLM processing unit 10028 shown in FIG. 1B, and instead of the large language model of the large language model server 19001, the large language model of the local LLM processing unit 10028 may be used. In this case, in the above description of Example 2, the large language model of the large language model server 19001 may be replaced with the large language model of the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010).

[0126] Also in this case, it is possible to reduce the sense of discomfort felt by the user from the conversation with the character displayed on the artificial intelligence response output device 10010. Note that when using the large language model of the local LLM processing unit 10028 instead of the large language model of the large language model server 19001, the need to consider the usage fee according to the number of processing tokens is reduced, but even for the large language model of the local LLM processing unit 10028, by reducing the number of processing tokens, consumption resources such as the power required for inference can be reduced. In this case, it is possible to provide the user with a character conversation service that consumes less power.

[0127] In addition, in the above description of Example 2, an example of recording and holding the conversation history with the character in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) was described. On the other hand, the conversation history with the character may be recorded and held in the second server 19002 connected to the Internet 19000 or other cloud servers. In this case, when a new conversation between the user and the character starts, the character conversation device (artificial intelligence response output device 10010) communicates with the second server 19002 or other cloud servers to obtain (download) the past conversation history between the character and the user, and holds it in the storage unit 1170 or the memory 1109 of the character conversation device (artificial intelligence response output device 10010), and it may be used for creating an instruction sentence for the large language model. Since the specific method of using the past conversation history for the instruction sentence to the large language model is as described in each figure of Example 2, repeated description is omitted.

[0128] In addition, each time a conversation between the user and the character takes place, or at a predetermined time such as when the conversation between the user and the character ends, the character conversation device (artificial intelligence response output device 10010) may transmit (upload) the character conversation history up to that point to the above-mentioned second server 19002 or other cloud servers. That is, the character conversation device (artificial intelligence response output device 10010) uploads the conversation history with the character to the second server 19002 or other cloud servers at a predetermined timing. When the user starts a conversation with the character, the character conversation device (artificial intelligence response output device 10010) may download the latest conversation history from the second server 19002 or other cloud servers and use it to generate an instruction text for the large language model. In this way, even if the character conversation device (artificial intelligence response output device 10010) used by the user the previous day and the character conversation device (artificial intelligence response output device 10010) that the user will use now are different individual devices, the same character can be displayed. When the user has multiple conversations with the same character at different times between these different individual devices, a conversation can be realized as if the memory of the character was pseudo-continuously inherited from the previous conversation, which is more suitable for the user.

[0129] The process described above, in which the character conversation device (artificial intelligence response output device 10010) uploads and downloads the conversation history with the character to the second server 19002 or other cloud servers to pseudo-transfer the memory of the character, is also effective when dealing with the database 19200 including the conversation histories of a plurality of characters described in FIGS. 2H and 2I. That is, if the database 19200 described in FIG. 2I is configured to be uploaded and downloaded to the second server 19002 or other cloud servers, not only for one character but also for a plurality of characters, when the user has multiple conversations with each character of the plurality of characters at different timings between different individual devices, a conversation can be realized as if the memory of each character was pseudo-transferred from the previous conversation, which is more suitable for the user.

[0130] <Example 3> Next, Example 3 of the present invention improves the character conversation device (artificial intelligence response output device 10010) and the character conversation system described in each figure of Example 2. In this example, the differences from Example 2 will be described, and repeated descriptions of the same configurations as these examples will be omitted.

[0131] Similar to Example 2, the character in Example 3 provides the user with the service of a large language model, which is artificial intelligence and can be helpful to the user. Therefore, the character can be an artificial intelligence (AI) assistant for the user. In this case, the character conversation device and the character conversation system in this example may also be referred to as an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, and an AI assistant response output system.

[0132] Using FIG. 3A, an example of the character conversation device and the character conversation system according to Embodiment 3 of the present invention will be described. In the character conversation system of Embodiment 3, a large language model server 20001 is provided instead of the large language model server 19001 in FIG. 2A and is connected to the Internet 19000.

[0133] Here, the large language model server 20001 is a server equipped with a large language model artificial intelligence, but in addition to the text information of natural language that could be processed by the large language model server 19001, it is a multimodal large language model artificial intelligence that can also process information of types other than text information of natural language.

[0134] Also, as an example, the artificial intelligence response output device 10010, which is a character conversation device, will be described as having the same configuration as the character conversation device (artificial intelligence response output device 10010) of Embodiment 2.

[0135] Also in Embodiment 3, the artificial intelligence response output device 10010, which is a character conversation device, can communicate with the large language model of the large language model server 20001 via the Internet 19000 using an API.

[0136] In the character conversation system of Embodiment 3, there is a mobile information processing terminal 20010 used by the user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.

[0137] Here, using FIG. 3B, an example of the mobile information processing terminal 20010 will be described. The mobile information processing terminal 20010 includes a display panel 20011 which is a touch operation input panel, a control unit 20012, an external power input interface 20013, a power source 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, an attitude sensor 20018, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, etc.

[0138] The display panel 20011 is equipped with a touch operation input sensor and can receive touch operation inputs by the finger of the user 230. The display panel 20011 performs display using a liquid crystal panel or an organic EL panel and can display images. The display panel 20011 may also be referred to as a display unit.

[0139] The communication unit 20020 may be composed of a communication interface of the Wi-Fi method, a communication interface of the Bluetooth method, a mobile communication interface such as 4G or 5G, etc. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the character conversation device (artificial intelligence response output device 10010). The mobile information processing terminal 20010 is equipped with a control unit such as a CPU and a memory, and the control unit controls the display panel 20011, the communication unit 20020, etc. Also, by any one of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with the communication device 19011 connected to the Internet 19000. Thereby, the mobile information processing terminal 20010 can communicate with various servers connected to the Internet 19000.

[0140] The power supply 20014 converts the AC current input from the outside through the external power supply input interface 20013 into a DC current and supplies the necessary DC current to each part of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. Also, when power is not supplied from the outside through the external power supply input interface 20013, the secondary battery 20015 supplies power to each part that requires power.

[0141] The video signal input unit 20023 connects to an external video output device and inputs video data. Various digital video input interfaces can be considered for the video signal input unit 20023. For example, it may be configured with a video input interface compliant with the HDMI (Registered Trademark) (High-Definition Multimedia Interface) standard, a video input interface compliant with the DVI (Digital Visual Interface) standard, or a video input interface compliant with the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be various USB interfaces and the like.

[0142] The audio signal input unit 20024 connects to an external audio output device and inputs audio data. The audio signal input unit 20024 may be configured with an audio input interface compliant with the HDMI standard, an optical digital terminal interface, or a coaxial digital terminal interface, etc. The audio signal input unit 20024 may also be various USB interfaces and the like. In the case of an interface compliant with the HDMI standard, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an integrated interface with integrated terminals and cables.

[0143] The audio output unit 20021 is capable of outputting audio based on the audio data input to the audio signal input unit 20024. The audio output unit 20021 is also capable of outputting audio based on the audio data stored in the storage unit 20016. The audio output unit 20021 may be composed of speakers. Also, the audio output unit 20021 may output built-in operation sounds or error warning sounds. Alternatively, a configuration that outputs digital signals to external devices, such as the Audio Return Channel function defined in the HDMI standard, may be used as the audio output unit 20021.

[0144] The microphone 20022 is a microphone that picks up sounds around the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may record the voices of people such as the user's voice, and the generated audio signals may be subjected to speech recognition processing by a control unit 20012 described later to obtain character information from the audio signals.

[0145] The imaging unit 20025 is a camera having an image sensor. The camera may be provided on the front surface on the display panel 20011 side of the mobile information processing terminal 20010, or may be provided on the back surface on the display panel 20011 side. Both the front camera and the back camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a back camera.

[0146] The storage unit 20016 is a storage device that records various information such as various data such as video data, image data, and audio data. The storage unit 20016 may be composed of a magnetic recording medium recording device such as a hard disk drive (HDD) or a semiconductor element memory such as a solid state drive (SSD). For example, various information such as various data such as video data, image data, and audio data may be recorded in the storage unit 20016 in advance when the product is shipped. Further, the storage unit 20016 may record various information such as various data such as video data, image data, and audio data obtained from an external device or an external server via the communication unit 20020. The video data, image data, etc. recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc. recorded in the storage unit 20016 may be output to an external device or an external server via the communication unit 20020.

[0147] The video control unit 20017 performs various controls on the video signal input to the display panel 20011. The video control unit 20017 may also be referred to as a video processing circuit and may be composed of hardware such as an ASIC, an FPGA, or a video processor. Note that the video control unit 20017 may also be referred to as a video processing unit or an image processing unit. The video control unit 20017 performs controls such as video switching, for example, which video signal among the video signals stored in the memory 20026 and the video signals (video data) input to the video signal input unit 20023 is to be input to the display panel 20011. Further, the video control unit 20017 may perform control to perform image processing on the video signal input from the video signal input unit 20023, the video signal stored in the memory 20026, etc. Examples of the image processing include scaling processing such as enlarging, reducing, and deforming an image, brightness adjustment processing for changing brightness, contrast adjustment processing for changing the contrast curve of an image, and retinex processing for decomposing an image into light components and changing the weighting for each component.

[0148] The attitude sensor 20018 is a sensor composed of a gravity sensor or an acceleration sensor, or a combination thereof, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operations of the connected components.

[0149] The non-volatile memory 20027 stores various data used in the mobile information processing terminal 20010. The data stored in the non-volatile memory 20027 includes, for example, various operation data to be displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data of objects to be operated by the user's operations, layout information, etc. The memory 20026 stores video data to be displayed on the display panel 20011, control data of the device, etc. The control unit 20012 may read various software from the storage unit 20016, expand it, and store it in the memory 20026.

[0150] The control unit 20012 controls the operations of each connected unit. Further, the control unit 20012 may perform arithmetic processing based on the information acquired from each unit within the mobile information processing terminal 20010 in cooperation with the program stored in the memory 20026.

[0151] Next, with reference to FIG. 3C, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 20001. Also in Embodiment 3, the character conversation device (artificial intelligence response output device 10010) expands the character operation program stored in the storage unit 1170 or the like into the memory 1109, and by the control unit 1110 executing the character operation program, various processes described below can be realized.

[0152] In Embodiment 2, the action performed by the user 230 on the character conversation device (artificial intelligence response output device 10010) was mainly a call by the voice of the user 230. In the character conversation device (artificial intelligence response output device 10010) of Embodiment 2, a series of operations were performed starting from the process of picking up the voice of the user 230 with the microphone. On the other hand, in the character conversation device (artificial intelligence response output device 10010) of Embodiment 3, it is assumed that a series of operations performed by the character conversation device (artificial intelligence response output device 10010) starting from the process of picking up the voice of the user 230 with the microphone, which was described in Embodiment 2, can also be executed. In addition to this, in the character conversation device (artificial intelligence response output device 10010) of Embodiment 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by a user operation via the operation input unit 1107 in FIG. 1B. Here, examples of the operation input unit 1107 in FIG. 1B include a mouse, a keyboard, a touch panel, and the like.

[0153] Also, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the user 230 can perform an action on the character conversation device (artificial intelligence response output device 10010) by a touch operation of the user that can be detected by the touch operation input sensor of the display unit 10011 in FIG. 1B.

[0154] Also, the user 230 can input an operation input to the character conversation device (artificial intelligence response output device 10010) by operating the mobile information processing terminal 20010 and communicating from the mobile information processing terminal 20010 to the character conversation device (artificial intelligence response output device 10010).

[0155] Also, an information storage image such as a two-dimensional code storing information that the user wants to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the imaging unit 1180 in FIG. 1B of the character conversation device (artificial intelligence response output device 10010) may capture the display. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may extract information from the information storage image such as the two-dimensional code captured by the imaging unit 1180 and obtain the information. Also, an image that the user wants to transmit to the character conversation device (artificial intelligence response output device 10010) may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and the imaging unit 1180 in FIG. 1B of the character conversation device (artificial intelligence response output device 10010) may capture the display. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) may perform image recognition processing on the image captured by the imaging unit 1180 and obtain the result of the image recognition processing.

[0156] As described above, in the character conversation device (artificial intelligence response output device 10010) of Example 3, the types of actions that can be performed by the user 230 on the character conversation device (artificial intelligence response output device 10010) are more numerous than those in Example 2. As a result, the character conversation device (artificial intelligence response output device 10010) of Example 3 can obtain the results of actions performed by the user 230 other than the user's voice, and based on this, generate an instruction text (prompt) to be sent to the large language model server 20001. Thereby, it is possible to more suitably include information of types other than the text information of natural language extracted from the user's voice in the instruction text to be sent to the large language model server 20001. Information of types other than the text information of natural language extracted from the user's voice is, for example, images, videos, audio, and the like.

[0157] Next, the character conversation device (artificial intelligence response output device 10010) of this example sends the instruction text to the large language model server 20001 using the API. Also in this example, the instruction text may be metadata storing information described by a notation using tags such as the markup format of a markup language, a notation using predetermined symbols such as the Markdown format, or an object notation of a predetermined script such as JSON. Also in this example, as types of the instruction text, there are a setting instruction text storing instructions such as initial settings and a user instruction text reflecting instructions from the user. Type identification information for identifying whether the instruction text is a setting instruction text or a user instruction text may be stored in a part other than the main message of the instruction text. At this time, the instruction text includes text information of natural language as the main message. Further, in this example, in addition to the text information of natural language, non-natural language information sources such as images, videos, or audio can be included in the main message of the instruction text as information of types other than the text information of natural language. A specific method for including non-natural language information sources in the instruction text will be described later.

[0158] The large-scale language model server 20001 of this embodiment has a multimodal large-scale language model that can process non-natural language information sources in addition to natural language text information. The large-scale language model server 20001 receives an instruction sentence from a character conversation device (artificial intelligence response output device 10010). Based on the instruction sentence, the multimodal large-scale language model executes an inference and generates a response including natural language text information that is the result of the inference. Here, since the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, in addition to natural language text information, the response can include non-natural language information sources such as images, videos, or sounds.

[0159] The character conversation device (artificial intelligence response output device 10010) receives the response from the large-scale language model server 20001 and extracts the natural language text information stored as the main message in the response and non-natural language information sources such as images, videos, or sounds. The character operation program of the character conversation device (artificial intelligence response output device 10010) generates natural language speech that is the response to the user using speech synthesis technology based on the natural language text information extracted from the aforementioned response, and may output it from the audio output unit 1140, which is a speaker, so that it sounds like the voice of the character 19051 being displayed on the display screen.

[0160] Also, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display natural language characters that are the response to the user on the display screen of the character conversation device (artificial intelligence response output device 10010) based on the natural language text information extracted from the aforementioned response. At this time, the characters may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed instead of the video of the character 19051. These specific processes may be executed by the video control unit 1160.

[0161] Also, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the image of the non-natural language information source extracted from the above-mentioned response on the display screen of the character conversation device (artificial intelligence response output device 10010) for presenting to the user based on the information of the image. At this time, the image may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed instead of the video of the character 19051. These specific processes may be executed by the video control unit 1160.

[0162] Also, the character operation program of the character conversation device (artificial intelligence response output device 10010) may display the video of the non-natural language information source extracted from the above-mentioned response on the display screen of the character conversation device (artificial intelligence response output device 10010) for presenting to the user based on the information of the video. At this time, the video may be displayed together with the character 19051, may be superimposed on the video of the character 19051, or may be displayed instead of the video of the character 19051. These specific processes may be executed by the video control unit 1160.

[0163] Also, the character operation program of the character conversation device (artificial intelligence response output device 10010) may output the voice generated based on the voice information of the non-natural language information source extracted from the above-mentioned response from the voice output unit 1140 which is a speaker.

[0164] According to the character conversation device (artificial intelligence response output device 10010) shown in FIG. 3C described above, or the character conversation system including the character conversation device (artificial intelligence response output device 10010) and the large language model server 20001, it is not necessary to install the large language model itself that requires huge amounts of data for learning and computing resources in the character conversation device (artificial intelligence response output device 10010) itself. Moreover, through the API, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large language model can be utilized. In response to an action on the character from the user, in addition to an answer based on the text of natural language, an answer based on a non-natural language information source can be given, making it possible to conduct a more suitable conversation.

[0165] Next, with reference to FIG. 3D, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 20001. Specifically, FIG. 3D shows the natural language text of the main message of the instruction text sent from the character conversation device (artificial intelligence response output device 10010) to the large language model server 20001 and an example of a non-natural language information source such as an image, and an example of the natural language text of the main message of the server response as the response. In this embodiment, as the non-natural language information source, an image, a video, an audio, etc. can be used, but in FIG. 3D, an example of an image is shown as the non-natural language information source.

[0166] Also, in FIG. 3D, it is shown that the exchange of instructions and responses is made in time series from the setting instruction text, the first round of user instruction text and its response to the second round of user instruction text and its response. Here, the instructions and responses shown in FIG. 3D include non-natural language information sources 20061 and 20062 that were not shown in FIG. 2D of Embodiment 2. In the example of FIG. 3D, both the non-natural language information source 20061 and the non-natural language information source 20062 are images.

[0167] Here, in FIG. 3D, for simplicity of explanation, it is shown with an image of the non-natural language information source 20061 pasted in the instruction text. However, there are multiple methods for transmitting or designating the data of the non-natural language information source 20061 in the instruction text sent from the character conversation device (artificial intelligence response output device 10010) to the large language model server 20001. The character conversation device (artificial intelligence response output device 10010) may use any one of these multiple methods or switch between them. An example of each method will be described below.

[0168] The first method of transmitting or designating non-natural language information source data in the instruction text is used, for example, when the non-natural language information source to be specified is a non-natural language information source existing in a location such as a server connected to a network such as the Internet. As a specific method of the first method, a non-natural language information source file existing on a network such as the Internet is specified by using information such as tags and symbols in the instruction text with the location information (such as a so-called URL) of the network such as the Internet and the file name.

[0169] For example, using a tag for specifying an image in a markup language or the like <img src=""****”"> by describing the location information and file name information of the image file in the **** part, an image existing on a network such as the Internet may be specified. Also, a tag for specifying a video in a markup language or the like <video src=""****”">By using [it] and describing the location information and file name information of the video file in the **** part, a video existing on a network such as the Internet may be specified. Also, <audio src=""****”">By using [specific method] and describing the location information and file name information of the audio file in the **** part, it is possible to specify the audio existing on a network such as the Internet. Also, in the case of the JSON format notation, by preparing keys such as img_src and describing the location information and file name information of the image file in the value, it is possible to specify the image existing on a network such as the Internet. In the case of video files and audio files, respective keys and values may be prepared. The specific example of the format is just an example, and other unique formats may be used. In any case, the information specifying the location information and file name information of the non-natural language information source file may be stored in the instruction sentence.

[0170] When storing the information specifying the location information and file name information of the non-natural language information source file in the instruction sentence as in the first method, it is not necessary to store the data of the non-natural language information source file itself in the instruction sentence itself. Therefore, the data amount of the instruction sentence can be reduced. When the large language model server 20001 receives the instruction sentence in which the non-natural language information source data is specified by the first method, it may acquire the non-natural language information source file located at a place such as a server connected to a network such as the Internet by using the location information and file name information of the non-natural language information source file stored in the instruction sentence.

[0171] Here, when the character conversation device (artificial intelligence response output device 10010) specifies the non-natural language information source data in the instruction sentence by the first method, how to input the location information and file name information will be explained. In FIG. 3C, in this embodiment, it was explained that the types of actions that can be performed by the user 230 on the character conversation device (artificial intelligence response output device 10010) have increased compared to Embodiment 2, in addition to the voice of the user 230. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data and file name information by a user operation (for example, mouse, keyboard, touch panel) via the operation input unit 1107 in FIG. 1B.

[0172] Also, in the character conversation device (artificial intelligence response output device 10010), the control unit 1110 may cooperate with the memory 1109 to execute a WEB browser program and display the GUI of the WEB browser program on the display screen of the character conversation device (artificial intelligence response output device 10010). A user operation on the GUI of the WEB browser program may be received by a user operation via the operation input unit 1107 (for example, mouse, keyboard, touch panel) or a touch operation input sensor of the display unit 10011, and non-natural language information source data such as an image, video, or audio selected on the browser screen of the WEB browser program may be used as the data to be specified in the instruction text. In this case, the WEB browser program may acquire the location information and file name information of the non-natural language information source data and deliver them to the character motion program.

[0173] Also, the user 230 may operate the mobile information processing terminal 20010 to communicate with the character conversation device (artificial intelligence response output device 10010) from the mobile information processing terminal 20010, and input location information such as a URL for specifying non-natural language information source data to the character conversation device (artificial intelligence response output device 10010). Also, in the manner described with reference to FIG. 3C, an information storage image such as a two-dimensional code may be displayed on the display panel 20011 of the mobile information processing terminal 20010, and image recognition processing may be performed on the image captured by the imaging unit 1180 of the character conversation device (artificial intelligence response output device 10010) to obtain the result of the image recognition processing, and location information such as a URL for specifying non-natural language information source data, file name information, etc. may be input.

[0174] Note that the use of the first method for transmitting or specifying non-natural language source data in an instruction text is not limited to the case where the non-natural language source file exists in a location such as a server previously connected to a network such as the Internet. For example, when it is desired to include non-natural language source data such as images, videos, and audio stored in the storage unit 1170 of the character conversation device (artificial intelligence response output device 10010) in the instruction text, the character conversation device (artificial intelligence response output device 10010) uploads the non-natural language source data to a second server 19002 via the Internet 19000, and may include in the instruction text the location information (such as a so-called URL) of the non-natural language source data of the second server 19002 on the Internet and the file name. In this case, the second server 19002 functions as a so-called intermediate server.

[0175] Similarly, when it is desired to include non-natural language source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in the instruction text, the mobile information processing terminal 20010 may upload the non-natural language source data to the second server 19002 via the Internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit to the character conversation device (artificial intelligence response output device 10010) the location information (such as a so-called URL) of the non-natural language source data of the second server 19002 on the Internet and the file name, and the character operation program of the character conversation device (artificial intelligence response output device 10010) may include in the instruction text the location information (such as a so-called URL) of the non-natural language source data uploaded to the second server 19002 and the file name that have been acquired.

[0176] Furthermore, a media server accessible from other servers via the Internet 19000 may be constructed within the character conversation device (artificial intelligence response output device 10010) in cooperation with the memory 1109 and the storage unit 1170 by the character operation program of the character conversation device (artificial intelligence response output device 10010). In this case, when the character conversation device (artificial intelligence response output device 10010) designates non-natural language information source data in the instruction text by the first method, the location information on the Internet (such as a so-called URL) indicating the media server constructed inside the character conversation device (artificial intelligence response output device 10010) itself and the file name of the corresponding non-natural language information source data may be stored in the instruction text.

[0177] Next, a second method of designating the transmission or designation of non-natural language information source data in the instruction text is, for example, a method of simply storing (attaching) the non-natural language information source data itself in the instruction text (prompt) and transmitting it. Generally, non-natural language information source data such as images, videos, and voices has a larger data volume than text information which is natural language. Therefore, in this case, the data volume of the instruction text (prompt) itself becomes larger than that of the first method. The character operation program of the character conversation device (artificial intelligence response output device 10010) stores the non-natural language information source data to be stored (attached) in the instruction text (prompt) in the memory 1109 once, and when transmitting the instruction text (prompt), it may be output from the memory 1109 via the communication unit 1132 and stored (attached) in the instruction text (prompt) to the large language model server 20001. The non-natural language information source data itself stored in the memory 1109 by the character operation program of the character conversation device (artificial intelligence response output device 10010) may be acquired by the communication unit 1132 via the Internet 19000, may be acquired by the communication unit 1132 from the mobile information processing terminal 20010, or may be read from the storage unit 1170 and stored in the memory 1109.

[0178] By the method described above, the character conversation device (artificial intelligence response output device 10010) can transmit or specify non-natural language information source data by an instruction sentence.

[0179] Since the large language model server 20001 is a multimodal large language model that can process non-natural language information sources together with natural language text information, as shown in the example of Fig. 3D, in the first round of the user instruction sentence, the non-natural language information source 20061, which is an image of a swimming pool and the pool side, and the text information in natural language are acquired. As an inference result, as a response to the first round of the user instruction sentence, text information in natural language as shown in the figure can be output.

[0180] Also, since the large language model server 20001 is a multimodal large language model that can process non-natural language information sources together with natural language text information, as shown in the response to the second round of the user instruction sentence in the example of Fig. 3D, the large language model server 20001 can include the non-natural language information source 20062 generated by the inference of the multimodal large language model in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). In Fig. 3D, the non-natural language information source 20062 shows an example of an image in which a circle image is attached to the image of the swimming pool and the pool side, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Fig. 3D and may be a video or audio.

[0181] For the method of including a non-natural language information source other than natural language text information in the response from the large language model server 20001, a method conforming to the first method or the second method in which the above-described character conversation device (artificial intelligence response output device 10010) transmits or specifies non-natural language information source data in an instruction sentence may be used.

[0182] Specifically, as a method conforming to the above-described first method, the large language model server 20001 may store, in the instruction sentence, information specifying the location information and file name information of the non-natural language information source file in the response. The non-natural language information source 20062 itself, such as an image, video, or audio, may be held by the large language model server 20001, or the non-natural language information source 20062 may be transferred to the second server 19002 that functions as an intermediate server and held there. In either case, the large language model server 20001 may store, in the response, information specifying the location information and file name information of the non-natural language information source file in the instruction sentence. The character conversation device (artificial intelligence response output device 10010) that has obtained the response may access the large language model server 20001 or the second server 19002 using the location information and file name information of the non-natural language information source file described in the instruction sentence to obtain the non-natural language information source 20062.

[0183] Also, specifically, as a method conforming to the above-described second method, the large language model server 20001 may store (attach) the file data itself of the non-natural language information source 20062 in the response and transmit it to the character conversation device (artificial intelligence response output device 10010). The character conversation device (artificial intelligence response output device 10010) may obtain the data of the non-natural language information source 20062 stored (attached) in the instruction sentence and use it for various outputs to the user 230.

[0184] According to the operations of the character conversation device (artificial intelligence response output device 10010) and the character conversation system of Example 3 described above with reference to FIG. 3D, transmission and reception of instructions and responses for realizing a conversation using images, videos, and audio, which are non-natural language information, are performed between the character displayed on the character conversation device (artificial intelligence response output device 10010) and the user 230. As a result, it becomes possible to realize a more advanced and natural conversation as shown in each message of FIG. 3D.

[0185] Next, with reference to FIG. 3E, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 20001. Specifically, FIG. 3E shows an example of the main message of the instruction text transmitted from the artificial intelligence response output device 10010 to the large language model server 20001, which is the basis for the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response thereto.

[0186] FIG. 3E shows an example of a case where, after the series of conversations shown in FIG. 3D and after the continuation of the series of conversations has ended, the user 230 speaks to the character 19051 again to start a new conversation. In the example of FIG. 3E, the processing using the conversation history as described in FIGS. 2F, 2G, and 2I of Embodiment 2 is not performed. Therefore, similar to FIG. 2E of Embodiment 2, FIG. 3E shows a response of the content in a state where it does not remember at all the name of the large language model itself included in the setting instruction text, the role to be played, the characteristics of the conversation, the name of the user, the conversation history, etc.

[0187] Next, with reference to FIG. 3F, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 20001. Specifically, FIG. 3F shows an example of the main message of the instruction text transmitted from the artificial intelligence response output device 10010 to the large language model server 20001, which is the basis for the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response that is the response thereto.

[0188] Figure 3F shows an example of a case where, after the end of the continuation of the series of conversations shown in Figure 3D, user 230 speaks to character 19051 again to start a new conversation. Here, in Figure 3F, the method of storing a message explaining the history of past conversations in the setting instruction text, which was described in Figure 2F of Example 2, is also applied to the character conversation device (artificial intelligence response output device 10010) of Example 3. Specifically, in Figure 3F, the message that is the content of the setting instruction text in Figure 3D is stored as a reset message, and following the reset message, a message explaining the history of past conversations is stored as a conversation history message.

[0189] Since the large language model server 20001 of Example 3 is a multimodal large language model that can process non-natural language information sources in addition to natural language text information, there may be cases where non-natural language information source data is transmitted or specified in past instructions and responses. Therefore, in the example of Figure 3F, the conversation history message reflects not only the natural language text information in past instructions and responses but also the transmission or specification of non-natural language information source data in past instructions and responses. The specific method of transmitting or specifying non-natural language information source data in the instruction of Figure 3F is the same as the transmission or specification of non-natural language information source data as described in Figure 3D, so repeated explanations are omitted.

[0190] In the example of Figure 3D, the method of transmitting or specifying non-natural language information source data may include cases where the non-natural language information source data itself is stored (attached) in the instruction and cases where the non-natural language information source data is not stored (attached) in the instruction. This is the same for the instruction of Figure 3F.

[0191] Next, with reference to FIG. 3G, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 20001. Specifically, FIG. 3G shows an example of the main message of the instruction text transmitted from the artificial intelligence response output device 10010 to the large language model server 20001, which is the basis of the conversation between the character 19051 displayed on the artificial intelligence response output device 10010 and the user 230, and the main message of the server response as the response thereto.

[0192] FIG. 3G shows an example of a series of conversations from the first round of user instruction text following the initial setting instruction text to the third round of user instruction text and its response in the series of conversations shown in FIG. 3F. In FIG. 3G, it is shown that the exchange of instructions and responses is made in chronological order. Since the content of the setting instruction text is as shown in FIG. 3F, repeated description is omitted.

[0193] As described above, even when using the large language model server 20001 having a multimodal large language model capable of processing non-natural language information sources in addition to the natural language text information of Embodiment 3, after the series of conversations has ended and the user 230 speaks to the character 19051 again to start a new conversation, if the generation process and transmission process of the setting instruction text in FIG. 3F are performed, the response to the subsequent user instruction text will, as shown in FIG. 3G, reflect the settings and conversation history such as the role, name, conversation characteristics, personality, and / or conversation characteristics of the character at the time of the previous conversation. As a result, from the user's perspective, it is recognized that the identity of the settings and memories such as the role, name, conversation characteristics, or personality of the character at the time of the previous conversation is more securely ensured, which is more preferable.

[0194] Next, with reference to FIG. 3H, an example of the operation of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described. This can also be said to be an explanation of an example of the operation of a character conversation system including the artificial intelligence response output device 10010 and the large language model server 20001. Specifically, FIG. 3H is an explanatory diagram of a database 20200 for managing character settings and character conversation histories for a plurality of characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Here, for the settings of the plurality of characters displayed on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010), the example described in FIG. 2H of Embodiment 2 is used. Therefore, repeated explanations of the settings of the plurality of characters and the like are omitted.

[0195] Also, the database 20200 for managing character settings and character conversation histories shown in FIG. 3H has the same format as the database 19200 shown in FIG. 2I of Embodiment 2. In FIG. 3H, only the differences from the database 19200 shown in FIG. 2I will be described. Also, the content of the character "Koto" in the database will be described, and the content of other characters will be omitted.

[0196] Here, as described above, since the large language model server 20001 of Example 3 is a multimodal large language model that can process non-natural language information sources in addition to natural language text information, not only natural language text information but also the transmission or specification of non-natural language information source data is included in the instruction text from the character conversation device (artificial intelligence response output device 10010) and the response from the large language model server 20001. Therefore, in the database 20200 shown in FIG. 3H, in the conversation history data, not only the natural language text information included in these instruction texts and responses but also the information on the transmission or specification of non-natural language information source data is recorded. The specific method of transmitting or specifying non-natural language information source data in the recording of the conversation history is the same as the transmission or specification of non-natural language information source data described in FIG. 3D, so repeated explanations are omitted.

[0197] In the example of FIG. 3D, the method of transmitting or specifying non-natural language information source data may include storing (attaching) the non-natural language information source data itself in the instruction text and not storing (attaching) the non-natural language information source data in the instruction text. This is the same for the conversation history in FIG. 3H. However, in the conversation history of FIG. 3H, as a method of specifying non-natural language information source data, when specifying the location information and file name information of the non-natural language information source file of a server (the second server 19002 functioning as an intermediate server or other cloud servers) existing on a network such as the Internet, if the conversation history period is long, the non-natural language information source file on the server may be deleted. Then, it may not be possible to acquire the non-natural language information source file at a later date using the location information and file name information, and the information in the conversation record may be missing.

[0198] To prevent this, when the character conversation device (artificial intelligence response output device 10010) converts and records the instruction text and response messages into the conversation history, the non-natural language information source file itself specified in the instruction text and response should be obtained from a server or the like on the network using the location information and file name information, and stored in the storage unit 1170. Furthermore, the location information and file name of the non-natural language information source file should be rewritten into the location information on the Internet (such as a so-called URL) indicating the inside of the media server of the media server constructed within the character conversation device (artificial intelligence response output device 10010) by the character operation program of the character conversation device (artificial intelligence response output device 10010), and then recorded in the conversation record. In this way, as long as the character conversation device (artificial intelligence response output device 10010) itself does not delete the non-natural language information source file from the storage unit 1170, the non-natural language information source will not be missing from the conversation record information, which is more suitable for the preservation of the conversation record.

[0199] By using the database of FIG. 3H described above, even when the character conversation device (artificial intelligence response output device 10010) is configured to switch and display the characters displayed on the display unit 10011 from among a plurality of character candidates, from the user's perspective, there is less discomfort felt from the conversation with each character, and it is possible to share the memory with each of the plurality of characters, obtaining the effect of FIG. 2I of Example 2, that is, a more enjoyable character conversation experience. Also, when the large language model server 20001 is a multimodal large language model capable of processing non-natural language information sources in addition to natural language text information, this effect can also be demonstrated.

[0200] Note that in the character conversation device (artificial intelligence response output device 10010) or the character conversation system of Example 3, a multimodal large language model artificial intelligence capable of processing non-natural language information other than natural language text information in addition to natural language text information is used in the large language model server 20001.

[0201] Here, communication between the character conversation device (artificial intelligence response output device 10010) and the large language model server 20001 is carried out using an API. In a multimodal large language model, in addition to processing the text information of natural language into units of words separated by sentences called tokens, there may also be a form in which the usage fee of the API is billed according to the data volume of the non-natural language information source.

[0202] Therefore, in order to provide the character conversation service by the character conversation system according to this embodiment to the user at a lower cost, the following modified examples may be used.

[0203] As a first modified example, in the recording of the conversation history in the database of FIG. 3H, the transmission or designation information of the non-natural language information source data is also recorded. However, the character and the user exchange conversations about the natural language information source data in the text information of natural language, and the content is recorded in the text information of natural language. Then, in the recording of the conversation history in the database of FIG. 3H, even if the recording of the transmission or designation information of the natural language information source data is omitted, the conversation itself about the natural language information source data will be recorded to a certain extent as the text information of natural language. Therefore, if a certain amount of information reduction is allowed, in the recording of the conversation history in the database of FIG. 3H, the recording of the transmission or designation information of the natural language information source data may be omitted. In this case, the transmission or designation information of the natural language information source data is also omitted from the conversation history message of the setting instruction text in FIG. 3F. Thereby, the data volume of the non-natural language information source communicated using the API can be reduced.

[0204] Next, as a second modification example, in the recording of the conversation history in the database of FIG. 3H, instead of transmitting non-natural language source data or recording designated information, an example is to record text information in natural language that describes the content of the non-natural language source data. The text information in natural language that describes the content of the non-natural language source data can be obtained, for example, by starting a conversation between the large language model of the large language model server 20001 and the character conversation device (artificial intelligence response output device 10010) separately from the conversation as a character, and having the large language model server 20001 describe the content of the non-natural language source data with a specified number of characters limit. Also, the content of the non-natural language source data can be obtained by having a conversation with another large language model of another server that can be used at a lower cost than the large language model of the large language model server 20001, with a specified number of characters limit. Further, if alternative text data is prepared from the time when the non-natural language source data is acquired, the alternative text data may be used as the text information in natural language that describes the content of the non-natural language source data. As a specific example of the alternative text data of the non-natural language source data, tags of markup language <img src=""”alt="****""> , <video src=""”alt="****""> 、 <audio src=""”alt="****"">It is the text information described in the **** part such as

[0205] Also, in the case of JSON format notation, in the object stored in association with the location information of the non-natural language information source data and the file name information, which are the key and value indicating the location information of the non-natural language information source data, further, a key corresponding to the alternative text and the value which is the alternative text data itself may be associated and stored.

[0206] Also in this case, in the recording of the conversation history in the database of FIG. 3H, the recording of the transmission or designation information of the natural language information source data can be omitted, and the transmission or designation information of the natural language information source data is also omitted from the conversation history message of the setting instruction text in FIG. 3F. Thereby, the data amount of the non-natural language information source communicated using the API can be reduced.

[0207] Next, as a third modification example, at the time of the first pass of the user instruction text in FIG. 3D, instead of storing the transmission or specified information of the non-natural language information source data in the user instruction text, it is an example of replacing it with text information in natural language that describes the content of the non-natural language information source data. For example, in the first pass of the user instruction text in FIG. 3D, instead of the transmission or specified information of the non-natural language information source data 20061 in the user instruction text, an explanatory text such as "This image is of a swimming pool, a sheet on the pool side, and a parasol. There is water in the swimming pool. There is a drink on the table beside the sheet" is stored as text information in natural language. At this time, the explanatory text may be obtained by having a conversation with another large language model of another server that can be used at a lower cost than the large language model of the large language model server 20001 to explain the content of the non-natural language information source data within a specified number of characters. Also, the explanatory text may be obtained from a server of other various services that can obtain an overview and content description of non-natural language information source data such as images, videos, and sounds. Further, if alternative text data is prepared from the time of acquisition of the non-natural language information source data, the alternative text data may be used as text information in natural language that describes the content of the non-natural language information source data.

[0208] Next, an example of a display example of the character conversation device (artificial intelligence response output device 10010) according to Embodiment 3 of the present invention will be described with reference to FIG. 3I. In the example of FIG. 3I, it shows an example of displaying the response from the large language model to the instruction text from the user described in each of FIGS. 3A to 3H on the display unit 10011 of the character conversation device (artificial intelligence response output device 10010). Specifically, it is an example of displaying the text 10063 of the natural language information source data, which is the response from the large language model, the image 10064 of the non-natural language information source data, and / or the video 10065 of the non-natural language information source data on the display unit 10011 together with the video of the character 19051. The text 10063, image 10064, and / or video 10065, which are the responses from the large language model, may be displayed superimposed in front of the video of the character 19051 as shown in FIG. 3I.

[0209] Also, the text 10063, image 10064, and / or video 10065, which are responses from the large language model, may be displayed together with the video of the character 19051 without overlapping the video of the character 19051. The display in Fig. 3I is an example. For example, when the user 230 operates via the touch operation input sensor of the operation input unit 1107 or the display unit 10011 to minimize the volume of the voice output of the voice output unit 1140 of the character conversation device (artificial intelligence response output device 10010) or sets the voice output to OFF, etc., the user 230 cannot confirm the response from the large language model by voice. Therefore, in this case, the control unit 1110 may control to start a display mode of displaying the text 10063, image 10064, and / or video 10065, which are responses from the large language model, together with the video of the character 19051 as shown in Fig. 3I.

[0210] In this way, even when the user wants to refrain from voice output, the user 230 can use the character conversation device (artificial intelligence response output device 10010) more suitably. Note that the ON / OFF of the display mode in which the text 10063, image 10064, and / or video 10065, which are responses from the large language model, are displayed together with the video of the character 19051 may be configured to be manually switchable by an operation of the user 230 via the touch operation input sensor of the operation input unit 1107 or the display unit 10011. According to the display example in Fig. 3I, in the character conversation device (artificial intelligence response output device 10010) corresponding to multimodality, it is possible to more suitably output the response from the large language model.

[0211] According to the character conversation device and the character conversation system according to Embodiment 3 described above, in addition to the effects in the character conversation device and the character conversation system according to Embodiment 2, a more advanced conversation experience including non-natural language information in addition to natural language information can be provided to the user by using a multimodal large language model. Further, according to the character conversation device and the character conversation system according to Embodiment 3, the character conversation service can be provided to the user at a lower cost.

[0212] In the above description of Embodiment 3, an example of using the large language model owned by the large language model server 20001 as the large language model has been described. In contrast, the character conversation device (artificial intelligence response output device 10010) may include the local LLM processing unit 10028 shown in FIG. 1B, and a multimodal large language model owned by the local LLM processing unit 10028 may be used. In this case, instead of the multimodal large language model owned by the large language model server 20001, the multimodal large language model owned by the local LLM processing unit 10028 may be used.

[0213] In this case, in the above description of Embodiment 3, the multimodal large language model of the large language model server 20001 may be replaced with the multimodal large language model of the local LLM processing unit 10028 of the character conversation device (artificial intelligence response output device 10010). Also in this case, by using the multimodal large language model, a more advanced conversation experience including non-natural language information in addition to natural language information can be provided to the user. When using the multimodal large language model of the local LLM processing unit 10028 instead of the multimodal large language model of the large language model server 20001, the need to consider usage fees according to the number of processing tokens and the data volume of non-natural language information sources is reduced. However, even if it is the multimodal large language model of the local LLM processing unit 10028, by reducing the number of processing tokens and the data volume of non-natural language information sources, consumption resources such as power required for inference can be reduced. In this case, a character conversation service with less power consumption can be provided to the user.

[0214] Note that the configuration of uploading and downloading the conversation history with the character and the data of the database including the conversation history with the character described in Embodiment 2 to the second server 19002 or other cloud servers can also be used in the example of using the multimodal large language model described in Embodiment 3. Also in this case, when a user has conversations with a plurality of characters at different times between different individual devices, and the characters of each of the plurality of characters are different, a conversation can be realized as if the memory of each character was pseudo-continuously inherited from the previous conversation, which is more suitable for the user.

[0215] <Embodiment 4> Next, Example 4 of the present invention is an improvement of the artificial intelligence response output device 10010, the character conversation device, or the system thereof described in each figure of Example 2 or Example 3. In this example, the differences from Example 2 or Example 3 will be described, and the repeated description of the same configurations as those in these examples will be omitted.

[0216] Similar to the above-described embodiments, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. A system including the artificial intelligence response output device 10010 and a large language model server may also be referred to as an artificial intelligence response output system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system.

[0217] Using FIG. 4A, an example of the operation using the database in the character conversation device (artificial intelligence response output device 10010) of Example 4 of the present invention will be described. The database according to Example 4 shown in FIG. 4A is an extension of the database described in FIG. 2I or FIG. 3I. Specifically, the database shown in FIG. 4A assumes a case where a plurality of different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores the initial setting instructions and conversation histories corresponding to each user and character in the database.

[0218] In the example of FIG. 4A, for user 1 with user ID 1, the initial setting instructions and conversation histories of each character of character Koto with character ID 1, character Tom with character ID 2, and character Necco with character ID 3 are stored. In addition to this, for each of user 2 with user ID 2 and user 3 with user ID 3, the initial setting instructions and conversation histories of each character of character Koto with character ID 1, character Tom with character ID 2, and character Necco with character ID 3 are also stored.

[0219] These initial setting instructions and conversation history data are stored as separate data in different areas for each combination of user and character. For the sake of explanation in FIG. 4A, the data stored in each area is denoted as data 11, 12, 13, 21, 22, 23, 31, 32, and 33. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses the initial setting instructions and conversation history stored in different areas for each combination of user and character based on the user who is currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system, so that the consistency of the character's personality and the continuity of memory can be more suitably maintained for each of different users.

[0220] Specifically, consider the situation where user 1 has previously conversed with character Tom using the character conversation device (artificial intelligence response output device 10010), and user 2 is unaware of that conversation. Subsequently, when user 2 converses with character Tom. At this time, if the artificial intelligence response output device 10010 uses a database of initial setting instructions or conversation history data that does not identify the user, the response output from the artificial intelligence response output device 10010 may be based on a conversation history that is not in user 2's memory, and the conversation between user 2 and the character of the artificial intelligence response output device 10010 may become inconsistent.

[0221] On the other hand, even in the same situation, if the database shown in FIG. 4A is used, the control unit 1110 of the character conversation device (artificial intelligence response output device 10010) identifies the user by ID, stores the initial setting instruction text and the conversation history in different areas for each user, and uses the initial setting instruction text and the conversation history stored in different areas for each user to generate the artificial intelligence response. Thereby, the initial setting instruction text and the conversation history used for generating the artificial intelligence response for each user are based on the operation or the course of the conversation of that user, and are managed separately from the operation or the course of the conversation of other users. Thereby, the consistency of the conversation history between each user and each character of the artificial intelligence response output device 10010 can be made more suitable.

[0222] Note that the database of the initial setting instruction text and / or the conversation history described in FIG. 4A may be stored in the storage unit 1170 of the artificial intelligence response output device 10010 and used by the control unit 1110. Also, not limited to this, the database of the initial setting instruction text and / or the conversation history may be stored in a server on the network. For example, when the artificial intelligence response output device 10010 uses the large language model of the large language model server 19001 or the multimodal large language model of the large language model server 20001 in generating the artificial intelligence response, the database of the initial setting instruction text and / or the conversation history described in FIG. 4A may be stored in these servers themselves. In this way, the process of re-including the initial setting instruction text and the conversation history in the instruction text and transmitting them from the artificial intelligence response output device 10010 to these servers can be omitted, and the number of transmission tokens for using the large language model can be saved.

[0223] When storing the initial setting instructions and / or the conversation history database described with reference to FIG. 4A, the user ID, character ID, and user instructions for subsequent conversations may be transmitted from the artificial intelligence response output device 10010 to these servers. The large language model on these servers uses the user ID and character ID obtained from the artificial intelligence response output device 10010 to retrieve the corresponding initial setting instructions and conversation history from the initial setting instructions and / or conversation history database of FIG. 4A. The large language model on these servers may perform inference using the initial setting instructions and conversation history and the user instructions for subsequent conversations transmitted from the artificial intelligence response output device 10010, generate an artificial intelligence response, and transmit it to the artificial intelligence response output device 10010. In this way, it is possible to obtain the effect of more suitably maintaining the consistency of the character's personality and the continuity of memory for each of different users while saving the number of transmission tokens for using the large language model.

[0224] Next, with reference to FIG. 4B, an example of the operation using the database in the character conversation device (artificial intelligence response output device 10010) according to Embodiment 4 of the present invention will be described. The database according to Embodiment 4 shown in FIG. 4B is an extension of the database described with reference to FIG. 1C or FIG. 2L. Specifically, the database shown in FIG. 4B assumes a case where a plurality of different users use the same character conversation device (artificial intelligence response output device 10010) or the same character conversation system, and stores in the database the data of response templates corresponding to each user and character.

[0225] In the example of FIG. 4B, for User 1 with a user ID of 1, response template data for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are stored. In addition to this, for each of User 2 with a user ID of 2 and User 3 with a user ID of 3, response template data for each of the characters Koto with a character ID of 1, Tom with a character ID of 2, and Necco with a character ID of 3 are also stored.

[0226] These response template data are stored as different data in different regions for each combination of user and character. For the sake of explanation in FIG. 4B, the data stored in each region are denoted as response template data 101, 102, 103, 201, 202, 203, 301, 302, 303. For example, the response template data 101 stores a database such as a table corresponding to the response templates for condition numbers 1 to 7 of Character 1: Koto shown in FIG. 2L. The data 201 in FIG. 4B stores a database such as a table corresponding to the response templates for condition numbers 1 to 7 of Character 2: Tom shown in FIG. 2L.

[0227] The data 301 in FIG. 4B stores a database such as a table corresponding to the response templates for condition numbers 1 to 7 of Character 3: Necco shown in FIG. 2L. The data 102, 202, 302 in FIG. 4B store response templates modified for User 2 in the same format. The data 103, 203, 303 in FIG. 4B store response templates modified for User 3 in the same format. The control unit 1110 of the character conversation device (artificial intelligence response output device 10010) uses the response template data stored in different regions for each combination of user and character based on the user currently using (logging in to) the character conversation device (artificial intelligence response output device 10010) or its system.

[0228] By doing so, even for the same character, it becomes possible to make responses using different response templates for each user. That is, even for the same character, depending on the relationship between the character and the user, it may be more appropriate to change the content of the response template. For example, depending on the relationship between the age set for the character and the age of the user registered in the artificial intelligence response output device 10010 or the system, the user may be older than the character, the same age as the character, or younger than the character. At this time, it is more appropriate or natural for the conversation between the user and the character if the content of the response template for the character to an older user, the response template for the character to a user of the same age, and the response template for the character to a younger user are each changed. That is, by performing the operation using the database of FIG. 4B, it becomes possible to produce a more appropriate or natural conversation by varying the content of the response template for each relationship between the character and the user.

[0229] Incidentally, the response template sentence database (response template sentence DB) of FIG. 4B described above is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 may use this. However, the response template sentence database (response template sentence DB) shown in FIG. 4B may be provided on the side of the large language model server 19001 or the side of the large language model server 20001. In this case, the control unit of the large language model server 19001 or the control unit of the large language model server 20001 may generate a response using the response template sentence database (response template sentence DB). Instead of the response generated by the large language model stored in each server, the control unit of the large language model server 19001 or the control unit of the large language model server 20001 may transmit the response generated using the response template sentence database (response template sentence DB) to the artificial intelligence response output device 10010. In this way, even when the artificial intelligence response output device 10010 is not provided with the response template sentence database (response template sentence DB), it is possible to generate a response using the response template sentence database (response template sentence DB).

[0230] According to the character conversation device and the character conversation system according to the fourth embodiment described above, it is possible to produce a more suitable or more natural conversation according to the relationship between the character and the user, the conversation history, and the like.

[0231] <Embodiment 5> Next, the fifth embodiment of the present invention is an improvement of the artificial intelligence response output device 10010 or the artificial intelligence response output system described with reference to the drawings of the first, second, and third embodiments. Specifically, it is an example of performing a process of switching the response generation process of the artificial intelligence response output device 10010 from the response generation process by a large language model on the network to the response generation process by a local large language model (such as the local LLM processing unit 10028) provided in the artificial intelligence response output device 10010 or the response generation process by the response template sentence database. In this embodiment, the differences from these embodiments will be described, and repeated descriptions of the same configurations as these embodiments will be omitted.

[0232] Similar to the above embodiments, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. A system including the artificial intelligence response output device 10010 and a large language model server may also be referred to as an artificial intelligence response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system.

[0233] Using FIG. 5A, an example of the switching process of the response generation process in the artificial intelligence response output device 10010 according to Embodiment 5 of the present invention will be described. In the table of FIG. 5A, Examples 1 to 9 are shown for an example of the switching process of the response generation process in the artificial intelligence response output device 10010. In the table of FIG. 5A, the column of "Switching Summary" shows the summary of the switching process for each example. The column of "State before switching of LLM (API-connected LLM) on the network" shows the state before the response generation process by the large language model on the network (such as the large language model provided by the large language model server 19001 in FIG. 1 and the multimodal large language model provided by the large language model server 20001) is switched to another response generation process. The column of "Switching Occurrence Condition" shows the condition under which the switching process of the response generation process occurs. The column of "Switching Destination from LLM (API-connected LLM) on the Network" shows the switching destination to which the response generation process of the artificial intelligence response output device 10010 is switched from the large language model on the network (such as the large language model provided by the large language model server 19001 and the multimodal large language model provided by the large language model server 20001) that is connected using an API. The control unit 1110 of the artificial intelligence response output device 10010 may perform control to switch to the large language model, database, or correspondence shown in the "Switching Destination from LLM (API-connected LLM) on the Network" when the condition shown in the "Switching Occurrence Condition" occurs in the state of "State before switching of LLM (API-connected LLM) on the network" shown in FIG. 5A.

[0234] Next, each example shown in the table of FIG. 5A will be described. Example 1 is an example of performing switching according to the connection availability state of the network of the artificial intelligence response output device 10010, as shown in the "Switching Overview". In Example 1, as the "state before switching the LLM (API-connected LLM) on the network", it is shown that the connection state of the network of the artificial intelligence response output device 10010 is a connectable state. Here, in Example 1, as the "switching occurrence condition", it is shown that "when the network connection becomes unavailable". That is, this means when the connection via the network between the artificial intelligence response output device 10010 and the large language model on the network (the large language model connected using the API) becomes unavailable. Specifically, the unavailability may be due to a communication unavailable state in the connection path from the artificial intelligence response output device 10010 to the Internet 19000. Or, the unavailability may be due to a communication unavailable state in the Internet 19000. Or, the unavailability may be due to a situation where the large language model on the network (the large language model connected using the API) itself cannot be connected to the Internet 19000. Also, in Example 1, as the "switching destination from the LLM (API-connected LLM) on the network", "local LLM" is shown. Specifically, this means performing a switching process to the response generation process by the local LLM processing unit 10028 included in the artificial intelligence response output device 10010. That is, in Example 1, even if the connection to the large language model on the network (the large language model connected using the API) becomes unavailable for some reason and the response generation process by the large language model on the network (the large language model connected using the API) cannot be used, it is switched to the response generation process by the local LLM processing unit 10028 included in the artificial intelligence response output device 10010. Thereby, even if there is a performance difference as a large language model, it is possible to continue the response generation process using the large language model.

[0235] Next, Example 2 in FIG. 5A will be described. Example 2 is obtained by changing the "switching destination from the LLM on the network (API-connected LLM)" in Example 1 from "local LLM" to "response template sentence DB (database)". Since the response generation process using the said "response template sentence DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, repeated description will be omitted. That is, in Example 2, if for some reason the connection to the large language model on the network (large language model connected using an API) becomes unavailable and the response generation process by the large language model on the network (large language model connected using an API) cannot be used, by switching to the response generation process using the response template sentence database, it becomes possible to generate a response by a simpler process and output the response to the user.

[0236] Next, Example 3 in FIG. 5A will be described. Example 3 is obtained by changing the "switching destination from the LLM on the network (API-connected LLM)" in Example 1 from "local LLM" to "non-response handling". The said "non-response handling" means that even when there is a user input requesting a response from the large language model via the touch panel, the microphone 1139, or the operation input unit 1107, no response to this input is generated, or even when there is a user input requesting a response from the large language model, no response to this is output. That is, in Example 3, if for some reason the connection to the large language model on the network (large language model connected using an API) becomes unavailable and the response generation process by the large language model on the network (large language model connected using an API) cannot be used, it becomes possible to simplify the handling in such a case.

[0237] Next, Example 4 in FIG. 5A will be described. Example 4 is an example of performing switching due to the response delay of the LLM on the network, as shown in the "Switching Overview". In Example 4, as the "state before switching the LLM (API-connected LLM) on the network", a state where a response from the LLM on the network can be obtained within a predetermined time is shown. Here, in Example 4, as the "switching occurrence condition", a case where a response from the LLM on the network cannot be obtained within a predetermined time and exceeds the predetermined time is shown. Also, in Example 4, as the "switching destination from the LLM (API-connected LLM) on the network", "local LLM" is shown. The "local LLM" of the switching destination is the same as in Example 1, so repeated explanations are omitted. That is, in Example 4, for some reason, the response from the LLM (large language model connected using an API) on the network exceeds the predetermined time, and even when the response generation process by the LLM (large language model connected using an API) on the network cannot be used smoothly, the response generation process by the local LLM processing unit 10028 included in the artificial intelligence response output device 10010 is switched to. Thereby, even if there is a performance difference as a large language model, it is possible to continue the response generation process using the large language model.

[0238] Next, Example 5 in FIG. 5A will be described. Example 2 is obtained by changing the "switching destination from the LLM (API-connected LLM) on the network" in Example 4 from "local LLM" to "response fixed-form sentence DB (database)". The response generation process by the "response fixed-form sentence DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, so repeated explanations are omitted. That is, in Example 5, for some reason, the response from the LLM (large language model connected using an API) on the network exceeds the predetermined time, and when the response generation process by the LLM (large language model connected using an API) on the network cannot be used smoothly, by switching to the response generation process using the response fixed-form sentence database, it becomes possible to generate a response by a simpler process and output the response to the user.

[0239] Next, Examples 6 to 9 in FIG. 5A will be described. Examples 6 to 9 are examples where switching is performed when the upper limit of API usage or usage fee is reached, as shown in the "Switching Outline". Here, as described in Example 2, the provider of the large language model often recovers the cost used for training the large language model from the user of the terminal as the usage fee of the API of the terminal. At that time, in the natural language model, the usage fee of the API is often billed in a form based on the number of processed units of words that divide a sentence, called tokens. Here, various billing methods and limiting methods can be considered for the API usage fee. As one example, an example can be considered where the user defines the upper limit of the amount of the large language model usage service that can be received in the normal state using the number of token processes.

[0240] In this case, until the user reaches the usage amount (or the corresponding usage fee), the user can receive the service of using the large language model at a predetermined API usage fee. When the upper limit of the usage amount (or the corresponding usage fee) is reached, certain restrictions may occur, such as the inability to receive the large language model usage service in the normal state (performance or frequency).

[0241] Examples 6 to 9 in FIG. 5A are examples of switching control of the response generation process by the control unit 1110 of the artificial intelligence response output device 10010 when such a restriction occurs in the large language model usage service. Specifically, in Example 6, the "state before switching the LLM (API-connected LLM) on the network" is a state where the API usage amount and the API usage fee are below the predetermined upper limit. This means that the usage amount of the LLM (API-connected LLM) on the network has not reached the predetermined upper limit. At this time, the user can use the LLM (API-connected LLM) on the network in the normal state.

[0242] Here, in Example 6, as the "switching occurrence condition", it is shown that when the usage volume of the API or the usage fee of the API reaches a predetermined upper limit. This means when the usage volume of the LLM (API-connected LLM) on the network reaches a predetermined upper limit. Also, in Example 6, as the "switching destination from the LLM (API-connected LLM) on the network", a second LLM on a network different from the LLM (which may be referred to as the first LLM) that was being used in the normal state is shown. Examples of the second LLM on the network include an LLM with a lower fee than the first LLM that was being used in the normal state. Since it becomes a service with a lower fee, it is conceivable that the performance of the second LLM is lower than the performance of the first LLM. Even in this case, there are sufficient advantages if the large language model can be used inexpensively even after the upper limit of the usage volume / usage fee of the first LLM is reached.

[0243] Next, Example 7 in FIG. 5A will be described. Example 7 is a modification where the "switching destination from the LLM (API-connected LLM) on the network" in Example 6 is changed from a second LLM on a network different from the LLM (which may be referred to as the first LLM) that was being used in the normal state to a "local LLM". In Example 7, even when the usage volume of the API or the usage fee of the API reaches a predetermined upper limit, that is, even when the usage volume of the LLM (API-connected LLM) on the network reaches a predetermined upper limit, by switching to a response generation process using the local LLM that is not restricted by the usage volume of the LLM on the network, the API usage volume, or the API usage fee, etc., it becomes possible to continue performing the response generation process using the large language model.

[0244] Next, Example 8 in FIG. 5A will be described. Example 8 is obtained by changing the "switching destination from the LLM on the network (API-connected LLM)" in Example 7 from "local LLM" to "response template DB (database)". Since the response generation process by the "response template DB (database)" is the same as the process described in FIG. 1C, FIG. 2L, or FIG. 4B, repeated description will be omitted. In Example 8, even when the usage amount of the API or the usage fee of the API reaches a predetermined upper limit, that is, even when the usage amount of the LLM (API-connected LLM) on the network reaches a predetermined upper limit, the response generation process using the response template database that is not restricted by the usage amount of the LLM on the network, the usage amount of the API, or the usage fee of the API, etc. is switched to. As a result, it becomes possible to generate a response by a simpler process and output the response to the user.

[0245] Next, Example 9 in FIG. 5A will be described. Example 9 is obtained by changing the "switching destination from the LLM on the network (API-connected LLM)" in Example 7 from "local LLM" to "non-response handling". The "non-response handling" means a response that does not generate a response to the user or does not output a response to the user. In Example 9, when the usage amount of the API or the usage fee of the API reaches a predetermined upper limit, that is, when the usage amount of the LLM (API-connected LLM) on the network reaches a predetermined upper limit, it becomes possible to simplify the response when the response generation process by the large language model on the network (large language model connected using the API) cannot be used.

[0246] According to the switching control of the response generation process of the artificial intelligence response output device 10010 shown in Examples 1 to 9 of FIG. 5A described above, even in a situation where the response generation process by the LLM on the network (large language model connected using the API) cannot be used as usual, more suitable switching or response can be performed according to each situation.

[0247] Note that the switching control for Examples 1 to 9 in FIG. 5A may be performed by combining a plurality of examples. For example, the switching control for Examples 1 to 3 may be combined with any one of the controls for Examples 4 to 9 respectively. Similarly, the control for Example 4 or Example 5 may be combined with any one of the controls for Examples 1 to 3 or Examples 6 to 9 respectively. Similarly, the control for Examples 6 to 9 may be combined with any one of the controls for Examples 1 to 5 respectively.

[0248] Next, with reference to FIGS. 5B to 5D, an example of a display example of an AI assistant or a character will be described when the artificial intelligence response output device 10010 of Example 5 is configured as an AI assistant device or a character conversation device.

[0249] First, FIG. 5B is a display example of an AI assistant or a character in the artificial intelligence response output device 10010 when performing the switching control of Example 3 in FIG. 5A. In the example of FIG. 5B, the display state of the AI assistant or the character is changed according to whether the network connection state of the artificial intelligence response output device 10010 is network-connectable or network-unavailable. Since the states where the artificial intelligence response output device 10010 is network-connectable and network-unavailable are as described in FIG. 5A, repeated explanations will not be given.

[0250] In the example of FIG. 5B, when the artificial intelligence response output device 10010 is network-connectable, it displays the AI assistant or character in the normal awake state. When it is not network-connectable, it displays the AI assistant or character in the "sleeping" state. In the switching control of Example 3 in FIG. 5A, when the artificial intelligence response output device 10010 is not network-connectable, it does not generate a response or output a response even if there is an input of an instruction sentence from the user. At this time, the user may feel uncomfortable when the AI assistant or character displayed by the artificial intelligence response output device 10010 is in the normal awake state. However, if the AI assistant or character displayed by the artificial intelligence response output device 10010 is displayed in the sleeping state, the user can understand that "the reason why the AI assistant or character does not respond is that it is sleeping", and it is possible to further reduce the discomfort felt by the user.

[0251] In the case of FIG. 5B(2), before the user input for requesting a response from the large language model is made via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 1001, it is desirable for the user to understand that "the reason why the AI assistant or character does not respond is that it is sleeping". Therefore, in the case of (2) in FIG. 5B where it is not network-connectable, the start timing of the state where the AI assistant or character is displayed in the "sleeping" state is desirably immediately after the control unit 1110 of the artificial intelligence response output device 10010 determines that it is not network-connectable, which is before the user input for requesting a response from the large language model.

[0252] Next, as another display example, the display example in FIG. 5C will be described. The display example in FIG. 5C is an example of changing the display state of the AI assistant or character according to the state of "switching destination from LLM on the network (API-connected LLM)" in the table in the switching control of FIG. 5A. Specifically, in FIG. 5C, (1) when the artificial intelligence response output device 10010 can be connected to a large language model on the network (a large language model connected using an API) and can utilize the response generation process by the large language model on the network (referred to as the normal state in this figure), the display example of the AI assistant or character, (2) the artificial intelligence response output device 10010 switches to a response generation process by an LLM with lower performance than the large language model on the network (a large language model connected using an API) or a response fixed-form sentence database, and the display example of the AI assistant or character in this state, and (3) the artificial intelligence response output device 10010 switches to the non-response handling described in FIG. 5A, and the display example of the AI assistant or character in this state are shown.

[0253] In the example of FIG. 5C, for example, (1) when the artificial intelligence response output device 10010 is in the "normal state", the artificial intelligence response output device 10010 displays the AI assistant or character in a particularly problem-free state. Note that the "normal state" in FIG. 5C may be considered as a state other than the states in (2) and (3). Also, for example, (2) when the artificial intelligence response output device 10010 switches to a response generation process by an LLM with lower performance than the large language model on the network (a large language model connected using an API) or a response fixed-form sentence database, the artificial intelligence response output device 10010 displays the AI assistant or character in a "sleepy" state. Note that "displaying the AI assistant or character in a'sleepy' state" may also be expressed as "display indicating that the AI assistant or character is feeling sleepy".

[0254] The response generation process in (2) has lower performance than the response generation process by the large language model on the network in the normal state of (1) (the large language model connected using the API). Therefore, by displaying the AI assistant or character in a "sleepy" state, it is possible to implicitly convey to the user that the response performance of the AI assistant or character is low. As a result, it becomes possible to further reduce the discomfort felt by the user towards the low-performance response. Note that the switching condition for the artificial intelligence response output device 10010 to switch to the response generation process by an LLM or response template database that is less performant than the large language model on the network (the large language model connected using the API) is as described in Fig. 5A, so repeated explanation is omitted.

[0255] Also, in the case of Fig. 5C(2), before performing a user input to request a response from the large language model from the user via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 10010, it is desirable to implicitly convey to the user that the response performance of the AI assistant or character is low. Therefore, the start timing of the state of displaying the AI assistant or character in the "sleepy" state in Fig. 5C(2) is preferably immediately after the point in time when the artificial intelligence response output device 10010 switches to the response generation process by an LLM or response template database that is less performant than the large language model on the network (the large language model connected using the API) before the user input to request a response from the large language model from the user.

[0256] Also, for example, when the artificial intelligence response output device 10010 switches to the non-response handling described in FIG. 5A, the artificial intelligence response output device 10010 displays the AI assistant or character in a "sleeping" state. As also described in FIG. 5B, by displaying the AI assistant or character displayed by the artificial intelligence response output device 10010 in a "sleeping" state, the user can understand that "the reason the AI assistant or character does not respond is because it is sleeping.", and it becomes possible to further reduce the sense of discomfort felt by the user. Regarding the switching occurrence conditions for the artificial intelligence response output device 10010 to switch to the non-response handling described in FIG. 5A, since it is as described in Example 3 or Example 9 of FIG. 5A, repeated explanations are omitted. In the case of FIG. 5C(3), before the user input for requesting a response from the large language model via the touch panel, microphone 1139, or operation input unit 1107 of the artificial intelligence response output device 10010, it is desirable for the user to understand that "the reason the AI assistant or character does not respond is because it is sleeping.". Therefore, the start timing of the state of displaying the AI assistant or character in FIG. 5C(3) in a "sleeping" state is preferably immediately after the point in time when the artificial intelligence response output device 10010 switches to the non-response handling described in FIG. 5A, which is before the user input for requesting a response from the large language model.

[0257] In the display example of FIG. 5C, the artificial intelligence response output device 10010 makes a display that implicitly reflects the technical explanation of the state of the artificial intelligence response output device 10010 regarding the response generation process as a change in the state of the AI assistant or character, without directly explaining it to the user. Thereby, it is possible to further reduce the sense of discomfort felt by the user compared to the case of directly explaining the technical explanation of the state of the artificial intelligence response output device 10010 regarding the response generation process to the user. Also, even though the state of the artificial intelligence response output device 10010 regarding the response generation process has changed, it is possible to further reduce the sense of discomfort felt by the user compared to the case where the display state of the AI assistant or character remains the same as the normal state.

[0258] However, depending on the user, there may be cases where they want to know more precisely the description of the technical state in each state. Therefore, an example of a display for dealing with such users will be described with reference to FIG. 5D. Among the rows of the table shown in FIG. 5D, the rows of the device state and the display state description are exactly the same as those in FIG. 5C, so repeated explanations will be omitted. Also, the display examples of the AI assistant or character shown in the row of the display examples of the AI assistant or character are almost the same as those in FIG. 5C, but differ in that a question mark (?) is displayed in the display example. The said question mark (?) is a mark that the user operates when requesting an explanation from the artificial intelligence response output device 10010, and may be referred to as a help mark.

[0259] In the example of FIG. 5D, when the user selects the said question mark (?) by a user operation via a touch panel or the like provided in the operation input unit 1107 or the display unit 10011 of FIG. 1B, the display of the AI assistant or character of the artificial intelligence response output device 10010 is changed to the display example shown in the row of the display example after the user operation. Specifically, regardless of whether the device state is in any of the states (1), (2), or (3), an explanation of the technical state in each state is displayed. For example, in the example of FIG. 5D, if the state of the device is the normal state (1), a display may be made to explain that it is a normal state with no particular technical limitation, such as "It is in the normal state." Also, if the state of the device is the state of using the low-performance LLM or the response template database (2), a display may be made to technically explain that it is in the low-performance state, such as "It is in the low-performance mode." The said display may also be considered as a display explaining the factor that the display of the AI assistant or character is shown in the "sleepy" state.

[0260] In this case, a more technically detailed explanation may be provided. Specifically, a display such as "Low-performance LLM usage mode" or "Template response mode" may be presented. Also, if the device state is (3) non-response handling, a display stating "Network connection unavailable" and technically explaining the factor for switching to non-response handling may be given. If the factor for switching to non-response handling is that the response from the LLM (large language model connected using an API) on the network has exceeded a predetermined time, a display such as "Response from LLM is delayed" may be shown. Further, if the factor for switching to non-response handling is that the usage volume of the LLM on the network, API usage volume, or API usage fee has reached the limit, a display such as "Usage volume of LLM has reached the limit", "Usage volume of API has reached the limit", or "API usage fee has reached a predetermined amount" may be presented. These displays may also be considered as explanations for the factor of the display of the AI assistant or character being in a "sleeping" state.

[0261] According to the display example of FIG. 5D described above, even if there are technical constraints in the response generation process in the artificial intelligence response output device 10010, first, without directly explaining to the user, by changing the display state of the AI assistant or character to implicitly indicate the device state, the discomfort felt by the user can be further reduced. This display is more suitable for users who do not require a technical explanation. Additionally, by displaying an operation mark for explaining the technical state, for the user who operates the mark, a display is provided to technically explain the state of the response generation process (normal state or state with technical constraints) in the artificial intelligence response output device 10010. This enables a more suitable display for users who want to accurately know the technical state.

[0262] In the examples of FIGS. 5B, 5C, and 5D, an example of the display state of the AI assistant or character when in the "non-response handling" state is shown as the "sleeping" state. However, this is just an example, and the embodiments of this example are not limited to this. Instead of the "sleeping" state, other display states that implicitly indicate a situation where a response cannot be made, such as "on break", may be used. Also, in the examples of FIGS. 5C and 5D, an example of the display state of the AI assistant or character when in a state of using a low-performance LLM or a response template database is shown as the "sleepy" state. However, this is just an example, and the embodiments of this example are not limited to this. Other display states that implicitly indicate a low response performance of the AI assistant or character, such as "hungry", may be used.

[0263] According to the artificial intelligence response output device and the artificial intelligence response output system according to Embodiment 5 described above, according to the connection state between the large language model on the network and the artificial intelligence response output device, the response delay state from the large language model on the network, or the usage amount of the large language model on the network, etc., it is possible to more preferably switch the response generation process used by the artificial intelligence response output device. Also, when configuring the artificial intelligence response output device according to Embodiment 5 as an AI assistant device or a character conversation device, it is possible to perform a display that causes less discomfort to the user.

[0264] <Example 6> Next, Embodiment 6 of the present invention is an improvement of the artificial intelligence response output device 10010 or the artificial intelligence response output system described in each figure of Embodiments 1 to 5. Specifically, it is an example of more preferably combining the response generation process by the large language model on the network or the response generation process by the local large language model (such as the local LLM processing unit 10028) provided in the artificial intelligence response output device 10010 with the response generation process by the response template database to generate a response output. In this example, the differences from these examples will be described, and repeated descriptions of the same configurations as these examples will be omitted.

[0265] Similar to the above embodiments, the artificial intelligence response output device 10010 may also be referred to as an artificial intelligence response output device, a character conversation device, an AI assistant device, an AI assistant display device, or an artificial intelligence interface device. A system including the artificial intelligence response output device 10010 and a large language model server may also be referred to as an artificial intelligence response output system, a character conversation system, an AI assistant system, an AI assistant display system, or an artificial intelligence interface system.

[0266] With reference to FIG. 6, an example of the response generation process in the artificial intelligence response output device 10010 according to Embodiment 6 of the present invention will be described. FIG. 6 shows an example of a flowchart of the response generation process in the artificial intelligence response output device 10010 according to Embodiment 6 of the present invention according to Embodiment 6. Specifically, a time axis on which time progresses from top to bottom, a processing flow, and an example of response output are shown. The output of the response shown in the response output example may be performed via the display by the display unit 10011 of the artificial intelligence response output device 10010 or the voice output by the voice output unit 1140.

[0267] In the example of FIG. 6, first, at time t0, there is a user input requesting a response from the large language model from the user via the touch panel, the microphone 1139, or the operation input unit 1107 of the artificial intelligence response output device 10010, and the control unit 1110 of the artificial intelligence response output device 10010 acquires the user input (step 600). Next, at time t1, the control unit 1110 starts preparing for response output using the response template sentence database stored in the storage unit 1170 and starts response output using the response template sentence database (step 601). In the example of FIG. 6, at time t2, the response output using the response template sentence database has started, and as shown in the figure, the template sentence response is being output and has not been completed. "Good morning" in the figure indicates the output up to the middle of the sentence that continues as "Good morning. …".

[0268] At time t3, before the response output using the response template sentence database is completed, the control unit 1110 generates an instruction sentence based on the user input acquired in step 600, and transmits the generated instruction sentence to a large language model on the network or a local large language model (such as the local LLM processing unit 10028) provided in the artificial intelligence response output device 10010, and starts a request for a response from the large language model (step 602). Further, at time t4, before the response output using the response template sentence database is completed, the control unit 1110 starts acquiring a response from the large language model (step 603).

[0269] At time t5, a response output example in which the response output using the response template sentence database is completed is shown. For example, FIG. 6 shows an example in which the display of the response output "Good morning. Today is the 〇th day of 〇 month." is completed at time t5 using the template sentence stored in the response template sentence database and the date information stored in the memory. Here, at time t4 before the response output using the response template sentence database is completed, the control unit 1110 has already started acquiring a response from the large language model. Therefore, at time t6 following time t5 when the response output display using the response template sentence database is completed, the control unit 1110 starts the response output from the large language model following the response output using the response template sentence database (step 604). After that, at time t7, a response from the large language model is output following the response output using the response template sentence database. When the response output from the large language model is completed, the response output according to the processing flow shown in FIG. 6 is completed (step 605).

[0270] Next, the effects of the processing flow shown in FIG. 6 of the present invention will be described. Processing of large language models requires a lot of computing resources. Generally, even if inference, which requires fewer computing resources than learning, is processed using a GPU (Graphics Processing Unit), it may take several seconds to more than ten seconds from when the control unit starts a response request to the large language model until a response can be obtained from the large language model. This period corresponds to the period from time t3 to time t4 shown in FIG. 6. Also, from time t0 when there was user input until time t4, since the control unit 1110 has not obtained a response output from the large language model, it cannot output a response from the large language model to the user.

[0271] Therefore, in a processing flow where there is no start of preparation for response output using the response fixed sentence database in step 601 shown in FIG. 6 and no start of response output using the response fixed sentence database, the user may have to continue waiting without a response from the artificial intelligence response output device 10010 for a period exceeding several seconds to more than ten seconds from time t0 when the user input was made until time t4. For example, when the artificial intelligence response output device 10010 is configured as an AI assistant device or a character conversation device, etc., this waiting time may give the user a sense of discomfort.

[0272] In contrast, in the processing flow according to Embodiment 6 of the present invention shown in FIG. 6, the control unit 1110 starts processing of response output using a response fixed sentence database that requires fewer computing resources than processing of the large language model before the start of obtaining a response from the large language model. As a result, the user does not have to continue waiting without a response from the artificial intelligence response output device 10010 from time t0 until time t4. For the user, whether it is a response output using the response fixed sentence database or a response output from the large language model, it is still a response from the artificial intelligence response output device 10010.

[0273] Therefore, in the processing flow shown in FIG. 6, by providing step 601 before step 603, the response of the artificial intelligence response output device 10010 to the user can be pseudo-accelerated. Thereby, it is possible to further reduce the discomfort of the user caused by the length of the waiting time. Also, in step 604, by outputting the response from the large language model following the response using the response fixed sentence database, the user can be made to recognize as if these outputs are a series of more natural outputs.

[0274] According to the artificial intelligence response output device and the artificial intelligence response output system according to Example 6 described above, it is possible to shorten the response waiting time from the artificial intelligence response output device for the user, and to further reduce the discomfort felt by the user.

[0275] <Example 7> Example 7 of the present invention is an improvement of the artificial intelligence response output device 10010 or the artificial intelligence response output system described in Example 1 and the like. In this example, the differences from Example 1 will be described, and repeated descriptions of the same configurations as those of these examples will be omitted.

[0276] The artificial intelligence response output system according to Example 7 is an example in which a conversion process of converting the user's question sentence as needed is performed when transmitting an instruction sentence from the artificial intelligence response output device 10010 to the large language model. More specifically, when creating an instruction sentence (prompt) for the large language model artificial intelligence from the user's question sentence, a conversion process of converting the user's question sentence as needed is performed so that a more appropriate answer can be obtained for the user's question sentence.

[0277] With reference to FIG. 7, an example of the response output system according to Embodiment 7 of the present invention will be described. As shown in FIG. 7, the configuration of the response output system according to Embodiment 7 is basically the same as that of Embodiment 1. Also in Embodiment 7, the artificial intelligence response output device 10010 is communicably connected via the Internet 19000 to a large language model server 19001 or a large language model server 20001 equipped with a large language model, and a second server 19002 different from these servers. However, it is different from Embodiment 1 in that the second server 19002 includes a check database (check DB) 19010.

[0278] The check DB 19010 will be described in detail later. Also in Embodiment 7, a configuration in which the large language model server 19001 includes a large language model (LLM) will be described. In the following description, the large language model included in the large language model server 19001 may be simply referred to as a large language model or an LLM. Incidentally, the large language model may be a large language model on the network or a local large language model (such as the local LLM processing unit 10028) provided in the artificial intelligence response output device 10010.

[0279] By the way, when a user tries to ask a question to a large language model which is an artificial intelligence and obtain its response (answer), for example, the following procedure is performed. First, the user inputs a question to the large language model to the artificial intelligence response output device 10010. As an example, the user operates the operation input unit 1107 which is an input unit to input a question sentence which is text information in natural language. Also, the input of the question to the large language model may be operated by the operation input unit 1107 or may be voice input. The artificial intelligence response output device 10010 generates a response instruction sentence (prompt) for the large language model based on this question sentence and transmits it to the large language model server 19001 equipped with the large language model. When the large language model generates a response to the response instruction sentence (for example, an answer sentence to the user's question), the generated response is transmitted from the large language model server 19001 to the artificial intelligence response output device 10010.

[0280] When the artificial intelligence response output device 10010 receives this response, it performs processing on the user for the response of the large language model (LLM). As an example of the processing for the response of the large language model, the artificial intelligence response output device 10010 causes the display unit 10011 to display the answer sentence (also referred to as a response sentence) generated by the large language model. By such a procedure, the user can obtain an answer from the large language model to the question.

[0281] However, the user does not necessarily obtain the answer expected from the large language model. For example, depending on the way the question sentence is written, the content of the question may be interpreted in a meaning different from the user's intention, and there is a possibility that the answer desired by the user cannot be obtained from the large language model.

[0282] Therefore, in the response output system according to Example 7, when a user inputs a question sentence to the artificial intelligence response output device 10010, the artificial intelligence response output device 10010 executes a conversion process for the question sentence as necessary, and generates a response instruction sentence based on the question sentence after the conversion process (hereinafter sometimes referred to as the processed question sentence). As a result, the question sentence can be more appropriately interpreted by the large language model, and it becomes easier to obtain an answer desired by the user from the large language model.

[0283] Note that the method for the user to input a question sentence to the artificial intelligence response output device 10010 may be a method of converting a question input by voice using the microphone 1139 into text, in addition to character input using the operation input unit 1107 as an input unit. Also, the conversion process of the question sentence, the generation process of the instruction sentence, the output process of the answer from the large language model to the display unit 10011, etc. are executed by, for example, the control unit 1110. The control unit 1110 in this embodiment generates a response instruction sentence for the large language model based on the question sentence, and acquires a response sentence generated by the large language model for the response instruction sentence. Specifically, the control unit 1110 executes a conversion process for the question sentence, and generates a response instruction sentence based on the question sentence after the conversion process.

[0284] Hereinafter, the response output processing flow in the response output system according to Example 7, mainly the processing flow when obtaining a response from the large language model to the user's question, will be described in detail. FIGS. 8, 9, and 11 are diagrams showing an example of the processing flow of the response output system according to Example 7. FIG. 10 is a diagram showing an example of the check database stored in the second server according to Example 7. FIG. 12 is a diagram showing an example of the database collation priority table. Further, FIGS. 13 and 14 are diagrams showing an example of a modified example of the processing flow of the response output system according to Example 7. In the diagrams showing the processing flow of the response output system, the same step is denoted by the same reference numeral, and duplicate explanations for each step may be omitted.

[0285] First, referring to FIG. 8, the basic processing flow in the response output system according to Example 7 will be described. As shown in FIG. 8, first, the user operates, for example, the operation input unit 1107 to input a question sentence to the large language model. In step S810, when the control unit 1110 receives the user input (question sentence), in step S820, a check process is executed to determine whether conversion processing of the question sentence input by the user is necessary. Also in step S820, the control unit 1110 executes the conversion processing of the question sentence as necessary based on the result of this check process. These check process and conversion process will be described in detail later.

[0286] Next, proceeding to step S830, based on the question sentence input by the user or the question sentence after conversion processing (hereinafter referred to as the converted question sentence), an instruction sentence (prompt) for the large language model (LLM) provided in the large language model server 19001 is generated. In this example, a response instruction sentence for instructing the creation of an answer sentence as a response to the user's question is generated. Next, this response instruction sentence is transmitted to the large language model (step S840). Note that the above response instruction sentence, the conversion instruction sentence, the summary instruction sentence, and the additional instruction sentence described later are collectively referred to as the instruction sentence.

[0287] Next, when the large language model (LLM) receives the instruction sentence transmitted by the artificial intelligence response output device 10010, for example, the response instruction sentence (step S850), it generates a response based on this instruction sentence (step S860). In the example shown in FIG. 8, since the instruction sentence is the response instruction sentence, the large language model creates an answer sentence to the user's question as the response generation in step S860. Then, in step S870, the answer sentence created by the large language model is transmitted to the artificial intelligence response output device 10010.

[0288] When the control unit 1110 of the artificial intelligence response output device 10010 receives a response (e.g., an answer sentence) created by a large language model (step S880), in step S890, it executes processing for this response of the large language model. As an example of the processing for the response of the large language model, the control unit 1110 causes the display unit 10011 to display an answer sentence to the user's question.

[0289] As described above, in the response output system according to the seventh embodiment, when performing response output processing, a check process is performed on the question sentence input by the user, and a conversion process of the question sentence is performed as necessary. As a result, it becomes easier to obtain an answer intended by the user from the large language model for the user's question.

[0290] Next, the check process and conversion process of the question sentence in step S820 will be described in more detail. For example, as shown in FIG. 9, when the control unit 1110 receives user input in step S810, then, as the process of step S820, the control unit 1110 first executes a check process on the question sentence input by the user and determines whether a conversion process of the question sentence is necessary (step S8210).

[0291] As a result of this check process, if it is determined that the conversion process of the question sentence is not necessary (step S8210: No), the process proceeds to step S830, and an instruction sentence is generated based on the question sentence. On the other hand, if it is determined that the conversion process of the question sentence is necessary (step S8210: Yes), the process proceeds to step S8220, and the conversion process of the question sentence is executed. In this example, as the conversion process of the question sentence, a process of requesting the user to rephrase the question sentence is performed (step S2220).

[0292] To explain the check process in step S820 in more detail, the control unit 1110 executes, as the check process, a process of checking whether the question sentence corresponds to a predetermined check item. As an example, the control unit 1110 performs a process of checking whether the question sentence contains a specific expression set in advance. In other words, the fact that the question sentence contains the above specific expression is one of the above check items. Note that the content of the check item is not particularly limited as long as it can determine the necessity of the conversion process of the question sentence.

[0293] Here, the above specific expression refers to, in this example, words, grammar, etc. registered in advance in the check database (check DB) 19010 of the second server 19002. That is, a list of specific expressions is stored in the check DB 19010 of the second server 19002. As an example, as shown in FIG. 10, the specific expressions stored in the check DB 19010 include specific word groups such as "technical terms", "abbreviations", "homophones / homographs", and specific grammar such as "ambiguous sentence expressions" and "dialects". Note that regarding the specific word groups registered in the check DB 19010, in this example, it is assumed that their meanings or explanations are also stored.

[0294] In addition, the specific expressions stored in the check DB 19010 are classified into a plurality of groups including groups of "ambiguous sentence expressions", "abbreviations", "technical terms", "homophones / homographs", and "dialects", as shown in FIG. 10. Further, each group is set in a plurality of hierarchies. That is, the groups at each hierarchy are classified into a plurality of groups at the next hierarchy. For example, the group of "ambiguous sentence expressions" is classified into the groups of "ambiguous sentence expression classification 1" and "ambiguous sentence expression classification 2", and these groups of "ambiguous sentence expression classification 1" and "ambiguous sentence expression classification 2" are further classified into a plurality of groups. Similarly, the group of "abbreviations" is classified into the groups of "abbreviation classification 1" and "abbreviation classification 2", and the groups of "abbreviation classification 1" and "abbreviation classification 2" are classified into a plurality of groups.

[0295] As described above, in this example, a group of specific expressions is set in three levels. However, the number of levels of the group is not particularly limited and may be determined as appropriate. Further, the number of groups set as the same level is not particularly limited and may be determined as appropriate.

[0296] Then, the control unit 1110 executes a check process for the question sentence using such a check DB 19010. For example, as shown in FIG. 11, as the check process for the question sentence in step S820, the control unit 1110 first performs a database collation process for the question sentence in step S8211. In this collation process, the words and grammar in the question sentence input by the user are collated with the specific expressions registered in the check DB 19010.

[0297] Next, in step S8212, it is determined whether there are any words or grammar in the question sentence that match the specific expressions registered in the check DB 19010. In other words, the control unit 1110 performs a process of checking whether the words, grammar, etc. classified into the plurality of groups illustrated in FIG. 10 are included in the question sentence, that is, whether the specific expressions registered in the check DB 19010 are included in the question sentence.

[0298] When there are no words or grammar in the question sentence in the check DB 19010 (step S8212: No), it is determined that the question sentence does not correspond to the above check items, and the process proceeds to step S830 to generate an instruction sentence. On the other hand, when there are words or grammar in the question sentence in the check DB 19010 (step S8212: Yes), it is determined that the question sentence corresponds to the above check items, and the process proceeds to step S8220, and as a conversion process for the question sentence, a process of requesting the user to rephrase the question sentence is performed. The control unit 1110 causes the display unit 10011 to display information to the effect of requesting the user to rephrase the question sentence, for example.

[0299] In this case, it is preferable to also display the reason for requesting the user to rephrase the question. For example, it is preferable to also display that the question contains a specific expression and the words / grammar corresponding to the specific expression included in the question. Note that the method of requesting the user to rephrase the question is not particularly limited, and may be performed by voice or the like.

[0300] Also, for the rephrasing request as this conversion process, the question rephrased by the user corresponds to the "question after conversion". In the example shown in FIG. 9, when the user rephrases the question in response to the request in step S8220 and the control unit 1110 receives the rephrased question (question after conversion) (step S810), a check process for the question after conversion is performed to determine whether further conversion processing of the question after conversion is necessary (step S8210). Then, when it is determined that further conversion processing of the question after conversion is necessary, the user is again requested to rephrase the question after conversion (step S8220). Also, when it is determined that further conversion processing of the question after conversion is not necessary (step S8210: No), the process proceeds to step S830, and an instruction sentence is generated based on the question after conversion. The subsequent processing flow is the same as the example in FIG. 8, so the description is omitted.

[0301] By the way, in the artificial intelligence response output device 10010, a priority for collating the question with the specific expressions registered in the check DB 19010 is set in advance. In the artificial intelligence response output device 10010, for example, as shown in FIG. 12, a DB collation priority table for setting the collation priority between the question and the check DB 19010 is stored in advance. In this example, collation priorities 1 to 5 are set for the five groups of "ambiguous sentence expressions", "abbreviations", "technical terms", "homophones / homographs", and "dialects" in the check DB 19010.

[0302] In this case, in steps S8211 and S8212 shown in FIG. 11, the group set to priority 1 is processed (collated) first, and the group set to priority 5 is processed last. In the example shown in FIG. 12, in the case of the default setting, the "ambiguous sentence expression" in the question sentence is collated first, and the "dialect" in the question sentence is collated last.

[0303] Also, the priority of the group to be collated can be arbitrarily set, for example, according to the habits and preferences of the user's way of writing the question sentence. Furthermore, the types and numbers of the groups to be collated can also be arbitrarily set. Note that the group set as the group to be collated is referred to as the "set group".

[0304] In the example shown in FIG. 12, "User 1 personal settings" and "User 2 personal settings" take five groups of "ambiguous sentence expression", "abbreviation", "technical term", "homophone / homograph", and "dialect" as the "set group", and this is an example where the user arbitrarily sets the priority order of these five set groups. Also, "User 3 personal settings" takes three groups of "dialect", "ambiguous sentence expression", and "homophone / homograph" as the "set group", and this is an example where the user arbitrarily sets the priority order of the three set groups.

[0305] By appropriately setting the priority etc. of the group to be collated in this way, the conversion process of the question sentence can be efficiently executed. For example, by setting the priority of the group of specific expressions that are likely to be used as the question sentence to a higher level, it becomes easier to perform the collation between the question sentence and the check DB 19010 and the conversion process of the question sentence in parallel. Thereby, it is possible to shorten the time until an answer from the large language model for the question sentence is obtained.

[0306] The above described an example of the processing flow of the response output system of Example 7. However, the processing flow when obtaining the response of the large language model to the user's question is not limited to the above example. In the above example, as the conversion process of the question text, the case of performing the process of requesting the user to rephrase the question text was described. However, the conversion process of the question text is not limited to this process.

[0307] Next, with reference to FIGS. 13 to 15, another example of the conversion process of the question text will be described. The processing flow of FIG. 13 shows an example in which the control unit 1110 of the artificial intelligence response output device 10010 executes the conversion process of the question text. The control unit 1110 may also be referred to as a processor.

[0308] As shown in FIG. 13, in step S8212, when the words and grammar in the question text are in the check DB 19010 (step S8212: Yes), the process proceeds to step S8213, and as the conversion process of the question text, the control unit 1110 performs the conversion of the question text. For example, the control unit 1110 replaces a specific expression used in the question text with another expression (words and grammar), and converts it into a sentence that does not use the specific expression. In addition, in step S8212, when the words and grammar in the question text are not in the check DB 19010 (step S8212: No), the process proceeds to step S830, and as in the above example, an instruction sentence is generated based on the question text input by the user.

[0309] Also, in step S8212, when the words and grammar in the question text are in the check DB 19010 (step S8212: Yes), when the conversion of the question text by the control unit 1110 is performed in step S8213, then in step S8214, a process of asking the user whether the converted question text is as intended by the original question text is performed. For example, the confirmation items for confirming whether the converted question text is as intended by the original question text are displayed on the display unit 10011, and the user's answer to these confirmation items is requested.

[0310] Next, in step S8215, it is checked whether the converted question sentence matches the intention (user's intention) of the original question sentence. For example, based on the user's answer to the confirmation item displayed on the display unit 10011, it is determined whether the converted question sentence matches the user's intention. Specifically, a confirmation item "Does the converted question sentence match the intention of the original question sentence?" and a selection option of "Yes / No" as the user's answer to it are displayed. If "Yes" is selected, it is determined that the converted question sentence matches the user's intention, and if "No" is selected, it is determined that it does not match the user's intention.

[0311] Here, if the user answers that the converted question sentence matches the user's intention (step S8215: Yes), the process proceeds to step S830, where the control unit 1110 generates an instruction sentence based on the converted question sentence (converted question sentence after conversion), and transmits the generated instruction sentence to the large language model (step S840). On the other hand, if the user answers that the converted question sentence does not match the user's intention (step S8215: No), the process proceeds to step S8220, and further question sentence conversion processing is executed. In the example shown in FIG. 13, similar to the example shown in FIG. 9 and the like, as the question sentence conversion processing in step S8220, a process of requesting the user to rephrase the question sentence is performed.

[0312] Note that the question sentence rephrased by the user in response to this request also corresponds to the "converted question sentence after conversion" and is handled in the same manner as the above example. Also, in the example shown in FIG. 13, since the processing flow after the large language model receives the instruction sentence in step S850 is the same as the example shown in FIG. 8 and the like, the illustration is omitted. Also, step S820 is a system process, so other processes may be used. Furthermore, the processes of steps S8214 to S8220 do not necessarily have to be performed. For example, when the question sentence is converted in step S8213, the process may proceed to step S830, and a response instruction sentence may be created based on the converted question sentence after conversion.

[0313] The processing flows of FIGS. 14 and 15 show an example in which the large language model (LLM) of the large language model server 19001 performs the conversion process of the question text. In the example shown in FIG. 14, in step S8210, when the control unit 1110 determines that the conversion process of the question text is not necessary (step S8210: No), as in the above example, a response instruction text is generated based on the question text in step S830. Then, the generated instruction text (response instruction text) is transmitted to the large language model (step S840). When the large language model receives this response instruction text (step S850), it then proceeds to step S860 and generates a response to the response instruction text. Specifically, it generates an answer to the question.

[0314] Thereafter, as a response of the large language model to the instruction text (response instruction text), an answer to the question is transmitted to the control unit 1110 (step S870). When the control unit 1110 receives the response of the large language model to the response instruction text (step S880), it then performs processes such as presenting to the user the response of the large language model, in other words, the answer of the large language model to the question input by the user, in step S890. The user can confirm the answer of the large language model and, if satisfied with the answer, end the question to the large language model, input another new question, or, if not satisfied with the answer, input an additional question.

[0315] On the other hand, when it is determined that the conversion process of the question text is necessary (step S8210: Yes), it proceeds to step S831 and generates a conversion instruction text for instructing the large language model to perform the conversion process of the question text. Then, the generated instruction text (conversion instruction text) is transmitted to the large language model (step S840).

[0316] When the large language model receives this conversion instruction (step S850), it then proceeds to step S861 to generate a response to the conversion instruction. Specifically, as a conversion process for the question sentence created by the user, the question sentence is rephrased. For example, if the question sentence contains a specific expression as described above, the question sentence is converted into a sentence that does not use the specific expression. Of course, the content of the conversion process of the question sentence is not particularly limited. Thereafter, as a response of the large language model to the instruction (conversion instruction), the question sentence (converted question sentence) after the conversion process is transmitted to the control unit 1110.

[0317] When the control unit 1110 receives the converted question sentence from the large language model (LLM) as a response to the conversion instruction (step S880), it then, in step S900, checks with the user whether the question sentence rephrased by the large language model matches the original question sentence. In other words, it checks with the user whether the converted question sentence is as intended by the original question sentence.

[0318] Here, if an answer is received from the user that the converted question sentence matches the original question sentence (step S900: Yes), it proceeds to step S830, generates an instruction (response instruction) based on the converted question sentence converted by the large language model, and transmits the generated response instruction to the large language model (step S840).

[0319] On the other hand, if an answer is received from the user that the converted question sentence does not match the user's intention (step S900: No), it proceeds to step S910 and performs a process of requesting the user to rephrase the question sentence.

[0320] It should be noted that step S900 and step S910 in the processing flow of FIG. 14 can also be said to correspond to step S8214 to step S8220 in the processing flow of FIG. 13. Also, in the processing flow of FIG. 14, the processing from step S831 through step S861 and step S900 to step S910 corresponds to the conversion process of the question sentence.

[0321] Also, in each of the above-described processing flows of Example 7, when a question sentence is input by the user, a check process for the input question sentence is performed, and based on the result of the check process, a conversion process for the question sentence is performed as necessary. However, the check process for the question sentence does not necessarily have to be performed. That is, when a question sentence is input by the user, the conversion process for the question sentence may always be performed.

[0322] Specifically, in the example shown in FIG. 14, when a question sentence is input by the user in step S810, then a check process for the question sentence is performed in step S8210, and when conversion processing is necessary (step S8210: Yes), a conversion instruction sentence is generated (step S831). In contrast, in the processing flow of FIG. 15, when the control unit 1110 receives a question sentence input by the user (step S810), it creates a conversion instruction sentence in step S831 without performing a check process for the question sentence. Note that also in the processing flow of FIG. 15, the processing from step S831 through step S861 and step S900 to step S910 corresponds to the conversion process for the question sentence.

[0323] Here, the processing flow shown in FIG. 15 will be further described with reference to FIGS. 16A and 16B. For example, assume that the user inputs a question sentence "Why can't I run as fast as my older brother?" as shown in FIG. 16A (step S810). In this case, in step S830, the control unit 1110 generates a conversion instruction sentence in the form of a question sentence + a system instruction sentence as an example. The system instruction sentence is what instructs the response content (processing content) to the large language model, and in the example of FIG. 16A, it instructs to paraphrase the question sentence.

[0324] Example 1 is an example where the question sentence and the system instruction sentence in the conversion instruction sentence are each described in natural language. In the question sentence area of the conversion instruction sentence, "Why can't I run as fast as my older brother?" is described, and in the system instruction sentence area, "Please rephrase it in other words." is described. Example 2 is an example where the question sentence in the conversion instruction sentence is described in natural language and the system instruction sentence is described in code. In the question sentence area of the conversion instruction sentence, "Why can't I run as fast as my older brother?" is described, and in the system instruction sentence area, "paraphrase" is described.

[0325] After that, based on this conversion instruction sentence, the large language model generates a response, and the question sentence is converted (step S861). As described above, when the user's question sentence is "Why can't I run as fast as my older brother?", the large language model generates a converted question sentence such as "What is the reason why I run slower than my older brother?" as shown in FIG. 16B as an answer sentence obtained by rephrasing the above question sentence.

[0326] In addition, as processing for the response of this large language model (output to the user), the control unit 1110 displays, for example, as shown in FIG. 16B, the original question sentence "Why can't I run as fast as my older brother?" and the converted question sentence "What is the reason why I run slower than my older brother?" on the display unit 10011, and displays a request for the user to answer whether the converted question sentence and the original question sentence match (step S900).

[0327] By having the large language model perform the conversion process of the question sentence in this way, it is possible to confirm in advance whether the large language model properly understands the user's question sentence before creating an answer. Therefore, it becomes easier to obtain a more appropriate answer to the user's question from the large language model.

[0328] The processing flow shown in FIG. 17 is a modified example of the processing flow shown in FIG. 16, and is an example in which, as a conversion process for a question sentence, a process of creating a summary of the question sentence is performed. Specifically, in the processing flow shown in FIG. 17, when a question sentence is input by the user in step S810, the process proceeds to step S831 without performing a check process on the question sentence, and a summary instruction sentence for instructing the creation of a summary of the question sentence is created. Then, the generated instruction sentence (summary instruction sentence) is transmitted to the large language model (step S840).

[0329] When the large language model receives this summary instruction sentence (step S850), it then proceeds to step S861 and generates a response to the summary instruction sentence. Specifically, as a conversion process for the question sentence created by the user, a summary of the question sentence is created. In other words, the question sentence input by the user is converted into a sentence that concisely summarizes the key points. Thereafter, as a response of the large language model to the instruction sentence (summary instruction sentence), the question sentence (converted question sentence) subjected to the conversion process is transmitted to the control unit 1110.

[0330] Note that since the subsequent processing flow by the control unit 1110 has already been described, the description here is omitted. Also, in the example shown in FIG. 17, a check process for the question sentence input by the user is not performed, but similar to the example shown in FIG. 14, a check process for the question sentence may be performed. In the example of FIG. 17, for example, as a check process for the question sentence, it may be determined whether the number of words in the question sentence is equal to or more than a predetermined number, and when the number of words in the question sentence is equal to or more than the predetermined number, a summary instruction sentence may be created. That is, the check item may include that the number of words in the question sentence is equal to or more than a predetermined number. Of course, the content of the check process in this case is not particularly limited and may be determined as appropriate.

[0331] Here, with reference to FIGS. 18A and 18B, the processing flow shown in FIG. 17 will be further described. In step S810, assume that the user inputs a question sentence as follows: "The government has officially announced that starting from next year, it will target large enterprises for subsidies and reduce the amount of subsidies paid. Please explain the impact of this policy on society, including the keyword of corporate independence." In this case, in step S830, as an example, the summarization instruction sentence is generated in the form of the question sentence + the system instruction sentence. In this example, the system instruction sentence is for instructing the summarization of the question sentence.

[0332] Example 1 in FIG. 18A is an example where the question sentence and the system instruction sentence in the summarization instruction sentence are each described in natural language. In the question sentence area of the summarization instruction sentence, "The government has officially announced that starting from next year, it will target large enterprises for subsidies and reduce the amount of subsidies paid. Please explain the impact of this policy on society, including the keyword of corporate independence." is described, and in the system instruction sentence area, "Please summarize." is described. Also, example 2 in FIG. 18A is an example where the question sentence in the summarization instruction sentence is described in natural language and the system instruction sentence is described in code. In the question sentence area of the summarization instruction sentence, "The government has officially announced that starting from next year, it will target large enterprises for subsidies and reduce the amount of subsidies paid. Please explain the impact of this policy on society, including the keyword of corporate independence." is described, and in the system instruction sentence area, "summarize" is described.

[0333] After that, based on this summary instruction text, the large language model performs the conversion of the question text as the response generation (step S861). As described above, when the user's question text is "The government has officially announced that it will narrow the target to large enterprises for subsidies starting from next year and reduce the payment amount. Please explain the impact of this policy on society, including the keyword of corporate independence.", the large language model generates, as the answer text summarizing the above question text, for example, as shown in FIG. 18B, the converted question text "Subsidies will only be paid to large enterprises starting from next year, and the payment amount will be reduced from the current level. Please explain the impact of this on corporate independence."

[0334] Also, as a process for the response of this large language model (output to the user), for example, as shown in FIG. 18B, the control unit 1110 displays on the display unit 10011 the original question text or request text "Subsidies will only be paid to large enterprises starting from next year, and the payment amount will be reduced from the current level. Please explain the impact of this on corporate independence." and the converted question text "Subsidies will only be paid to large enterprises starting from next year, and the payment amount will be reduced from the current level. Please explain the impact of this on corporate independence.", and makes a display requesting the user to answer whether the converted question text and the original question text match (step S900).

[0335] By the way, it is not necessary to always enable the function of performing the above-described check process and conversion process of the question text. For example, when the user uses the system or when ending the system, the user may be able to select whether to enable this function. The function of performing the check process and conversion process of the question text is a function of correcting questions and requests, and is hereinafter also referred to as a question correction function. FIG. 19A is a diagram showing an example of a processing flow at the time of system startup, and FIG. 19B is a diagram showing an example of a processing flow at the time of system termination.

[0336] When the system starts up, as shown in an example in FIG. 19A, for example, when the operation input unit 1107 or the like is operated by the operator, the control unit 1110 receives a system startup request (step S1910). When receiving the system startup request, the control unit 1110 executes user authentication (step S1920). Next, in step S1930, the operator who operates the response output system checks whether the operator is a registered user. Note that the method for user authentication and for checking whether the operator is a registered user is not particularly limited, and an existing method may be adopted.

[0337] If the operator is a registered user (step S1930: Yes), then it is checked whether it is within the expiration date, that is, whether it is within the expiration date of the time-limited use contract of this system or function or service (step S1940). Note that if this system or function or service can be used without an expiration date, step S1940: Yes may be assumed and the process may proceed to step S1950. And if it is within the expiration date (step S1940: Yes), the process proceeds to step S1950, and the previously set user information for each user is read out, and the user information is reset or changed as necessary. The ON / OFF information of the above question correction function is included in this user information. That is, by changing the user information, the ON / OFF of the above question correction function is switched. Then, the process proceeds to step S1960, and main processing, for example, the processing of the answers to the questions described in FIG. 7 and the like, is started.

[0338] Note that if it is not within the expiration date in step S1940, that is, if the expiration date has passed (step S1940: No), the process proceeds to step S1970, and the user is notified that the expiration date has passed. Also, if the operator is not a registered user in step S1930 (step S1930: No), the process proceeds to step S1980, and it is checked whether the operator is a rental user.

[0339] Here, when the operator is not a rental user, for example, when the operator is a new user or the like (step S1980: No), the process proceeds to step S2000 to execute user registration processing. On the other hand, when the operator is a rental user (step S1980: Yes), it is confirmed in step S1990 whether to perform user registration for the rental user. When there is an answer from the rental user to perform user registration (step S1990: Yes), the process proceeds to step S2000 to perform user registration processing.

[0340] When performing this user registration processing, next, the user is requested to set the ON / OFF of the question correction function as one of the user information (step S2010). And when the user selects "ON" for the ON / OFF setting of the question correction function (step S2010: Yes), the question correction function is set to "ON", and when the user selects "OFF" for the ON / OFF setting (step S2010: No), the question correction function is set to "OFF", and the user information is set. Then, the process proceeds to step S1960 to start the main processing.

[0341] Note that in step S1990, when there is an answer from the rental user not to perform user registration (step S1990: No), the process proceeds to step S2010 without performing user registration processing, and only the selection of the ON / OFF setting of the question correction function is requested from the rental user as the user. Incidentally, the ON / OFF setting information of the question correction function for the rental user is temporarily stored only during the use of this system.

[0342] At the end of the system, as shown in an example in FIG. 19B, for example, when the operation input unit 1107 or the like is operated by a user including the rental user, the control unit 1110 receives a system termination request (step S2040). When the control unit 1110 receives the system termination request, it then confirms whether the user who used this system this time is a registered user (step S2050).

[0343] Here, if the user using the system is not a registered user, for example, if the user is a rental user or the like described above (step S2050: No), the process proceeds to step S2060. In step S2060, for the user who uses the system this time, user information including the ON / OFF setting of the question / request correction function is stored. Then, system termination processing is performed (step S2070). Also, if the user using the system is a registered user (step S2050: Yes), since the user information including the ON / OFF setting of the question correction function has already been stored, the system termination processing is performed in step S2070 without storing it again.

[0344] In this way, the ON / OFF setting of the question correction function can be performed, for example, when setting user information in the initial setting, but it may also be possible to change the setting at any timing during system use. For example, by selecting a predetermined menu item displayed on the display unit 10011 during system use, a user setting menu screen for setting user information is displayed, and the ON / OFF setting of the question / request correction function can be changed on this user setting menu screen. FIG. 20 is a diagram showing an example of the user setting menu screen.

[0345] In the example shown in FIG. 20, the ON / OFF of the question / request correction function can be selected, and the upper limit number of repetitions can be set when the question / request correction function is set to ON. As a result, the question / request correction function is executed within the range of the number of times set as the upper limit number of repetitions. Although the accuracy of the answer to the question text improves as the number of repetitions of the question / request correction function increases, it takes a long time to obtain an answer. Depending on the user and the content of the question, there may be cases where an answer in a shorter time is required rather than high accuracy. By setting the upper limit number of repetitions, it becomes easier to obtain an answer that meets the user's requirements even in such cases.

[0346] Next, with reference to the processing flow of FIG. 21, the processing flow of the response output system according to Example 7 will be further described. FIG. 21 is a modified example of the processing flow shown in FIG. 11, and is an example of determining whether or not a "technical term" as a specific expression is included in the question text as a check process.

[0347] Also in the example shown in FIG. 21, when the control unit 1110 receives a user input (step S810), it performs a database collation process on the question text in step S8211. In this collation process, the question text input by the user is collated with the specific expressions registered in the check DB 19010. In this example, as the collation process, the question text is collated with the technical terms registered in the check DB 19010.

[0348] Next, in step S8212, the control unit 1110 determines whether or not there is a technical term registered in the check DB 19010 in the question text. If the question text does not include a technical term as a specific expression (step S8212: No), the process proceeds to step S830, where a response instruction text is generated, and the generated response instruction text is transmitted to the large language model (step S840).

[0349] On the other hand, if the question text includes a technical term (step S8212: Yes), the process proceeds to step S8230, and it is determined whether or not the question by the user is asking about the meaning of the technical term as a specific expression. If the question by the user is asking about the meaning of the above technical term (step S8230: Yes), for example, with reference to the check DB 19010, the meaning or explanation of the above technical term is output to the user. Specifically, the explanation of the technical term or the like is displayed on the display unit 10011. In this embodiment, as described above, in the check DB 19010, the meaning and explanation of specific phrases such as technical terms are also stored. Therefore, the control unit 1110 can output the term explanation or the like by referring to the check DB 19010. Of course, the control unit 1110 may output the term explanation or the like by referring to other databases connected via the Internet or the like.

[0350] By doing so, it is possible to answer the user's questions regarding the meaning of technical terms without using a large language model. Also, when the user's question is about the meaning of a technical term, it is impossible to present the answer the user desires if the large language model does not correctly understand the technical term. However, referring to the check DB 19010 makes it easier to obtain the answer the user desires.

[0351] In step S8230, when the question by the user is not about the meaning of a technical term (step S8230: No), the process proceeds to step S8240 to perform conversion processing on the question text. Specifically, the technical term as a specific expression included in the question text is converted into another commonly used term (general term). Thereafter, the process proceeds to step S830, a response instruction text is created based on the converted question text using the general term, and the created response instruction text is transmitted to the large language model (step S840). By converting the technical term into a general term in this way, it becomes easier to obtain the answer the user desires even when the large language model does not correctly understand the technical term.

[0352] Note that in the processing flow shown in FIG. 21, steps S8211, S8212, S8230, and S8240 correspond to the check processing and conversion processing of the question text. Also, in step S850, the processing after the large language model receives the response instruction text is the same as the processing flow in FIG. 11 and the like, so the illustration and description in FIG. 21 are omitted.

[0353] Referring to FIGS. 22A and 22B, the processing flow shown in FIG. 21 will be further described. FIGS. 22A and 22B are diagrams for explaining the processing when the user's question text includes a technical term. Also, FIG. 22A is a diagram for explaining the processing when the user's question is about the meaning of a technical term, and FIG. 22B is a diagram for explaining the processing when the user's question is not about the meaning of a technical term.

[0354] In step S810, assume that the user inputs a question sentence "Please tell me the meaning of the word 'fail-safe'." as shown in Fig. 22A. Then, the control unit 1110 checks the question sentence against the check DB 19010 (step S8211), and assumes that the technical term "fail-safe" is found in the check DB 19010 (step S8212: Yes). In this case, the control unit 1110 determines whether the user's question is asking about the meaning of the technical term "fail-safe" (step S8230). In this example, the control unit 1110 determines from the content of the question sentence that the user's question is asking about the meaning of the technical term "fail-safe" (step S8230: Yes), refers to the check database, etc., and outputs the term explanation of "fail-safe" to the user (step S825). As output to the user, the control unit 1110 causes the display unit 10011 to display a term explanation such as "Fail-safe means that when a failure occurs in a device or system, it always controls to the safe side."

[0355] Also, in step S810, assume that the user inputs a question sentence "Please tell me an example of a fail-safe around me." as shown in Fig. 22B. Then, the control unit 1110 checks the question sentence against the check DB 19010 (step S8211), and assumes that the technical term "fail-safe" is found in the check DB 19010 (step S8212: Yes). Also in this case, the control unit 1110 determines whether the user's question is asking about the meaning of the technical term "fail-safe" (step S8230). In this example, the control unit 1110 determines from the content of the question sentence that the user's question is not asking about the meaning of the technical term "fail-safe" (step S8230: No).

[0356] Therefore, the control unit 1110 then creates a question text (converted question text) in which the technical terms in the question text are replaced with general terms (step S8240). As an example, the control unit 1110 creates a converted question text of "Please tell me a familiar example of a mechanism that maintains safety when a failure occurs in a device or system." from the question text input by the user. Further, a response instruction text based on the created converted question text is generated (step S830), and the generated response instruction text is transmitted to the large language model (step S840).

[0357] After that, when the large language model receives the response instruction text (step S850), it creates an answer text for the converted question text as a response based on this response instruction text (step S860). As an example, the large language model creates an answer text of "A familiar example is that when a heating appliance detects shaking or tipping over, it automatically stops operating." as a response to the above converted question text. The created answer text is transmitted to the control unit 1110 (step S870). When the control unit 1110 receives the transmitted answer text (step S880), as processing for the response of the large language model, an answer text to the user's question is output to the user (step S890).

[0358] By the way, it is also conceivable that the answer text created by the large language model for the user's question text contains technical terms that the user does not know. In such a case, an explanation of the technical terms may be added to the answer text as necessary.

[0359] FIG. 23 is a modified example after the control unit 1110 receives the response from the large language model in step S880. In FIG. 23, since the processing flow before the response transmission by the large language model (step S870) is the same as the processing flow in FIG. 11 etc., the illustration and description in FIG. 23 are omitted.

[0360] As shown in FIG. 23, when the control unit 1110 receives an answer sentence as a response from the large language model (step S880), it determines whether an explanation of a technical term is required for this answer sentence (step S920). The determination method here is not particularly limited, but as an example, a method similar to the above-described check process for the question sentence can be adopted. Specifically, the answer sentence is compared with the check DB 19010, and based on the determination result of whether the answer sentence contains a technical term as a specific expression registered in the check DB 19010, it is determined whether an explanation of the technical term is necessary. That is, if the answer sentence contains a technical term as a specific expression, it is determined that an explanation of the technical term is required (step S920: Yes), and if the answer sentence does not contain a technical term as a specific expression, it is determined that an explanation of the technical term is not required (step S920: No).

[0361] Here, when it is determined that an explanation of the technical term is not required (step S920: No), the control unit 1110 outputs the answer sentence to the user for the user's question transmitted from the large language model as processing for the response of the large language model (step S891). The case where it is determined that an explanation of the technical term is not required is, for example, when the technical term included in the answer sentence has already been explained to the user in the past, or when it is used as a general term in a certain expression depending on the usage of the expression, etc.

[0362] On the other hand, when it is determined that an explanation of the technical term is required (step S920: Yes), the process proceeds to step S832, and the control unit 1110 creates an instruction sentence (additional instruction sentence) instructing the large language model to add an explanation of the technical term to the answer sentence, and transmits the created additional instruction sentence to the large language model (step S841).

[0363] When the large language model receives the additional instruction sentence transmitted from the control unit 1110 (step S851), it creates an explanation of the technical term as a response to this additional instruction sentence (step S861). Then, the created term explanation is transmitted to the control unit 1110 (step S871).

[0364] When the control unit 1110 receives a response transmitted from the large language model, that is, a term explanation (step S881), as processing for the response of the large language model, it outputs the term explanation to the user together with the answer sentence that has already been created (step S892).

[0365] With reference to the explanatory diagram of FIG. 24, the processing flow of FIG. 23 will be further described. As shown in FIG. 24, for example, a user inputs a question sentence "What kind of surgery is Tommy John surgery?" As a response to this question sentence, the large language model creates an answer sentence "Tommy John surgery is a surgical method mainly for reconstructing ligaments by excising the damaged elbow ligaments caused by pitching motions or the like and transplanting the normal tendon excised from the palmaris longus muscle of the patient to the affected part." and transmits the created answer sentence to the control unit 1110 (step S870).

[0366] When the control unit 1110 receives this answer sentence (step S880), as an example, it collates the answer sentence with the check DB 19010. At this time, if "palmaris longus muscle" included in the answer sentence is registered as a technical term in the check DB 19010, the control unit 1110 determines that an explanation of the technical term "palmaris longus muscle" is necessary (step S920), generates an additional instruction sentence such as "Please explain the term palmaris longus muscle." (step S832), and transmits the generated additional instruction sentence to the large language model (step S841).

[0367] When the large language model receives the additional instruction sentence (step S851), as a response to this additional instruction sentence, it creates a term explanation "The palmaris longus muscle is a muscle that performs palmar flexion of the wrist joint and tension of the palmar aponeurosis in the human upper limb." (step S861). Then, it transmits the created term explanation (step S871). At this time, the answer sentence may be transmitted together as needed.

[0368] When the control unit 1110 receives a glossary explanation (step S881), as processing for the response of the large language model (answer to the question / with glossary explanation), it displays the user's question text "What kind of surgery is Tommy John surgery?" the answer text of the large language model "Tommy John surgery is a surgical method mainly for reconstructing ligaments by excising the ligaments of the elbow damaged by pitching motions etc. and transplanting the normal tendon removed from the palmaris longus muscle of the patient to the affected part." and the glossary explanation by the large language model "The palmaris longus muscle is a muscle that performs palmar flexion of the wrist joint and tension of the palmar aponeurosis in the human upper limb." on the display unit 10011 (step S892).

[0369] In addition, in step S920, if no specialized terms registered in the check DB 19010 are found in the answer text (step S920: No), the control unit 1110 displays the user's question text and the answer text of the large language model on the display unit 10011 as processing for the response of the large language model (answer to the question / without glossary explanation) (step S891).

[0370] By adopting such a response output processing flow, the effort required for the user to question the large language model can be saved. For example, if the answer text of the large language model contains specialized terms unknown to the user, the user will need to further question the large language model about the meaning of the specialized terms. However, by appropriately adding explanations of specialized terms to the answer text, the effort of the user to ask further questions can be saved.

[0371] However, even if the response contains technical terms, it is not always necessary to add explanations of the technical terms. For example, as shown in Fig. 25, the necessity of explaining technical terms according to the situation may be preset. In this example, even if the response contains technical terms, if the technical terms in the response are also included in the question, the explanation of the technical terms is set to "not required". That is, even if the response contains technical terms, if the technical terms in the response are also included in the user's question, it is determined that the user knows this technical term, and the term explanation is not added. Therefore, in the case of the setting example shown in Fig. 25, the term explanation is added to the response only when the response contains technical terms but the technical terms in the response are not included in the user's question. Note that the user may be allowed to arbitrarily change the setting of the necessity of the term explanation.

[0372] Next, with reference to the processing flow of Fig. 26, the response output processing flow of the response output system of Example 7 will be further described. Fig. 26 is a modified example of the processing flow shown in Fig. 11 and the like, and is an example of determining whether the question contains "abbreviation" as a specific expression as a check process.

[0373] Also in the example shown in Fig. 26, when the control unit 1110 receives a user input (step S810), it performs a database collation process on the question text in step S8211. In this collation process, the question text input by the user is collated with the specific expressions registered in the check DB 19010. In this example, as the collation process, the question text is collated with the "abbreviation" registered in the check DB 19010.

[0374] Next, in step S8212, it is determined whether the question text contains the abbreviations registered in the check DB 19010. If the question text does not contain the abbreviation as a specific expression (step S8212: No), a response instruction text is generated in step S830, and the generated response instruction text is transmitted to the large language model (step S840).

[0375] On the one hand, when the question sentence contains an abbreviation as a specific expression (step S8212: Yes), the process proceeds to step S8250, and the control unit 1110 executes the conversion process of the question sentence. In this example, a process of asking the user for details of the abbreviation included in the question sentence is performed. In other words, the control unit 1110 performs a process of requesting the user to input an explanation of the abbreviation included in the question sentence or a phrase when the abbreviation is expressed without being abbreviated (for example, expressing the abbreviation ETC as Electronic Toll Collection without abbreviation). Then, when receiving the user input for this inquiry, that is, the detailed information about the abbreviation (step S8260), the control unit 1110 generates a response instruction sentence based on the user's question sentence and the detailed information of the abbreviation (step S830), and transmits the created response instruction sentence to the large language model (step S840).

[0376] By adopting such a processing flow, even when the user's question sentence contains an abbreviation, it becomes easier to obtain a more appropriate answer to the question from the large language model. This is particularly effective when the abbreviation included in the question sentence has multiple meanings or when the abbreviation is familiar to the user but not generally widely recognized.

[0377] Note that in the processing flow of FIG. 26, steps S8211, S8212, and S8250 correspond to the check process and conversion process of the question sentence. Also, since the processing flow after the large language model receives the response instruction sentence in step S850 is the same as the processing flow in FIG. 11 etc., the illustration and description in FIG. 26 are omitted.

[0378] Next, with reference to the processing flow of FIG. 27, the response output processing flow of the response output system of Example 7 will be further described. FIG. 27 is a modified example of the processing flow shown in FIG. 11 etc., and is an example of determining whether a "dialect" as a specific expression is included in the question sentence as a check process.

[0379] Even in the example shown in FIG. 27, when the control unit 1110 receives a user input (step S810), it performs a database collation process for the question sentence in step S8211. In this collation process, the question sentence input by the user is collated with the specific expressions registered in the check DB 19010. In this example, as the collation process, the question sentence is collated with the "dialect" registered in the check DB 19010.

[0380] Next, in step S8212, it is determined whether the question sentence includes a dialect as a specific expression. And when the question sentence does not include a dialect as a specific expression (step S8212: No), the control unit 1110 generates a response instruction sentence in step S830 and transmits the generated response instruction sentence to the large language model (step S840).

[0381] On the other hand, when the question sentence includes a dialect (step S8212: Yes), it proceeds to step S8270, and the control unit 1110 executes the conversion process of the question sentence. In this example, the control unit 1110 identifies the dialect region that the user is likely to use based on the pre-registered user information and the conversation between the user and the large language model before the question, etc., and performs a process of converting the dialect of the question sentence into standard language based on the identified dialect region. After that, a response instruction sentence is generated based on the question sentence (converted question sentence) after the conversion process (step S830), and the generated response instruction sentence is transmitted to the large language model (step S840). That is, whether or not a dialect is used in the question sentence by the user, the control unit 1110 generates a response instruction sentence based on the question sentence described in standard language and transmits it to the large language model.

[0382] Similar to the case of the processing flow shown in FIG. 11 etc., when the large language model receives a response instruction sentence (step S850), it generates a response (answer sentence) corresponding to the received response instruction sentence (step S860), and transmits the generated response (answer sentence) to the control unit 1110 (step S870).

[0383] When the control unit 111 receives a response transmitted from the large language model (step S880), it then determines whether to use a dialect in the user output (step 930).

[0384] In this example, if the question sentence input by the user uses a dialect, it is determined to use a dialect in the response sentence output to the user (step S930: Yes). In other words, if the response instruction sentence is based on the converted question sentence, it is determined to use a dialect in the user output. In this case, the process proceeds to step S940, and the control unit 1110 executes the conversion process of the response sentence. Specifically, a process of converting the standard language of the response sentence into a dialect corresponding to the specified dialect area is performed. After that, as a process for the response of the large language model, the control unit 1110 outputs a response sentence using the dialect (converted response sentence) to the user (step S892).

[0385] On the other hand, in step S930, if the question sentence input by the user uses the standard language, it is determined to use the standard language in the response sentence output to the user (step S930: No). In other words, if the response instruction sentence is based on the original question sentence input by the user, it is determined to use the standard language in the user output. In this case, the process proceeds to step S893, and the control unit 1110 outputs a response sentence using the standard language to the user as a process for the response of the large language model (step S893).

[0386] By adopting such a processing flow, even when the user's question sentence contains a dialect, the question is asked after converting it so that the large langu...

Claims

1. An input unit for receiving a question text input by a user, a control unit for generating a response instruction text for a large language model based on the question text and obtaining a response text generated by the large language model for the response instruction text, and an output unit for outputting based on the response text obtained by the control unit, wherein the control unit performs conversion processing on the question text and generates the response instruction text based on the question text after the conversion processing, a response output system.

2. In the response output system according to claim 1, the control unit performs a check process to check whether the question text corresponds to a predetermined check item, and performs the conversion process of the question text when the question text corresponds to the check item, a response output system.

3. The response output system according to claim 2, wherein the check item includes that the question text contains a preset specific expression, and the control unit performs the conversion process of the question text when the specific expression is included in the question text in the check process, a response output system.

4. The response output system according to claim 3, wherein the specific expression is classified into a plurality of groups, and the control unit as the check process, checks whether the question text contains a specific expression classified into a preset set group among the plurality of groups, and performs the conversion process of the question text when the question text contains the specific expression classified into the set group, a response output system.

5. The response output system according to claim 4, wherein a priority is preset for each of the set groups, and the control unit checks whether the question text contains the specific expression classified into each set group in the order according to the priority, a response output system.

6. The response output system according to claim 1, wherein the control unit as the conversion process of the question text, generates a conversion instruction text for requesting conversion of the question text, and performs a process of obtaining, as the question text after the conversion process, the question text converted by the large language model according to the conversion instruction text, a response output system.

7. The response output system according to claim 1, wherein the control unit as the conversion process of the question text, Generate a summary instruction sentence that requests a summary of the question sentence, and perform a process of obtaining, as the question sentence after the conversion process, a summary of the question sentence created by the large language model according to the summary instruction sentence. Response output system.

8. The response output system according to claim 1, wherein the control unit as the conversion process of the question sentence performs a process of requesting the user to re-create the question sentence. Response output system.

9. The response output system according to claim 3, wherein a voice input unit that receives the question sentence as voice input, and an operation input unit that receives the question sentence as character input, and the control unit when the question sentence input from the voice input unit by the user contains a dialect as the specific expression as the conversion process of the question sentence performs a process of requesting the user to input the question sentence from the operation input unit. Response output system.

10. The response output system according to claim 1, wherein the control unit after executing the conversion process of the question sentence, asks the user whether the question sentence after the conversion process matches the user's intention, when there is an answer from the user that the question sentence after the conversion process matches the user's intention, generates the response instruction sentence based on the question sentence after the conversion process. Response output system.

11. The response output system according to claim 19, wherein the control unit when there is an answer from the user that the question sentence after the conversion process does not match the user's intention in response to the inquiry, further performs a process of requesting the user to re-create the question sentence. Response output system.

12. The response output system according to claim 1, wherein the control unit as the conversion process of the question sentence performs a first translation process of translating the question sentence input by the user in the first language into a second language different from the first language, and a second translation process of translating the question sentence after the first translation process from the second language into the first language. Response output system.

13. The response output system according to claim 12, wherein the control unit asks whether the question sentence after the second translation process matches the user's intention, when there is an answer from the user that the question sentence after the conversion process matches the user's intention, generates the response instruction sentence based on the question sentence after the second translation process. Response output system.

14. The response output system according to claim 13, wherein the control unit when there is an answer that the question sentence after the second translation process does not match the user's intention for the inquiry, further performs a process of instructing the user to remake the question sentence. Response output system.

Citation Information

Patent Citations

  • Structural unit for tank construction

    JP1977008512A