Response output device, response output method, and information processing device

JP2026131463APending Publication Date: 2026-08-14MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0007】 本発明によれば、より好適な応答出力技術を提供できる。これ以外の課題、構成および効果は、以下の実施形態の説明において明らかにされる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026131463000001_ABST
    Figure 2026131463000001_ABST
Patent Text Reader

Abstract

To provide a more suitable artificial intelligence response output technology. According to this invention, it will contribute to Sustainable Development Goals (SDGs) "9. Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation" and "11. Make cities and human settlements inclusive, safe, resilient and sustainable." [Solution] A response output device comprising a sensor that receives gesture operation input from a user that mimics natural language characters, a control unit, and an output unit, wherein the control unit generates an instruction sentence including natural language and gesture transmission data generated based on the operation input to the sensor, transmits the instruction sentence to a multimodal language model, acquires character data of the gesture recognition result as a response generated by inference performed by the multimodal language model based on the instruction sentence, determines the processing for the gesture operation input from the user based on the acquired character data and a correspondence table held by the response output device, and performs control by outputting an output from the output unit based on the determined processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a response output system.

Background Art

[0002] Regarding response output technology using artificial intelligence such as a language model, for example, it is disclosed in Patent Document 1.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the disclosure of Patent Document 1, considerations regarding configurations for more suitably providing response output technology using artificial intelligence to users were not sufficient.

[0005] An object of the present invention is to provide a more suitable response output technology.

Means for Solving the Problems

[0006] To solve the above problems, for example, the configuration described in the claims may be adopted. The present application includes multiple means for solving the above problems, but one example is the following configuration: A response output device comprising a sensor that receives gesture operation input from a user that mimics natural language characters, a control unit, and an output unit, wherein the control unit generates an instruction sentence including natural language and gesture transmission data generated based on the operation input input to the sensor, transmits the instruction sentence to a multimodal language model, acquires character data of the gesture recognition result as a response generated by inference performed by the multimodal language model based on the instruction sentence, determines the processing for the gesture operation input from the user based on the acquired character data and a correspondence table held by the response output device, and performs control by outputting an output from the output unit based on the determined processing. [Effects of the Invention]

[0007] According to the present invention, a more suitable response output technology can be provided. Other problems, configurations, and effects will be clarified in the following description of embodiments. [Brief explanation of the drawing]

[0008] [Figure 1A] This figure shows an example of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 1B] This figure shows an example of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 1C] This figure shows an example of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 1D] This figure shows an example of a server according to one embodiment of the present invention. [Figure 1E] This is an explanatory diagram of an example of an artificial intelligence response output system according to one embodiment of the present invention. [Figure 1F] This figure shows an example of an optical system and retroreflector for an aerial levitation image display device according to one embodiment of the present invention. [Figure 1G]This figure shows an example of an aerial levitation image display device according to one embodiment of the present invention. [Figure 1H] This figure shows an example of an optical system for an aerial levitation image display device according to one embodiment of the present invention. [Figure 1I] This figure shows an example of a retroreflective plate for an aerial levitation image display device according to one embodiment of the present invention. [Figure 1J] This figure shows an example of an aerial levitation image display device according to one embodiment of the present invention. [Figure 1K] This figure shows an example of an aerial levitation image display device according to one embodiment of the present invention. [Figure 2A] This is an explanatory diagram of an example of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 2B] This is an explanatory diagram of an example of a mobile information processing terminal according to one embodiment of the present invention. [Figure 2C] This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 2D] This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 2E] This is an explanatory diagram illustrating an example of a display example of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 3A] This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 3B] This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 3C] This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 3D] This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. [Figure 3E] This is an explanatory diagram illustrating an example of table information according to one embodiment of the present invention. [Figure 3F] This is an explanatory diagram of an example of a conversation used to describe one embodiment of the present invention. [Figure 3G] It is an explanatory diagram of an example of a conversation example used for explaining one embodiment of the present invention. [Figure 4A] It is a diagram showing an example of an artificial intelligence response output device and a system according to one embodiment of the present invention. [Figure 4B] It is a diagram showing an example of a processing flow of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4C] It is a diagram showing an example of an interface in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4D] It is a diagram showing an example of an interface in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4E] It is a diagram showing a configuration example of an instruction sentence in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4F] It is a diagram showing a configuration example of a correspondence table in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4G] It is a diagram showing an example of a processing flow of an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4H] It is a diagram showing a display example in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4I] It is a diagram showing a display example in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4J] It is a diagram showing a configuration example of an interface in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4K] It is a diagram showing a configuration example of an interface in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4L] It is a diagram showing a configuration example of an interface in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4M] It is a diagram showing a configuration example of an interface in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4N] It is a diagram showing a display example in an artificial intelligence response output device according to one embodiment of the present invention. [Figure 4O]This figure shows an example of the sensor configuration in an artificial intelligence response output device according to one embodiment of the present invention. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described in detail below with reference to the drawings. However, the present invention is not limited to the examples described herein, and various modifications and alterations are possible by those skilled in the art within the scope of the technical ideas disclosed herein. Furthermore, in all the figures used to illustrate the present invention, components having the same function are given the same reference numerals, and repeated descriptions may be omitted.

[0010] Furthermore, if the artificial intelligence response output device according to each embodiment of the present invention has a display screen, it may be called a display device. If the artificial intelligence response output device has a voice output function, it may be called a voice output device. The artificial intelligence response output device may simply be called an information processing device. A system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence response output system. Also, if the artificial intelligence response output device provides a response service of a large-scale language model, which is artificial intelligence, to the user and assists the user, the artificial intelligence response output device or the display output of the artificial intelligence response output device can become an artificial intelligence (AI) assistant for the user. Therefore, in this case, the artificial intelligence response output device may be called an AI assistant device or an AI assistant display device. Similarly, in this case, a system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an AI assistant system or an AI assistant display system. Also, in this case, since the artificial intelligence response output device becomes an interface between the user and artificial intelligence, it may be called an artificial intelligence interface device. In this case, a system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence interface system.

[0011] <Example 1> As Embodiment 1 of the present invention, an artificial intelligence response output device and system that outputs a response from a large-scale language model artificial intelligence will be described.

[0012] An example of the artificial intelligence response output device 10010 of the present invention will be described using Figure 1A. Furthermore, an example of a system including the artificial intelligence response output device 10010 and the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 will be described in the case where the artificial intelligence response output device 10010 cooperates with the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 through communication or other means.

[0013] In the example shown in Figure 1A, the artificial intelligence response output device 10010 has a display unit 10011. In the example shown in Figure 1A, the display unit 10011 may be a flat panel display, a screen that projects images from the back, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which pixels emit light themselves. Furthermore, the display unit 10011 may be equipped with a touch operation input sensor and configured as a touch panel.

[0014] In the example shown in Figure 1A, the voice output unit 1140 of the artificial intelligence response output device 10010 is composed of a speaker. The artificial intelligence response output device 10010 is also equipped with a microphone 1139 that can pick up the user's voice. Through voice input from the microphone 1139 and user operation input via the operation input unit described later, the artificial intelligence response output device 10010 can acquire user input that forms the basis of instructions (prompts) for the large-scale language model, which is an artificial intelligence.

[0015] The artificial intelligence response output device 10010 may be equipped with a local large-scale language model. In this case, the response of the large-scale language model may be output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.

[0016] Furthermore, the artificial intelligence response output device 10010 may not have a local large-scale language model, but communicate with an external large-scale language model server 19001 and / or a multimodal large-scale language model server 20001, and output the response received from the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 as the display output of the display unit 10011 and / or the audio output of the audio output unit 1140.

[0017] Alternatively, the artificial intelligence response output device 10010 may also include a local large-scale language model and be configured to communicate with an external large-scale language model server 19001 having a large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be switched and output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140. Alternatively, the response generated based on both the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.

[0018] The configuration when the artificial intelligence response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example in Figure 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as a wireless example, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may have both wired and wireless sections, or it may pass through routers or repeaters. Similarly, the communication path from the communication unit 1132 to the Internet 19000 may also have both wired and wireless sections, or it may pass through routers or repeaters. The artificial intelligence response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 that is different from these servers, via the communication device 19011 and the internet 19000. The configuration including the artificial intelligence response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.

[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" should be understood as encompassing the local large-scale language model provided by the artificial intelligence response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.

[0020] In the example in Figure 1A, the display unit 10011 shows an example where each element is displayed in two display areas: an instruction display area 10051 where the user inputs an instruction (prompt) to a large-scale language model which is artificial intelligence, and an artificial intelligence response display area 10061 which displays the response from the large-scale language model. In the example in Figure 1A, the instruction display area 10051 shows an example where an icon 10052 representing the user, text such as natural language or software code 10053 as a component of the instruction, an image 10054 as a component of the instruction, a video 10055 as a component of the instruction, etc. In the example in Figure 1A, the artificial intelligence response display area 10061 shows an example where an icon 10062 representing artificial intelligence or an artificial intelligence assistant, text such as natural language or software code 10063 as a component of the response from artificial intelligence, an image 10064 as a component of the response from artificial intelligence, a video 10065 as a component of the response from artificial intelligence, etc. Note that the display example of the display unit 10011 of the artificial intelligence response output device 10010 shown in Figure 1A is merely an example. Depending on the implementation example in which the artificial intelligence response output device 10010 is used, a different display from the example shown in Figure 1A may be used.

[0021] Here, we will explain large-scale language models. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT have been made publicly available. These technologies can be used in this embodiment as well. These large-scale language models are artificial intelligence models that have been generated through extensive pre-training on natural language contained in a large number of documents and texts that exist in the human world. The number of parameters of these artificial intelligence models exceeds hundreds of millions. Furthermore, in addition to this, there are also models that incorporate reinforcement learning based on human feedback. An example of a base model is a model called Transformer. As an example of training these models, for example, Reference 1 is publicly available.

[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf

[0023] These large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization. More advanced models can perform natural language question answering (also called dialogue or conversation), natural language suggestion generation, and programming code generation. Because these artificial intelligence models have a very large number of parameters, training requires enormous amounts of data and computing resources. Therefore, training this level of artificial intelligence for a specific application is extremely resource-inefficient. For this reason, foundation models that can be applied to various uses are generated through large-scale pre-training. For example, the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 shown in Figure 1A may be equipped with such large-scale language models and configured to be usable by various terminals via an API (Application Programming Interface). Alternatively, the artificial intelligence response output device 10010 shown in Figure 1A may be equipped with a local large-scale language model and configured to use it itself. These large-scale language models can be generated by performing large-scale pre-training separately, and the generated large-scale language models can be duplicated and provided to the large-scale language model server 19001, the multimodal large-scale language model server 20001, and the artificial intelligence response output device 10010. In this way, instead of performing pre-training for each application or terminal, duplicating the large-scale language model, which is the foundation model generated by performing large-scale pre-training, and using it on individual servers and terminals allows for the sharing of resource consumption used for training, resulting in better resource efficiency.

[0024] Furthermore, even if a large-scale language model is generated as a foundational model through extensive pre-training, it may be configured to perform additional training, such as transfer learning, on individual servers or devices, depending on the application and purpose.

[0025] Furthermore, large-scale language models can pre-train on natural language and perform input / output processing targeting natural language. In addition, multimodal large-scale language model artificial intelligence capable of processing not only natural language text information but also other types of information is also applicable to the embodiments of the present invention. Figure 1A shows a large-scale language model server 20001 having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model artificial intelligence include GPT-4 (see Reference 2) and Gato (see Reference 3), which are publicly available. These technologies may also be used in this embodiment. These multimodal large-scale language models are artificial intelligence models generated by performing large-scale pre-training on natural language and other types of information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. Furthermore, there are also models that incorporate reinforcement learning based on human feedback. Hereinafter, information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.

[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf

[0027] Next, using Figure 1B, we will describe an example configuration of an artificial intelligence response output device 10010 that receives user input to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as the large-scale language model to the user input.

[0028] The artificial intelligence response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a position sensor 1114, a local LLM processing unit 10028, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, a clock 1196, and the like. The artificial intelligence response output device 10010 may have a large screen, such as a so-called monitor or television.

[0029] The display unit 10011 may be a flat panel display, a screen that projects images from the back, or a display that projects an optical image into the air to show a floating image. If the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Alternatively, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which pixels emit light themselves. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be equipped with a touch operation input sensor and configured to accept touch operation input from the user 230's finger. In this case, the display unit 10011 may be configured as a touch panel. Through the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire user input that forms the basis of instructions (prompts) to the large-scale language model, which is artificial intelligence.

[0030] The communication unit 1132 may be configured with a Wi-Fi® communication interface, a Bluetooth® communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 1132 of the artificial intelligence response output device 10010 can communicate with the communication device 19011 connected to the internet 19000. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless sections, and may also pass through routers or repeaters. In the case of a wired connection, the communication unit 1132 may have an Ethernet® connection interface as hardware and communicate using a LAN communication method. This allows the artificial intelligence response output device 10010 to communicate with various servers connected to the internet 19000.

[0031] The artificial intelligence response output device 10010 is equipped with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011 and the communication unit 1132, etc.

[0032] The power supply 1106 converts the AC current input from an external source via the external power input interface 1111 into DC current and supplies the necessary DC current to each part of the artificial intelligence response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. In addition, the secondary battery 1112 supplies power to each part that requires power via the external power input interface 1111 when power is not supplied from an external source.

[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit such as a remote controller, or an infrared light receiving unit, and inputs signals for operations other than touch operations by the user to the touch operation input sensor of the display unit 10011. Separately from the user who touches the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the artificial intelligence response output device 10010. Through the user's operation input via the operation input unit 1107, the artificial intelligence response output device 10010 can obtain user input that forms the basis of instructions (prompts) to the large-scale language model, which is artificial intelligence. Note that there may also be a modified configuration in which the touch operation input sensor of the display unit 10011 is included as part of the operation input unit 1107.

[0034] The video signal input unit 1131 receives video data by connecting an external video output device. Various digital video input interfaces are possible for the video signal input unit 1131. For example, it can be configured with an HDMI (High-Definition Multimedia Interface) standard video input interface, a DVI (Digital Visual Interface) standard video input interface, or a DisplayPort standard video input interface. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also use various USB interfaces.

[0035] The audio signal input unit 1133 receives audio data by connecting an external audio output device. The audio signal input unit 1133 may be configured as an HDMI standard audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI standard interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface with integrated terminals and cables.

[0036] The audio output unit 1140 is capable of outputting audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 is also capable of outputting audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured as a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output audio signals as digital signals to external devices, such as the Audio Return Channel function specified in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output audio signals as analog signals to external devices such as headphones.

[0037] Microphone 1139 is a microphone that picks up sounds from the vicinity of the artificial intelligence response output device 10010, converts them into signals, and generates audio signals. The microphone may be configured to record human voices, such as the user's voice, and the control unit 1110, described later, may perform speech recognition processing on the generated audio signal to obtain textual information from the audio signal. Through the audio input from microphone 1139, the artificial intelligence response output device 10010 can obtain user input that forms the basis of instructions (prompts) for the large-scale language model, which is an artificial intelligence.

[0038] The imaging unit 1180 is a camera having an image sensor. The camera may be provided on the front of the display unit 10011 side of the artificial intelligence response output device 10010, or on the back of the display unit 10011 side. Both a front camera and a rear camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a rear camera.

[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be composed of a magnetic recording medium such as a hard disk drive (HDD) or a semiconductor memory such as a solid-state drive (SSD). For example, the storage unit 1170 may have various types of information, such as video data, image data, and audio data, pre-recorded in it at the time of product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from external devices or external servers via the communication unit 1132. The video data, image data, etc., recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, etc., recorded in the storage unit 1170 may also be output to external devices or external servers via the communication unit 1132.

[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may also be called a video processing circuit and may be composed of hardware such as an ASIC, FPGA, or video processor. The video control unit 1160 may also be called a video processing unit or image processing unit. The video control unit 1160 performs video switching control, such as determining which video signal to input to the display unit 10011 from among the video signals stored in the memory 1109 and the video signals (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform image processing control on the video signals input from the video signal input unit 1131 and the video signals stored in the memory 1109. Examples of image processing include scaling processing such as enlarging, reducing, and transforming images; brightness adjustment processing to change the brightness; contrast adjustment processing to change the contrast curve of an image; and retinex processing which decomposes an image into its light components and changes the weighting of each component.

[0041] The attitude sensor 1113 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the attitude of the artificial intelligence response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each connected part.

[0042] The position sensor 1114 is a sensor that measures the position of the artificial intelligence response output device 10010 using a GNSS (Global Navigation Satellite System) such as GPS (Global Positioning System). If the position of the artificial intelligence response output device 10010 can be estimated based on information obtained through communication by the communication unit 1132, the position estimated by that method may be substituted for the measurement result of the position sensor 1114. The position sensor 1114 may also receive UTC (coordinated universal time) time information transmitted by the GNSS.

[0043] The non-volatile memory 1108 stores various data used by the artificial intelligence response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, data for various operations displayed on the display unit 10011 of the artificial intelligence response output device 10010, display icons, data and layout information for user-operated objects, etc. Memory 1109 stores video data and device control data displayed on the display unit 10011. The control unit 1110 may read various software from the storage unit 1170, expand it into memory 1109, and store it there.

[0044] The local LLM processing unit 10028 has memory capable of holding a large-scale language model (LLM) and can perform inference of the large-scale language model based on the control of the control unit 1110. The hardware can be a so-called GPU (Graphics Processing Unit). Another example of the hardware for the local LLM processing unit 10028 is a so-called NPU (Neural Network Processing Unit). The local LLM processing unit 10028 may perform training as well as inference. Note that the local LLM processing unit 10028 is not necessarily required if the execution of large-scale language model inference in the local environment of the artificial intelligence response output device 10010 is not necessary.

[0045] The control unit 1110 controls the operation of each connected part. The control unit 1110 may also work in cooperation with a program stored in memory 1109 to perform calculation processing based on information acquired from each part within the artificial intelligence response output device 10010. One of the control states of the control unit 1110 is, for example, the output of responses from the large-scale language model of the local LLM processing unit 10028, or responses from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001, acquired via the communication unit 1132, via the display unit 10011 or the audio output unit 1140, which is a speaker or the like. The specific configuration of the control unit 1110 is a CPU or the like, and may also be called a processor or control circuit.

[0046] Clock 1196 is a clock that counts time or time, which can be used in the control of the control unit 1110 of the artificial intelligence response output device 10010. Clock 1196 may function as an internal clock, rotating a counter value. Alternatively, clock 1196 may be a clock that counts UTC (coordinated universal time), a time used worldwide, or regional times such as JST (Japan Standard Time). Clock 1196 may operate autonomously as an internal clock, or it may calibrate its time using external information at predetermined timings. For example, if clock 1196 is a clock that counts UTC time, it may acquire time information such as NTP information for UTC from an NTP (Network Time Protocol) server or external device via the communication unit 1132, and calibrate its time using the acquired NTP information. Alternatively, the position sensor 1114 may acquire UTC time information received from GNSS, and clock 1196 may calibrate its time using the acquired UTC time information. For example, if clock 1196 is a clock that counts regional time such as JST, it may acquire regional time information such as TOT (Time Offset Table) from a server or external device via the communication unit 1132, and use the regional time information acquired by clock 1196 to calibrate the time. Alternatively, the position sensor 1114 may acquire location information including the latitude and longitude information of the artificial intelligence response output device 10010, which is received from GNSS, and UTC time information, and use the UTC time information and location information acquired by clock 1196 to calculate the time information for the region where the artificial intelligence response output device 10010 is located, and use this to calibrate the time.

[0047] Furthermore, when input is received from the user via the touch panel, microphone 1139, or operation input unit 1107 as described above, the control unit 1110 can perform the control to generate an instruction sentence based on that input and send it to the local large-scale language model of the local LLM processing unit 10028 of the artificial intelligence response output device 10010, the large-scale language model of the large-scale language model server 19001, or the multimodal large-scale language model of the large-scale language model server 20001, and to obtain a response from these large-scale language models.

[0048] In the examples shown in Figures 1A and 1B, an example was described in which the artificial intelligence response output device 10010 includes a display unit 10011. However, the artificial intelligence response output device 10010 according to the embodiment of the present invention does not necessarily have to include a display unit 10011. For example, even without a display unit 10011, the device can be configured to receive user input to the artificial intelligence via an audio signal input unit 1133 or a microphone 1139, and to output a response from the artificial intelligence, such as a large-scale language model, to the user input via an audio output unit 1140.

[0049] According to the artificial intelligence response output device and artificial intelligence response output system of Embodiment 1 of the present invention described above, it is possible to receive input from a user to artificial intelligence such as a large-scale language model and output a response to the user input generated by the inference of the artificial intelligence, such as a large-scale language model on a server device on the network or a local large-scale language model on the artificial intelligence response output device.

[0050] The artificial intelligence response output device 10010 according to the embodiment of the present invention can be implemented as a device in various forms. An example is shown in Figure 1C. For example, the artificial intelligence response output device 10010 may be a display device 181 such as a television, monitor, or display device as shown in Figure 1C(1). The display device 181 displays images on a display screen such as a display panel. The artificial intelligence response output device 10010 may be an information processing terminal 182 such as a smartphone, tablet terminal, or PC (personal computer) as shown in Figure 1C(2). The information processing terminal 182 displays images on a display screen such as a display panel. The artificial intelligence response output device 10010 may be a head-mounted display (HMD) 183 as shown in Figure 1C(3). In this case, the artificial intelligence response output device 10010 includes an eyepiece optical system that projects the image displayed by the display unit 10011 onto the user's eyes and generates a virtual image of the image. The artificial intelligence response output device 10010 may also be a head-up display (HUD) 184, as shown in Figure 1C(4), which generates a virtual image 192 by reflecting projected light onto a transparent member 191 such as glass, and displays an image on the virtual image. In this case, the artificial intelligence response output device 10010 includes an optical system for projecting light emitted from the display unit 10011 onto the transparent member 191. The artificial intelligence response output device 10010 may also be an aerial floating image display device 185, as shown in Figure 1C(5), which generates an optical image 195, which is a real image, in the air by reflecting light emitted from a display device 193 with an imaging optical plate 194 such as a retroreflective member or a corner reflector array, and displays an image on the optical image. The artificial intelligence response output device 10010 may also be a projection-type image display device 186, such as a projector, which projects a projected image 196 onto a screen or wall, as shown in Figure 1C(6). In this case, the artificial intelligence response output device 10010 includes a display panel which is a display unit 10011, a light source, an illumination optical system that guides light from the light source to the display unit 10011, and a projection optical system that projects the light transmitted or reflected by the display panel from the light source onto a screen, wall, or the like.

[0051] Next, an example of the configuration of the large-scale language model server 20001 according to an embodiment of the present invention will be described with reference to Figure 1D. While the configuration shown in this figure is described as an example of the configuration of the large-scale language model server 20001, the large-scale language model server 19001 and the second server 19002 may have similar configurations.

[0052] The large-scale language model server 20001 may include, for example, a user interface 22010, an input interface 22011, an output interface 22012, an external power input interface 22014, a power supply 22015, a storage unit 22020, a learning data database (DB) 22021, various databases (DB) 22022, a control unit 22030, a memory 22031, a non-volatile memory 22032, a communication unit 22035, an AI calculation unit 22040, an LLM processing unit 22041, various AI processing units 22042, a cooling unit 22050, and so on.

[0053] Specifically, user interface 22010 is a user interface for server users, such as administrators or maintenance personnel, to operate the large-scale language model server 20001. User interface 22010 includes, for example, an input interface 22011. Input interface 22011 is an interface for inputting operation inputs and control information for users to operate or control the server. For example, input interface 22011 may be a physical operation button, an operation input panel such as a touch panel, or a remote controller. Alternatively, or in addition to the above, an interface may be provided for connecting to an information terminal owned by the user and inputting control information from that information terminal. Specifically, a communication interface such as a USB interface using a Universal Serial Bus or an Ethernet interface may be provided. User interface 22010 also includes an output interface 22012. Output interface 22012 is an interface for outputting information from the server to a user, such as an administrator. For example, the information to be output may include the operating status of the server, such as the load, whether there are any errors, and the type of error if there are any. Furthermore, the information to be output may also be information related to the AI ​​calculation unit 22040. This information may include the type of large-scale language model being inferred by the LLM processing unit 22041, or information such as the execution status of the inference, including the progress rate. For example, the output interface 22012 may be an interface that connects to an information terminal owned by the user and outputs information to that terminal. Specifically, this may be a communication interface such as a USB interface or an Ethernet interface. Alternatively, or in addition to this, the output interface 22012 may be provided with a display unit. This allows information to be transmitted to the user via a displayed image or video. Note that if the user interface 22010 is a communication interface, it may share hardware with the communication unit 22035, which will be described later.

[0054] The power supply 22015 converts the AC current input from an external source via the external power input interface 22014 into DC current and supplies the necessary DC current to each part of the large-scale language model server 20001.

[0055] The storage unit 22020 is a storage device that records various types of information, such as image data, video data, audio data, and text data. This information may be stored in a database (DB) structure. It may also be composed of magnetic recording media such as hard disk drives (HDDs) or semiconductor memory such as solid-state drives (SSDs). In the example shown in this figure, the storage unit 22020 is configured within the large-scale language model server 20001. However, the storage unit 22020 may also be located outside the large-scale language model server 20001 and connected to the large-scale language model server 20001 via a communication interface. In the example shown in this figure, the storage unit 22020 includes various DBs 22021. The information stored in the various DBs 22021 may be output to, for example, the AI ​​calculation unit 22040, which will be described later. In this case, the AI ​​calculation unit 22040 can use the information stored in the various DBs 22021 for inference. Furthermore, the storage unit 22020 may record data such as image data, video data, audio data, and text data, which are the results of the inference performed by the AI ​​calculation unit 22040, in various DBs 22021. The information stored in the various DBs 22021 may also be output as a response to access from external devices connected via the communication unit 22035 and the Internet 19000, as described later. Additionally, the storage unit 22020 may record data such as image data, video data, audio data, and text data transmitted from external devices connected via the communication unit 22035 and the Internet 19000 in various DBs 22021.

[0056] Furthermore, when training an AI model, such as a large-scale language model used in the AI ​​computation unit 22040 described later, is performed on the large-scale language model server 20001, the storage unit 22020 may store a training data DB 22022, which is a database of training data, such as image data, video data, audio data, and text data, that will be the target data for the training.

[0057] The control unit 22030 controls the operation of each part within the large-scale language model server 20001. The control unit 22030 may also work in cooperation with a program stored in memory 22031 to perform calculations based on information obtained from each part within the large-scale language model server 20001. Control states by the control unit 22030 include receiving various control information and instructions from external devices via the communication unit 22035 and the internet 19000, and using this control information and instructions to cause the AI ​​calculation unit 22040 to perform inference of the large-scale language model. Furthermore, control states by the control unit 22030 include, for example, outputting the output from the large-scale language model of the AI ​​calculation unit 22040 (described later) to external devices via the communication unit 22035 and the internet 19000. The specific configuration of the control unit 22030 is a CPU, and may also be referred to as a processor or control circuit.

[0058] The non-volatile memory 22032 stores various data used by the large-scale language model server 20001. The data stored in the non-volatile memory 22032 includes, for example, various data output from the output interface 22012 of the large-scale language model server 20001.

[0059] Memory 22031 stores control data and other information used by the control unit 22030. The control unit 22030 may read various software / applications from the storage unit 22020, expand them into memory 22031, and store them there.

[0060] The communication unit 22035 is a communication interface such as Ethernet and connects to the internet 19000 via a router, switch, gateway device, etc. The communication path from the communication unit 22035 to the internet 19000 may include both wired and wireless sections, and may also pass through repeaters. If the communication unit 22035 is an Ethernet interface, communication may be performed using a LAN communication method up to a designated router, switch, or gateway device. This allows the communication unit 22035 to communicate with the communication unit 1132 of the artificial intelligence response output device 10010 and other devices such as servers via the internet 19000. The communication unit 22035 communicates with the communication unit 1132 of the artificial intelligence response output device 10010 and can send and receive instructions, control information, and artificial intelligence responses between the AI ​​calculation unit 22040 (described later) and the artificial intelligence response output device 10010 using an API.

[0061] Next, the AI ​​computing unit 22040 is a computing unit that performs inference on AI models such as neural networks. Specifically, it is a processor or circuit such as a GPU (Graphics Processing Unit). It may also be configured with multiple cores, with the GPU and memory for the GPU acting as cores. Another example of the hardware for the AI ​​computing unit 22040 is a so-called NPU (Neural Network Processing Unit). The AI ​​model on which the AI ​​computing unit 22040 performs inference is, for example, a neural network of various deep learning types. Specifically, it may be an LLM (Large Language Model), or an image recognition processing model, or an speech recognition processing model. The model can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Transformer, an Autoencoder, a Generative Adversarial Network (GAN), or a Diffusion Model.

[0062] In the example shown in Figure 1D, the AI ​​calculation unit 22040 includes an AI model of an LLM and an LLM processing unit 22041 that performs inference on the LLM. For example, the LLM on which the LLM processing unit 22041 performs inference is an AI model that has learned at least text information. Alternatively, the LLM on which the LLM processing unit 22041 performs inference may be a multimodal LLM that has learned not only text data but also media other than text data. Such media other than text data may be, for example, image data, video data, or audio data. The LLM processing unit 22041 obtains instruction statements and control information from external devices such as an artificial intelligence response output device 10010 connected to the Internet 19000 via communication via a communication unit 22035 using an API, and performs inference on the LLM based on this information. The result of the inference is output to the external device as an artificial intelligence response via communication via the communication unit 22035 using an API. This series of processes can be controlled, for example, by the control unit 22030.

[0063] Furthermore, the AI ​​calculation unit 22040 may include various AI models other than LLM, and may also include various AI processing units 22042 that perform inference on these various AI models. Examples of various AI models other than LLM include clustering AI models that have undergone unsupervised learning, regression prediction AI models that have undergone supervised learning, and classification AI models that have undergone supervised learning. The various AI processing units 22042 acquire control information from external devices such as artificial intelligence response output devices 10010 connected to the Internet 19000 via communication via the communication unit 22035 using an API, and perform inference on the various AI models based on this information. The results of the inference are output to the external devices as artificial intelligence responses via communication via the communication unit 22035 using an API. These series of processes can be controlled, for example, by the control unit 22030.

[0064] Furthermore, when training AI models such as large-scale language models on the large-scale language model server 20001, the AI ​​calculation unit 22040 can use image data, video data, audio data, text data, etc., stored in the training data DB 22022 of the storage unit 22020 as training data to train the AI ​​model.

[0065] The cooling unit 22050 is primarily configured to cool the AI ​​computing unit 22040. Training and inference of AI models require a lot of power and generate heat. Therefore, a configuration is needed to forcibly cool the AI ​​computing unit 22040 in addition to natural convection air cooling. Specifically, the cooling unit 22050 may be a forced convection device for air cooling such as a fan, or a forced convection device for liquid cooling such as a pump. Alternatively, the cooling unit 22050 may be a heat exchanger such as a heat pipe.

[0066] As described above using Figure 1D, the server configuration allows the AI ​​calculation unit 22040 to perform inference on an AI model based on control information received via communication through the communication unit 22035, and the result to be output as an artificial intelligence response via communication through the communication unit 22035.

[0067] Note that the server in Figure 1D is described as a single physical server as an example. However, this server may also be a virtualized server using multiple physical servers. This server may also be configured as a cloud server where various functions are distributed and virtualized on the cloud (cloud computing system).

[0068] Furthermore, the control unit 22030 may monitor the operating status of the power supply 22015, the load of the AI ​​calculation unit 22040, and the load of the forced convection device such as the fan or pump of the cooling unit 22050. This operating status information monitored by the control unit 22030 may be output through an output interface 22012, such as the display unit of the user interface 22010. This operating status information monitored by the control unit 22030 may also be output to a network such as the Internet 19000 via the communication unit 22035, and the outputted operating status information may be configured to be obtainable by an external device such as an artificial intelligence response output device 10010 connected to the network.

[0069] Next, using Figure 1E, we will describe an example of deploying and executing a client application and an LLM application in an artificial intelligence response output system. Figure 1E is a schematic diagram focusing on the deployment, execution, and communication of the client application and the LLM application in an artificial intelligence response output system configuration including an artificial intelligence response output device 10010 and a large-scale language model server 20001. For simplicity, other components besides the client application and the LLM application are omitted from the notation and description.

[0070] In the example shown in Figure 1E, the client application 8010 is executed in the artificial intelligence response output device 10010. Specifically, the client application 8010 is loaded into the memory 1109 shown in Figure 1B and executed by the control unit 1110. In addition, the server LLM application 8020 is executed in the large-scale language model server 20001. Specifically, the server LLM application 8020 is loaded into the memory 22031 shown in Figure 1D and executed by the control unit 22030. The server LLM application 8020 may also be an application that controls the entire large-scale language model possessed by the large-scale language model server 20001. For example, the server LLM application 8020 controls the execution of inference of the AI ​​model of the LLM provided in the LLM processing unit 22041 of the AI ​​calculation unit 22040 shown in Figure 1D. The server LLM application 8020 may also be an application that controls the input and output of information to and from the large-scale language model possessed by the large-scale language model server 20001. The client application 8010 and the server LLM application 8020 communicate via, for example, the communication unit 1132 shown in Figure 1B, the communication device 19011 shown in Figure 1A, a network such as the Internet 19000, and the communication unit 22035 shown in Figure 1D. For example, the client application 8010 sends an instruction to the server LLM application 8020. Upon receiving the instruction, the server LLM application 8020 performs control to execute inference of the LLM AI model provided in the LLM processing unit 22041 of the AI ​​calculation unit 22040 shown in Figure 1D. The server LLM application 8020 then sends the artificial intelligence response, which is the result of this inference, to the client application 8010. The client application 8010 can then output an output based on this artificial intelligence response to the user. In this way, in the artificial intelligence response output system of the example in Figure 1E, the client application 8010 and the server LLM application 8020 cooperate to output a more suitable artificial intelligence response to the user. Here, the large-scale language model possessed by the large-scale language model server 20001 may also be referred to as the server's large-scale language model.Server LLM application 8020 may also be called a server-wide large-scale language model application.

[0071] In the example shown in Figure 1E, the local LLM application 8015 is executed in the artificial intelligence response output device 10010. Specifically, the local LLM application 8015 is loaded into the memory 1109 shown in Figure 1B in the artificial intelligence response output device 10010 and executed by the control unit 1110. The local LLM application 8015 may also be an application that controls the entire large-scale language model possessed by the local LLM processing unit 10028 shown in Figure 1B. For example, the local LLM application 8015 controls the execution of inference for the large-scale language model possessed by the local LLM processing unit 10028 shown in Figure 1B. Alternatively, the local LLM application 8015 may also be an application that controls the input and output of information to and from the large-scale language model possessed by the local LLM processing unit 10028. The client application 8010 and the local LLM application 8015 communicate via a communication path, such as a bus within the artificial intelligence response output device 10010. For example, the client application 8010 sends an instruction to the local LLM application 8015. Upon receiving the instruction, the local LLM application 8015 controls the local LLM processing unit 10028, shown in Figure 1B, to perform inference on its large-scale language model. The local LLM application 8015 then sends the artificial intelligence response, which is the result of this inference, to the client application 8010. The client application 8010 can then output to the user based on this artificial intelligence response. In this way, the artificial intelligence response output system in the example shown in Figure 1E allows the client application 8010 and the local LLM application 8015 to work together to output a more suitable artificial intelligence response to the user. Here, the large-scale language model of the local LLM processing unit 10028 may also be called a local large-scale language model. The local LLM application 8015 may also be called a local large-scale language model application.

[0072] Furthermore, in the artificial intelligence response output system shown in the example in Figure 1E, the client application 8010, the local LLM application 8015, and the server LLM application 8020 work together to output a more suitable artificial intelligence response to the user.

[0073] In Figures 1A, 1D, and 1E, an example was shown in which the artificial intelligence response output device 10010 and the large-scale language model server 20001 are connected via the Internet 19000. That is, in Figure 1E, the server LLM application 8020 was connected to the client application 8010 via the Internet 19000. However, the network connecting the artificial intelligence response output device 10010 and the large-scale language model server 20001 is not limited to the Internet 19000, but may also be a so-called LAN (Local Area Network). The network connecting the artificial intelligence response output device 10010 and the large-scale language model server 20001 may also be a so-called local 5G network. In these cases, the artificial intelligence response output device 10010, which has the client application 8010 and the local LLM application 8015, may be on the same subnet mask as the server LLM application 8020 and the large-scale language model server 20001. Thus, an artificial intelligence response system, including an artificial intelligence response output device 10010 and a large-scale language model server 20001, built without using the internet 19000, can be used in local settings such as homes, factories, and hospitals. In this case, even if communication with the internet 19000 is interrupted for any reason, the client application 8010 of the artificial intelligence response output device 10010 can communicate with the server LLM application 8020 of the large-scale language model server 20001 via a LAN or local 5G network, enabling collaboration.

[0074] Here, a more specific configuration example of the artificial intelligence response output device 10010 according to an embodiment of the present invention being an aerial floating image display device will be explained using Figures 1F to 1K.

[0075] The following examples in Figures 1F to 1K relate to an image display device capable of transmitting an image generated by image light from an image light source through a transparent component that partitions a space, such as glass, and displaying it as an aerial floating image outside the transparent component. In the explanation of Figures 1F to 1K, the image floating in the air is referred to as an "aerial floating image." Instead of this term, it is also acceptable to use terms such as "aerial image," "spatial image," "spatial floating image," "aerial floating optical image of a displayed image," or "spatial floating optical image of a displayed image." The term "aerial floating image," which is mainly used in the explanation of the embodiments, is used as a representative example of these terms.

[0076] <Example Configuration of an Aerial Floating Image Display Device 1> First, using Figure 1F, we will explain an example of an optical system used in the artificial intelligence response output device 10010, which is an aerial floating image display device.

[0077] In the optical system shown in Figure 1F, the display device 1 comprises a liquid crystal display panel 11 and a light source device 13. The display device 1 outputs image light of a specific polarization. The surface of the display device 1 may be equipped with an absorbing polarizing plate 12 that transmits the image light of the specific polarization and absorbs the other polarization. By providing the absorbing polarizing plate 12, unwanted reflected light can be reduced, thereby reducing stray light and other unwanted light. The image light of the specific polarization output from the display device 1 is input to the polarization separation member 101B. The polarization separation member 101B is a member that selectively transmits the image light of the specific polarization. The polarization separation member 101B is not integrated with the transparent member 100, but has an independent plate-like shape. Therefore, the polarization separation member 101B may also be described as a polarization separation plate. The polarization separation member 101B may be configured, for example, as a reflective polarizing plate formed by attaching a polarization separation sheet to a transparent member. Alternatively, the transparent member may be formed with a metal multilayer film that selectively transmits specific polarizations and reflects polarizations of other specific polarizations. In Figure 1F, the polarization separation member 101B is configured to transmit image light of a specific polarization output from the display device 1.

[0078] The image light that has passed through the polarization separation member 101B is incident on the retroreflector 2. A λ / 4 plate 21 is provided on the image light incident surface of the retroreflector. The image light is polarized from one polarization to the other by passing through the λ / 4 plate 21 twice, once when it is incident on the retroreflector and once when it is emitted. Here, the polarization separation member 101B has the property of reflecting the polarization of the other polarization that has been polarized by the λ / 4 plate 21, so the image light after polarization conversion is reflected by the polarization separation member 101B. The image light reflected by the polarization separation member 101B passes through the transparent member 100, forming a real image, the floating image 3, on the outside of the transparent member 100.

[0079] Here, we will explain a first example of polarization design in the optical system shown in Figure 1F. For example, the display device 1 may be configured to emit P-polarized (P stands for parallel; polarization in which the electric field oscillates within the incident plane) image light to the polarization separation member 101B, and the polarization separation member 101B may be configured to reflect S-polarized (S stands for senkrecht; polarization in which the electric field oscillates perpendicular to the incident plane) and transmit P-polarized light. In this case, the P-polarized image light that reaches the polarization separation member 101B from the display device 1 passes through the polarization separation member 101B and heads towards the retroreflector 2. When the image light is reflected by the retroreflector 2, it passes through the λ / 4 plate 21 provided on the incident plane of the retroreflector 2 twice, so the image light is converted from P-polarized to S-polarized. The image light converted to S-polarized then heads towards the polarization separation member 101B again. Here, since the polarization separation member 101B has the characteristic of reflecting S-polarized light and transmitting P-polarized light, the S-polarized image light is reflected by the polarization separation member 101 and transmitted through the transparent member 100. Since the image light transmitted through the transparent member 100 is light generated by the retroreflector 2, it forms an aerial levitation image 3, which is an optical image of the display image of the display device 1, at a position that is mirror-image to the display image of the display device 1 with respect to the polarization separation member 101B. With such a polarization design, the aerial levitation image 3 can be suitably formed.

[0080] Next, a second example of polarization design in the optical system shown in Figure 1F will be described. For example, the display device 1 may be configured to emit S-polarized video light to the polarization separation member 101B, and the polarization separation member 101B may be configured to reflect P-polarized light and transmit S-polarized light. In this case, the S-polarized video light that reaches the polarization separation member 101B from the display device 1 passes through the polarization separation member 101B and heads towards the retroreflector 2. When the video light is reflected by the retroreflector 2, it passes through the λ / 4 plate 21 provided on the incident surface of the retroreflector 2 twice, so the video light is converted from S-polarized to P-polarized light. The video light converted to P-polarized light heads towards the polarization separation member 101B again. Here, since the polarization separation member 101B has the characteristic of reflecting P-polarized light and transmitting S-polarized light, the P-polarized video light is reflected by the polarization separation member 101 and transmitted through the transparent member 100. Since the image light transmitted through the transparent member 100 is light generated by the retroreflector 2, it forms a floating image 3, which is an optical image of the display image of the display device 1, at a position that is mirror-image to the display image of the display device 1 with respect to the polarization separation member 101B. With this polarization design, the floating image 3 can be suitably formed.

[0081] In Figure 1F, the image display surface of the display device 1 and the surface of the retroreflector 2 are arranged parallel to each other. The polarization separation member 101B is positioned at an angle α (e.g., 45°) relative to the image display surface of the display device 1 and the surface of the retroreflector 2. As a result, in the reflection by the polarization separation member 101B, the direction of propagation of the image light reflected by the polarization separation member 101B (the direction of the principal ray of the image light) differs from the direction of propagation of the image light incident from the retroreflector 2 (the direction of the principal ray of the image light) by an angle β (e.g., 90°). With this configuration, in the optical system of Figure 1F, image light is output at a predetermined angle shown toward the outside of the transparent member 100, forming the floating image 3, which is a real image. In the configuration of Figure 1F, when a user views from the direction of arrow A, the floating image 3 is visible as a bright image. However, when another person views from the direction of arrow B, the floating image 3 cannot be seen as an image at all. This characteristic makes it ideal for systems that display video requiring high security, or highly confidential video that should be kept hidden from the user.

[0082] Next, Figure 1F(2) shows an example of the surface shape of a typical retroreflector 2. The retroreflector 2 has a prism body in which regularly arranged triangular pyramidal recesses serve as reflective surfaces. Light rays incident on the arranged triangular pyramidal recesses are reflected by multiple reflective surfaces of the triangular pyramidal recesses and emitted as retroreflected light in the direction corresponding to the incident light, and the display device 1 displays a real image of a floating object in the air based on the image displayed on the display device 1.

[0083] The surface shape of the retroreflector in this embodiment is not limited to the examples described above. It may have various surface shapes that realize retroreflection. Specifically, retroreflective elements formed by periodically arranging triangular pyramidal prisms, hexagonal pyramidal prisms, other polygonal prisms, multi-vertex prisms, or combinations thereof may be provided on the surface of the retroreflector in this embodiment. Alternatively, retroreflective elements forming cube corners by periodically arranging these prisms may be provided on the surface of the retroreflector in this embodiment. These can also be expressed as corner reflector arrays or polyhedron reflector arrays. Alternatively, capsule lens type retroreflective elements formed by periodically arranging glass beads may be provided on the surface of the retroreflector in this embodiment. Since the detailed configuration of these retroreflective elements can be described using existing technology, a detailed explanation is omitted. Specifically, the technology disclosed in Japanese Patent Publication No. 2001-33609, Japanese Patent Publication No. 2001-264525, Japanese Patent Publication No. 2005-181555, Japanese Patent Publication No. 2008-70898, Japanese Patent Publication No. 2009-229942, etc., can be used.

[0084] As explained above, the optical system in Figure 1F can form a suitable aerial levitation image.

[0085] Next, Figure 1G shows an example of the configuration of an aerial levitation image display device 1000, which is one embodiment of the artificial intelligence response output device 10010. The aerial levitation image display device 1000 shown in Figure 1G is equipped with an optical system corresponding to the optical system in Figure 1F. The aerial levitation image display device 1000 shown in Figure 1G is installed vertically, for example, so that the side on which the aerial levitation image 3 is formed faces the front of the aerial levitation image display device 1000 (towards the user 230). That is, in Figure 1G, the transparent member 100 of the aerial levitation image display device 1000 is installed on the front of the device (towards the user 230). The aerial levitation image 3 is formed on the user 230 side relative to the surface of the transparent member 100 of the aerial levitation image display device 1000. The light of the aerial levitation image 3 travels in the direction toward the user (the -y direction in the figure). If the aerial operation detection sensor 1351, which will be described later, is provided as shown, it is possible to detect operation of the aerial levitation image 3 by the user 230's finger.

[0086] <Example Configuration of an Aerial Floating Image Display Device 2> Another example of the optical system configuration for the aerial floating image display device will be explained using Figure 1H. The optical system in Figure 1H is an optical system that uses a retroreflector 5, which is different from the retroreflector 2 used in Figure 1F. Components in Figure 1H that are denoted by the same reference numerals as in Figure 1F have the same function and configuration as those in Figure 1F. Such components will not be explained again in order to simplify the explanation.

[0087] Figure 1H shows an example of the main components and retroreflective components of an aerial floating image display device according to one embodiment of the present invention. A display device 1 that emits image light is provided obliquely to a transparent member 100 such as glass. The display device 1 comprises a liquid crystal display panel 11 and a light source device 13 that generates light.

[0088] The principal ray 9020, which represents the light beam emitted from the display device 1, travels toward the retroreflector 5 and is incident on the retroreflector 5 at an incident angle α (here defined as the angle with respect to the plane direction of the retroreflector 5). The incident angle α can be, for example, 45°. However, the incident angle α is not limited to 45°; for example, 45°±15° can also be used.

[0089] The retroreflector 5 is an optical component having optical properties that retroreflect light rays in at least some directions. Furthermore, since the reflected light rays have optical properties that form an image, the retroreflector 5 may also be described as an imaging optical component or imaging optical plate.

[0090] The specific configuration of the retroreflector 5 will be described in detail using Figure 1I, but the retroreflector 5 causes the principal ray 9020 to propagate in the z direction while being retroreflected in the x and y directions. As a result, the reflected ray 9021 travels away from the retroreflector 5 in an optical path that is mirror-symmetric with respect to the principal ray 9020 with respect to the retroreflector 5, passes through the transparent member 100, and forms a floating image 3 as a real image at the imaging plane.

[0091] The light beam forming the floating image 3 is a collection of light rays converging from the retroreflector 5 to the optical image of the floating image 3, and these light rays continue to travel in a straight line even after passing through the optical image of the floating image 3. Therefore, unlike the diffused image formed on a screen by a typical projector, the floating image 3 is an image with high directivity. Thus, in the configuration of Figure 1H, when a user views from the direction of arrow A, the floating image 3 is visible as a bright image. However, when another person views from the direction of arrow B, the floating image 3 cannot be seen as an image at all. This characteristic is suitable for use in systems that display images requiring high security or highly confidential images that should be hidden from people directly facing the user.

[0092] Next, an example of the configuration of the retroreflector 5 will be described using Figure 1I. The retroreflector 5 has a configuration in which multiple corner reflectors 9040 are arranged in an array on the surface of a transparent material. This may also be called a corner reflector array or a multifaceted reflector array. Light rays 9111, 9112, 9113, and 9114 emitted from the light source 9110 are reflected twice by the two mirror surfaces 9041 and 9042 of the corner reflectors 9040, becoming reflected light rays 9121, 9122, 9123, and 9124. This double reflection is retroreflection in the x and y directions, where the light is reflected back in the same direction as the incident direction (moving in a direction rotated 180°), and in the z direction, it is specular reflection in which the angle of incidence and the angle of reflection coincide due to total internal reflection.

[0093] In other words, the light rays 9111 to 9114 produce reflected light rays 9121 to 9124 on a straight line symmetrical in the z direction with respect to the corner reflector 9040, forming the aerial real image 9120. The light rays 9111 to 9114 emitted from the light source 9110 are four representative rays of diffused light from the light source 9110, and depending on the diffusion characteristics of the light source 9110, the light rays incident on the retroreflector 5 are not limited to these, but any incident light ray will cause similar reflection and form the aerial real image 9120. For the sake of clarity in the diagram, the position of the light source 9110 and the position of the aerial real image 9120 in the x direction are shown offset, but in reality, the position of the light source 9110 and the position of the aerial real image 9120 in the x direction are at the same position, and when viewed from the z direction, they are in overlapping positions.

[0094] In the optical system shown in Figure 1F, the retroreflector 2 has retroreflective properties in three axes. As a result, when a diffuse incident light beam is incident on the retroreflector 2, a convergent reflected light beam travels toward the side of the incident light beam where the light source is located relative to the retroreflector 2. This convergent reflected light beam forms an image in the air, creating a floating image 3. The direction of propagation of the principal ray of the convergent reflected light beam reflected from the retroreflector 2 is opposite to the direction of propagation of the principal ray of the diffuse incident light beam incident on the retroreflector 2.

[0095] In contrast, in the optical system shown in Figure 1H, the retroreflector 5 has retroreflective properties in two axes and specular reflection in the other axis. As a result, when a diffuse incident light beam is incident on the retroreflector 5, the convergent reflected light beam reflected by the corner reflector array travels toward the retroreflector 5 toward the side of the incident light beam away from the light source. This convergent reflected light beam forms an image in the air, creating a floating image 3.

[0096] The direction of propagation of the principal ray of the convergent reflected light beam reflected by the corner reflector array of the retroreflector 5 is not in the opposite direction to the direction of propagation of the principal ray of the diffuse incident light beam incident on the retroreflector 5. The component of the direction of propagation of the principal ray of the diffuse incident light beam incident on the retroreflector 5 in the direction of the plate-shaped surface of the retroreflector 5, and the component of the direction of propagation of the principal ray after it has been reflected by the retroreflector 5 and become a convergent reflected light beam, remain in a straight line before and after reflection by the corner reflector array.

[0097] In other words, the diffusive incident light beam is converted into a convergent reflected light beam by reflection at the retroreflector 5, but in the direction normal to the plate-shaped surface of the retroreflector 5, the light beam will travel through the retroreflector 5. Here, the diffusive incident light beam that enters the retroreflector 5 and the convergent reflected light beam that exits the retroreflector 5 are geometrically symmetrical with respect to the plate-shaped surface of the retroreflector 5.

[0098] The shape of the retroreflector (imaging optical plate) in the optical system shown in Figure 1H is not limited to the example described above. It may have various shapes that realize retroreflection. Specifically, it may be various cubic corner bodies, corner reflector arrays, slit mirror arrays, two-sided corner reflector arrays, multi-sided reflector arrays, or a shape in which combinations of their reflective surfaces are arranged periodically. Alternatively, a capsule lens type retroreflector element with glass beads arranged periodically may be provided on the surface of the retroreflector in this embodiment. The detailed configuration of these retroreflector elements can be described using existing technology, so a detailed explanation is omitted. Specifically, the technology disclosed in Japanese Patent Publication No. 2017-33005, Japanese Patent Publication No. 2019-133110, Japanese Patent Publication No. 2017-67933, WO2009 / 131128, etc., can be used.

[0099] Next, Figure 1J shows an example of another configuration of the aerial floating image display device 1000, which is one embodiment of the artificial intelligence response output device 10010. The aerial floating image display device 1000 shown in Figure 1J is equipped with an optical system corresponding to the optical system in Figure 1H. In the aerial floating image display device 1000 shown in Figure 1J, it is installed horizontally so that the side on which the aerial floating image 3 is formed faces upward.

[0100] In other words, in Figure 1J, the aerial levitation image display device 1000 has a transparent member 100 installed on its upper surface. The aerial levitation image 3 is formed above the surface of the transparent member 100 of the aerial levitation image display device 1000. The light of the aerial levitation image 3 travels diagonally upward. If the aerial operation detection sensor 1351, which will be described later, is provided as shown, it is possible to detect operation of the aerial levitation image 3 by the user 230's finger.

[0101] In Figure 1G, the display device 1 and the floating image 3 are symmetrical with respect to the plane of the polarization separation member 101. In contrast, in Figure 1J, the display device 1 and the floating image 3 are symmetrical with respect to the plane of the retroreflector 5. Also, the configuration in Figure 1G includes a retroreflector 2 and a λ / 4 plate 21, but these are not present in Figure 1J. Furthermore, in Figure 1G, the presence of an absorptive polarizer 12 is preferable, but in Figure 1J, the absorptive polarizer 12 is not particularly necessary.

[0102] <<Block diagram of the internal structure of the aerial levitation video display device>> Next, a block diagram of the internal configuration of the aerial floating image display device 1000, as described in Figure 1G or Figure 1J, will be explained. Figure 1K is a block diagram showing an example of the configuration of the aerial floating image display device 1000. In the embodiment of the present invention, the aerial floating image display device 1000 is one aspect of the artificial intelligence response output device 10010. Therefore, the configuration example of the aerial floating image display device 1000 in Figure 1K is a configuration example in which some configurations have been changed or added compared to the configuration example of the artificial intelligence response output device 10010 in Figure 1B. In Figure 1K, components that are denoted by the same reference numerals as in Figure 1B are the same as those in Figure 1B, so repeated explanations will be omitted.

[0103] The aerial levitation image display device 1000 includes, in place of or in addition to, the display unit 10011 in Figure 1B, a retroreflective unit 1101, an image display unit 1102, a light guide 1104, and a light source 1105. The aerial levitation image display device 1000 includes, in place of or in addition to, the operation input unit 1107 in Figure 1B, an aerial operation detection sensor 1351 and an aerial operation detection unit 1350. The aerial levitation image display device 1000 also includes a power supply 1106, an external power input interface 1111, a control unit 1110, an imaging unit 1180, etc. Furthermore, it may also include a secondary battery 1112, etc. The various processing units 9999 have the configuration already described in Figure 1B, and a detailed, repeated explanation is omitted in this figure. Specifically, these include memory 1109, non-volatile memory 1108, storage unit 1170, video control unit 1160, attitude sensor 1113, position sensor 1114, local LLM processing unit 10028, communication unit 1132, audio output unit 1140, microphone 1139, video signal input unit 1131, audio signal input unit 1133, etc.

[0104] Each component of the aerial floating image display device 1000 is arranged in the housing 1190. Note that the imaging unit 1180 and the aerial operation detection sensor 1351 shown in Figure 1K may be provided on the outside of the housing 1190.

[0105] The retroreflective section 1101 in Figure 1K corresponds to the retroreflective plate 2 in Figure 1F. The retroreflective section 1101 retroreflectively reflects light modulated by the image display section 1102. Of the reflected light from the retroreflective section 1101, the light output to the outside of the floating image display device 1000 forms the floating image 3. When the optical system in Figure 1H is applied, the retroreflective section 1101 corresponds to the retroreflective plate 5 in Figure 1H.

[0106] The video display unit 1102 in Figure 1K corresponds to the liquid crystal display panel 11 in Figures 1F, 1G, 1H, and 1J. The light source 1105 in Figure 1K corresponds to the light source device 13 in Figures 1F, 1G, 1H, and 1J. The video display unit 1102, light guide 1104, and light source 1105 in Figure 1K correspond to the display device 1 in Figures 1F, 1G, 1H, and 1J.

[0107] The video display unit 1102 is a display unit that generates an image by modulating transmitted light based on a video signal input under control by the video control unit 1160, which will be described later. For example, a transmissive liquid crystal panel is used as the video display unit 1102 (the liquid crystal display panel 11 described above), but it is not limited to this. Alternatively, for example, a reflective liquid crystal panel that modulates reflected light or a DMD (Digital Micromirror Device: registered trademark) panel may be used as the video display unit 1102.

[0108] The light source 1105 generates light for the image display unit 1102 and is a solid-state light source such as an LED (Light Emitting Diode) or a laser light source. The power supply 1106 converts the AC current input from an external source via the external power input interface 1111 into DC current and supplies power to the light source 1105. The power supply 1106 also supplies the necessary DC current to each part of the levitating image display device 1000. The secondary battery 1112 stores the power supplied from the power supply 1106. The secondary battery 1112 also supplies power to the light source 1105 and other components that require power via the external power input interface 1111 when external power is not supplied. In other words, if the levitating image display device 1000 is equipped with a secondary battery 1112, the user can use the levitating image display device 1000 even when external power is not supplied.

[0109] The light guide 1104 guides the light generated by the light source 1105 and illuminates the image display unit 1102. The combination of the light guide 1104 and the light source 1105 can also be called the backlight of the image display unit 1102. The light guide 1104 may be made mainly of glass. The light guide 1104 may be made mainly of plastic. The light guide 1104 may be made using mirrors.

[0110] The aerial operation detection sensor 1351 is a sensor that detects operation of the floating aerial image 3 by an object such as a user's finger. The aerial operation detection sensor 1351 senses, for example, the area that overlaps with the entire display range of the floating aerial image 3. Alternatively, the aerial operation detection sensor 1351 may sense only the area that overlaps with at least a portion of the display range of the floating aerial image 3.

[0111] Specific examples of the aerial operation detection sensor 1351 include distance sensors using invisible light such as infrared, invisible light lasers, and ultrasonic waves. The aerial operation detection sensor 1351 may also be configured by combining multiple sensors to detect coordinates on a two-dimensional plane. Furthermore, the aerial operation detection sensor 1351 may consist of a Time of Flight (ToF) LiDAR (Light Detection and Ranging) or an image sensor.

[0112] The aerial operation detection sensor 1351 only needs to be able to sense touch operations by a user's finger on an object displayed as an aerial floating image 3. Such sensing can be performed using existing technologies.

[0113] The aerial operation detection unit 1350 acquires a sensing signal from the aerial operation detection sensor 1351 and, based on the sensing signal, calculates whether or not the user's finger has made contact with an object in the aerial floating image 3, and the position where the user's finger and the object made contact (contact position). The aerial operation detection unit 1350 is composed of circuits such as an FPGA (Field Programmable Gate Array). In addition, some functions of the aerial operation detection unit 1350 may be implemented in software, for example, by an aerial operation detection program executed in the control unit 1110 or the video control unit 1160. The aerial operation detection sensor 1351 and the aerial operation detection unit 1350 may be configured as an integrated unit. The aerial operation detection unit 1350 and the control unit 1110 or the video control unit 1160 may be configured as an integrated unit.

[0114] The aerial operation detection sensor 1351 and the aerial operation detection unit 1350 may be built into the aerial floating image display device 1000, or they may be provided separately from the aerial floating image display device 1000. When provided separately from the aerial floating image display device 1000, the aerial operation detection sensor 1351 and the aerial operation detection unit 1350 are configured to transmit information and signals to the aerial floating image display device 1000 via wired or wireless communication lines or video signal transmission lines. This makes it possible to construct a system in which the aerial floating image display device 1000, which does not have an aerial operation detection function, is used as the main unit, and only the aerial operation detection function can be added as an option.

[0115] Alternatively, the aerial operation detection sensor 1351 may be a separate unit, with the aerial operation detection unit 1350 built into the aerial floating image display device 1000. A configuration in which only the aerial operation detection sensor 1351 is a separate unit offers advantages, such as when it is desired to have more freedom in positioning the aerial operation detection sensor 1351 relative to the installation location of the aerial floating image display device 1000.

[0116] The imaging unit 1180 is, for example, a camera with an image sensor, and captures images of the space near the aerial floating image 3, and / or the face, arms, fingers, etc., of the user 230. Multiple imaging units 1180 may be provided. For example, the imaging units 1180 may be provided as a stereo camera. By using multiple imaging units 1180, or by using an imaging unit with a depth sensor, the aerial operation detection unit 1350 can be assisted when detecting touch operations on the aerial floating image 3 by the user 230. The imaging unit 1180 may be provided separately from the aerial floating image display device 1000. If the imaging unit 1180 is provided separately from the aerial floating image display device 1000, it should be configured to transmit imaging signals to the aerial floating image display device 1000 via a wired or wireless communication connection path.

[0117] For example, if the aerial operation detection sensor 1351 is configured as an object intrusion sensor that detects whether or not an object has entered a plane (intrusion detection plane) that includes the display surface (display range) of the aerial floating image 3, the aerial operation detection sensor 1351 may not be able to detect information such as how far away an object that has not entered the intrusion detection plane (for example, a user's finger) is from the intrusion detection plane, or how close an object is to the intrusion detection plane.

[0118] In such cases, the distance between the object and the intrusion detection plane (floating image 3) can be calculated by using information such as depth calculation information of the object based on images captured by multiple imaging units 1180 and depth information of the object from a depth sensor. This various information, including depth calculation information, depth information, and the distance between the object and the intrusion detection plane, is then used for various display controls of the floating image 3.

[0119] Alternatively, instead of using the aerial operation detection sensor 1351, the aerial operation detection unit 1350 may detect touch operations on the aerial floating video 3 by the user 230 based on the image captured by the imaging unit 1180. In this case, the imaging unit 1180 may be referred to as the aerial operation detection sensor.

[0120] Furthermore, the imaging unit 1180 may capture an image of the user operating the floating video 3, and the control unit 1110 or the like may perform user identification processing. In addition, to determine if other people are standing around or behind the user operating the floating video 3 and whether they are peeking at the user's operation of the floating video 3, the imaging unit 1180 may capture an area that includes the user operating the floating video 3 and the area surrounding the user.

[0121] In addition to the configuration and operation described in Figure 1B, the control unit 1110 also controls the various parts that realize the aerial floating image display function and the aerial operation detection function, which are newly explained using Figure 1K.

[0122] As described above, the various processing units 9999 have the configuration already explained in Figure 1B, so a repeated explanation will be omitted.

[0123] As explained above, the aerial floating image display device 1000 is equipped with various functions. However, the aerial floating image display device 1000 does not need to have all of these functions; any configuration is acceptable as long as it has the function of forming the aerial floating image 3.

[0124] As described above using Figures 1F to 1K, the aerial floating image display device 1000 of this embodiment can be equipped with an aerial floating image display function. In other words, an aerial floating image display device equipped with an artificial intelligence response output function can be realized.

[0125] <Example 2> Next, as Embodiment 2 of the present invention, we will describe an example in which the artificial intelligence response output device 10010 described in Embodiment 1 is connected to the internet and operates by connecting to a server equipped with a large-scale language model artificial intelligence via the internet. In this embodiment, we will explain the differences from Embodiment 1, and repeating explanations of configurations similar to those in these embodiments will be omitted.

[0126] Using Figure 2A, an example of the connection state between the artificial intelligence response output device 10010 and the large-scale language model server 20001 of Embodiment 2 of the present invention will be described. The artificial intelligence response output device 10010 according to Embodiment 2 has the function of displaying a character or avatar, and may be called a character conversation device or an avatar display device. Furthermore, the system including the artificial intelligence response output device 10010 and the large-scale language model server 20001 according to Embodiment 2 may be called a character conversation system or an avatar conversation system. Here, the display unit 10011 displayed by the artificial intelligence response output device 10010 displays an image of character 19051. The image of character 19051 is an image generated by rendering a 3D model of the character in a virtual space. Note that since repeatedly writing "character or avatar" is redundant, in the following description of this embodiment, it will simply be written as "character".

[0127] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can become an artificial intelligence (AI) assistant for the user. In this case, the character conversation device or character conversation system in this embodiment may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.

[0128] In the example shown in Figure 2A, the voice output unit 1140 of the artificial intelligence response output device 10010 is composed of a speaker. The artificial intelligence response output device 10010 is also equipped with a microphone 1139 that can pick up the user's voice. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the internet 19000 via a communication unit 1132. In the example shown in Figure 2A, the communication between the communication unit 1132 and the communication device 19011 is shown as wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the internet 19000 may have both wired and wireless sections. The artificial intelligence response output device 10010 can communicate with the large-scale language model server 20001 via the communication device 19011 and the internet 19000. Furthermore, the artificial intelligence response output device 10010 can communicate with a second server 19002, which is different from the large-scale language model server 20001, via the communication device 19011 and the internet 19000. The configuration including the artificial intelligence response output device 10010 and the large-scale language model server 20001 may be considered as a single system.

[0129] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence. For example, it is a server with a multimodal large-scale language model that can process not only natural language text information but also other types of information.

[0130] The artificial intelligence response output device 10010 can communicate with the large-scale language model of the large-scale language model server 20001 via the internet 19000 using an API.

[0131] The artificial intelligence response output system of Example 2 includes a mobile information processing terminal 20010 used by user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.

[0132] Here, an example of a mobile information processing terminal 20010 will be described using Figure 2B. The mobile information processing terminal 20010 includes a display panel 20011 which is a touch operation input panel, a control unit 20012, memory 20026, non-volatile memory 20027, an external power input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, an attitude sensor 20018, a position sensor 20019, a local LLM processing unit 20028, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.

[0133] The display panel 20011 is equipped with a touch input sensor and can accept touch input from the user 230's finger. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel and can display images. The display panel 20011 may also be called a display unit.

[0134] The communication unit 20020 can be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the artificial intelligence response output device 10010. The control unit 20012 controls the communication unit 20020. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with the communication device 19011 connected to the Internet 19000. As a result, the mobile information processing terminal 20010 can communicate with various servers connected to the Internet 19000.

[0135] Power supply 20014 converts AC current input from an external source via the external power input interface 20013 into DC current and supplies the necessary DC current to each part of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied by power supply 20014. In addition, the secondary battery 20015 supplies power to each part that requires power via the external power input interface 20013 when external power is not supplied.

[0136] The video signal input section 20023 receives video data by connecting an external video output device. Various digital video input interfaces are possible for the video signal input section 20023. For example, it can be configured with an HDMI (High-Definition Multimedia Interface) standard video input interface, a DVI (Digital Visual Interface) standard video input interface, or a DisplayPort standard video input interface. Alternatively, analog video input interfaces such as analog RGB or composite video may be provided. The video signal input section 20023 may also use various USB interfaces.

[0137] The audio signal input unit 20024 receives audio data by connecting an external audio output device. The audio signal input unit 20024 may be configured as an HDMI audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface, etc. The audio signal input unit 20024 may also be various USB interfaces, etc. In the case of an HDMI interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface with integrated terminals and cables.

[0138] The audio output unit 20021 is capable of outputting audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 is also capable of outputting audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured as a speaker. In addition, the audio output unit 20021 may output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output as a digital signal to an external device, such as the Audio Return Channel function specified in the HDMI standard.

[0139] Microphone 20022 is a microphone that picks up sounds from the surrounding area of ​​the mobile information processing terminal 20010, converts them into signals, and generates audio signals. The microphone may be configured to record human voices, such as the user's voice, and the control unit 20012, described later, may perform speech recognition processing on the generated audio signal to obtain text information from the audio signal.

[0140] The imaging unit 20025 is a camera having an image sensor. The camera may be provided on the front of the display panel 20011 side of the mobile information processing terminal 20010, or on the back of the display panel 20011 side. Both a front camera and a rear camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a rear camera.

[0141] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be composed of a magnetic recording medium such as a hard disk drive (HDD) or a semiconductor memory such as a solid-state drive (SSD). For example, the storage unit 20016 may have various types of information, such as video data, image data, and audio data, pre-recorded in it at the time of product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices or external servers via the communication unit 20020. The video data, image data, etc., recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc., recorded in the storage unit 20016 may also be output to external devices or external servers via the communication unit 20020.

[0142] The video control unit 20017 performs various controls related to the video signals input to the display panel 20011. The video control unit 20017 may also be called a video processing circuit and may be composed of hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be called a video processing unit or image processing unit. For example, the video control unit 20017 controls video switching, such as determining which video signal to input to the display panel 20011 from among the video signals to be stored in memory 20026 and the video signals (video data) input to the video signal input unit 20023. The video control unit 20017 may also perform control to perform image processing on the video signals input from the video signal input unit 20023 and the video signals to be stored in memory 20026. Examples of image processing include scaling processing, such as enlarging, reducing, and transforming images; brightness adjustment processing, which changes the brightness; contrast adjustment processing, which changes the contrast curve of an image; and retinex processing, which decomposes an image into its light components and changes the weighting of each component.

[0143] The attitude sensor 20018 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each connected part.

[0144] The position sensor 20019 is a sensor that measures the position of the mobile information processing terminal 20010 using GPS (Global Positioning System) or the like. However, if the position of the mobile information processing terminal 20010 can be estimated based on information obtained through communication by the communication unit 20020, the position estimated by that method may be used instead of the measurement result of the position sensor 20019.

[0145] The non-volatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the non-volatile memory 20027 includes, for example, data for various operations displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for user-operated objects, etc. Memory 20026 stores video data and device control data displayed on the display panel 20011. The control unit 20012 may read various software from the storage unit 20016, expand it into memory 20026, and store it there.

[0146] The local LLM processing unit 20028 has memory capable of holding a large-scale language model (LLM) and can perform inference of the LLM based on the control of the control unit 20012. The hardware can be a so-called GPU (Graphics Processing Unit). The local LLM processing unit 20028 may perform training as well as inference. Note that the local LLM processing unit 20028 is not necessarily required if the execution of LLM inference in the local environment of the mobile information processing terminal 20010 is not needed.

[0147] The control unit 20012 controls the operation of each connected component. The control unit 20012 may also work in cooperation with a program stored in memory 20026 to perform calculations based on information acquired from various components within the mobile information processing terminal 20010. The specific configuration of the control unit 20012 may include a CPU, and it may also be referred to as a processor or control circuit.

[0148] Next, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 2 of the present invention will be described using Figure 2C. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Note that in Figure 2C, the illustration of communication paths such as the Internet 19000 shown in Figure 2A has been omitted. In Figure 2C, the user 230 of the artificial intelligence response output device 10010 is also illustrated. Note that the operation or processing of each part of the artificial intelligence response output device 10010 described in Embodiment 2 of the present invention may be controlled by the client application 8010 as described in Figure 1E. Note that the operation or processing of the large-scale language model server 20001 described in Embodiment 2 of the present invention may be controlled by the server LLM application 8020 as described in Figure 1E.

[0149] Here, we will describe an example of the sequence of operations of the artificial intelligence response output device 10010 (the first example of operation). The artificial intelligence response output device 10010 loads the character operation program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character operation program, thereby enabling the various processes described below. The character operation program may be the same program as the client application 8010 described above, or it may be a different program. Furthermore, the character operation program may be a part of the program that constitutes the client application 8010 described above.

[0150] <Example of operation 1> First, the artificial intelligence response output device 10010 is equipped with a microphone 1139. When user 230 speaks to character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. The character operation program executed by the control unit 1110 then extracts the text of the words spoken by user 230 from the audio signal. This text is in natural language. The extraction of the text of the words spoken by user 230 may be performed continuously for all words, or it may be started when the user speaks within a predetermined period following a trigger keyword. For example, the trigger keyword could be when the user says "Hello" followed by the character's name. For example, if character 19051's name is "Koto," then "Hello, Koto!" can be used as the trigger keyword.

[0151] The character operation program of the artificial intelligence response output device 10010 creates an instruction (prompt) based on the text of the words spoken by the user 230, and sends the instruction to the large-scale language model server 20001 using an API. Here, the instruction may be metadata containing information written using notation such as markup format of a markup language using tags, notation using predetermined symbols such as Markdown format, or object notation of a predetermined script such as JSON. The instruction contains natural language text information as the main message. The types of instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001 include setting instruction statements that store instructions such as initial settings, and user instruction statements that reflect instructions from the user. Type identification information that identifies whether the instruction is a setting instruction statement or a user instruction statement may be stored in a part of the instruction statement other than the main message. When the character motion program of the artificial intelligence response output device 10010 creates an instruction sentence (prompt) based on the text of the words spoken by the user 230, it creates a user instruction sentence and sends it to the large-scale language model server 20001.

[0152] Next, the large-scale language model of the artificial intelligence in the large-scale language model server 20001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response containing natural language text information based on the result. The large-scale language model server 20001 sends the response to the artificial intelligence response output device 10010 using an API. The response contains natural language text information as the main message. Here, the response may also be metadata containing information written in the same format as the instruction mentioned above (notation using tags such as the markup format of a markup language, notation using predetermined symbols such as the Markdown format, or object notation of a predetermined script such as JSON). If the same format as the instruction mentioned above is used in the response, type identification information may be stored in a part other than the main message to indicate that it is a different type of information from the setting instruction and the user instruction mentioned above. For example, information indicating that it is a response from the large-scale language model may be stored.

[0153] Next, the artificial intelligence response output device 10010 receives a response from the large-scale language model server 20001 and extracts the natural language text information stored as the main message in that response. Based on the natural language text information extracted from the aforementioned response, the character operation program of the artificial intelligence response output device 10010 uses speech synthesis technology to generate natural language speech that serves as a response to the user, and outputs it from the speaker-like speech output unit 1140 so that it sounds as if it were the voice of character 20051. This process may also be described as the character "uttering".

[0154] According to the first example of operation described above, using the artificial intelligence response output device 10010 shown in Figure 2C, or the artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001, the advanced natural language processing capabilities of the large-scale language model can be utilized via the API, enabling the character to provide a more appropriate response and engage in a more suitable conversation when the user speaks to it.

[0155] Next, we will explain another example (a second example of operation) of the sequence of operations of the artificial intelligence response output device 10010.

[0156] <Second example of operation> In the first operational example, the actions performed by user 230 to the artificial intelligence response output device 10010 were mainly voice calls from user 230. In the first operational example, a series of operations were performed starting with the process of capturing user 230's voice with a microphone. In contrast, in the second operational example, in addition to the series of operations performed by the artificial intelligence response output device 10010 starting with the process of capturing user 230's voice with a microphone, user 230 can also perform actions to the artificial intelligence response output device 10010 through user operation via the operation input unit 1107 in Figure 1B. Here, examples of the operation input unit 1107 in Figure 1B include a mouse, keyboard, and touch panel.

[0157] Furthermore, in the second example of operation, user 230 can perform an action on the artificial intelligence response output device 10010 based on user touch operations that can be detected by the touch operation input sensor of the display unit 10011 in Figure 1B.

[0158] Furthermore, in the second example of operation, user 230 can also input user 230's operation input to the artificial intelligence response output device 10010 by operating the mobile information processing terminal 20010 and communicating from the mobile information processing terminal 20010 to the artificial intelligence response output device 10010.

[0159] Alternatively, the mobile information processing terminal 20010 may display an information-storage image, such as a two-dimensional code containing information that the user wants to convey to the artificial intelligence response output device 10010, on its display panel 20011, and the imaging unit 1180 of the artificial intelligence response output device 10010 (Figure 1B) may capture this display. The control unit 1110 of the artificial intelligence response output device 10010 may extract information from the information-storage image, such as a two-dimensional code, captured by the imaging unit 1180, and obtain the information. Alternatively, the mobile information processing terminal 20010 may display an image that the user wants to convey to the artificial intelligence response output device 10010 on its display panel 20011, and the imaging unit 1180 of the artificial intelligence response output device 10010 (Figure 1B) may capture this display. The control unit 1110 of the artificial intelligence response output device 10010 may perform image recognition processing on the image captured by the imaging unit 1180 and obtain the result of the image recognition processing.

[0160] Thus, in the second operational example, the types of actions that the user 230 can perform on the artificial intelligence response output device 10010 are greater than those described in the first operational example. As a result, the second operational example can acquire the results of actions performed by the user 230 other than the user's voice and generate an instruction sentence (prompt) to send to the large-scale language model server 20001 based on these results. This makes it possible to more favorably include types of information other than natural language text information extracted from the user's voice in the instruction sentence sent to the large-scale language model server 20001. Examples of types of information other than natural language text information extracted from the user's voice include images, videos, and audio.

[0161] Next, in the second operational example, the artificial intelligence response output device 10010 sends an instruction to the large-scale language model server 20001 using an API. In this operational example as well, the instruction may be metadata containing information written using notation such as markup format of a markup language using tags, notation such as Markdown format using predetermined symbols, or object notation of a predetermined script such as JSON. In this operational example as well, there are two types of instruction: setting instruction statements that store instructions such as initial settings, and user instruction statements that reflect instructions from the user. Type identification information that identifies whether an instruction is a setting instruction statement or a user instruction statement may be stored in a part of the instruction statement other than the main message. In this case, the instruction statement includes natural language text information as the main message. Furthermore, in this embodiment, in addition to natural language text information, the main message of the instruction statement may include non-natural language information sources such as images, videos, or audio as a type of information other than natural language text information. A specific method for including non-natural language information sources in an instruction statement will be described later.

[0162] The large-scale language model server 20001 has a multimodal large-scale language model that can process non-natural language information sources in addition to natural language text information. The large-scale language model server 20001 receives an instruction sentence from the artificial intelligence response output device 10010. Based on the instruction sentence, the multimodal large-scale language model performs inference and generates a response that includes natural language text information as a result of the inference. Here, since the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.

[0163] The artificial intelligence response output device 10010 receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or audio stored as the main message in the response. The character operation program of the artificial intelligence response output device 10010 may use speech synthesis technology to generate natural language audio that serves as a response to the user based on the natural language text information extracted from the aforementioned response, and output it from the audio output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.

[0164] Furthermore, the character operation program of the artificial intelligence response output device 10010 may display natural language characters that serve as a response to the user on the display screen of the artificial intelligence response output device 10010, based on the natural language text information extracted from the aforementioned response. In this case, the characters may be displayed together with character 19051, superimposed on the image of character 19051, or displayed in place of the image of character 19051. The video control unit 1160 may perform these specific processes.

[0165] Furthermore, the character operation program of the artificial intelligence response output device 10010 may display an image on the display screen of the artificial intelligence response output device 10010 in order to present it to the user, based on the image information of the non-natural language information source extracted from the aforementioned response. In this case, the image may be displayed together with the character 19051, superimposed on the image of the character 19051, or displayed in place of the image of the character 19051. These specific processes can be executed by the image control unit 1160.

[0166] Furthermore, the character operation program of the artificial intelligence response output device 10010 may display the video information of the non-natural language information source extracted from the aforementioned response on the display screen of the artificial intelligence response output device 10010 in order to present it to the user. In this case, the video may be displayed together with the character 19051, superimposed on the video of the character 19051, or displayed in place of the video of the character 19051. These specific processes can be executed by the video control unit 1160.

[0167] Furthermore, the character operation program of the artificial intelligence response output device 10010 may output speech generated based on the speech information of the non-natural language information source extracted from the aforementioned response from the speech output unit 1140, which is a speaker.

[0168] According to the second operational example of the artificial intelligence response output device 10010 shown in Figure 2C, or the artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001, as described above, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via the API. In addition to responses based on natural language text, responses based on non-natural language information sources can be provided in response to user actions directed at the character, enabling more appropriate conversations.

[0169] <Example of operation> Next, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 2 of the present invention will be described using Figure 2D. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 2D shows an example of natural language text and non-natural language information sources such as images for the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, and an example of natural language text and non-natural language information sources such as images for the main message of the server response that is the response to it. In this embodiment, non-natural language information sources can include images, videos, audio, etc., but Figure 2D shows an example of an image as a non-natural language information source. The series of operations or processes described using this figure may be controlled by the cooperation of the client application 8010 and the server LLM application 8020, as described in Figure 1E.

[0170] Furthermore, Figure 2D shows the exchange of instructions and responses in chronological order, from the first round of setting instructions and user instructions and their responses to the second round of user instructions and their responses. Here, the instructions and responses shown in the example of Figure 2D include non-natural language information sources 20061 and 20062. In the example of Figure 2D, both non-natural language information sources 20061 and 20062 are images.

[0171] In Figure 2D, for the sake of simplicity, an image of the non-natural language information source 20061 is shown embedded within the instruction text. However, there are multiple methods for transmitting or specifying the data of the non-natural language information source 20061 in the instruction text sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001. The artificial intelligence response output device 10010 may use any one of these methods, or switch between them. An example of each method will be explained below.

[0172] The first method for transmitting or specifying non-natural language information source data in an instruction is used, for example, when the non-natural language information source to be specified is located on a server or other location connected to a network such as the Internet. A specific example of the first method is to specify a non-natural language information source file located on a network such as the Internet using information such as tags and symbols within the instruction, along with the network location information (so-called URL, etc.) and file name.

[0173] For example, a tag used to specify an image in a markup language. <img src=""****”"> By using this tag and writing the location and filename information of the image file in the **** part, you can specify an image that exists on a network such as the internet. Alternatively, you can use a tag that specifies a video in a markup language. <video src=""****”">You can also specify a video that exists on a network such as the internet by using the **** part and writing the location information and file name information of the video file. Alternatively, you can use a tag that specifies audio in a markup language. <audio src=""****”">You can specify audio files located on a network such as the internet by using the format and writing the location and filename information of the audio file in the **** section. Alternatively, if using JSON notation, you can specify images located on a network such as the internet by preparing a key such as img_src and writing the location and filename information of the image file as the value. For video and audio files, you just need to prepare the respective keys and values. The example given is just one example, and you may use other proprietary formats. In any case, the information specifying the location and filename information of the non-natural language information source file should be stored in the instruction statement.

[0174] As in the first method, if the instruction statement contains information specifying the location and filename of the non-natural language information source file, it is not necessary to store the data of the non-natural language information source file in the instruction statement. Therefore, the amount of data in the instruction statement can be reduced. In the first method, the large-scale language model server 20001 that receives an instruction statement specifying non-natural language information source data can use the location and filename information of the non-natural language information source file stored in the instruction statement to obtain the non-natural language information source file located on a server or other location connected to a network such as the Internet.

[0175] Here, we will explain how location information and file name information are input when the artificial intelligence response output device 10010 specifies non-natural language information source data in an instruction sentence using the first method. In Figure 2C, we explained that there are types of actions that the user 230 can take with the artificial intelligence response output device 10010 other than the user 230's voice. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, etc., through user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in Figure 1B.

[0176] Furthermore, in the artificial intelligence response output device 10010, the control unit 1110 may work with the memory 1109 to execute a web browser program and display the GUI of the web browser program on the display screen of the artificial intelligence response output device 10010. User operations on the GUI of the web browser program may be received via the operation input unit 1107 (for example, mouse, keyboard, touch panel) or by user touch operations detectable by the touch operation input sensor of the display unit 10011, and non-natural language information source data such as images, videos, and audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction statement. In this case, the web browser program should acquire the location information and file name information of the non-natural language information source data and pass it to the character operation program.

[0177] Alternatively, user 230 may operate the mobile information processing terminal 20010 to communicate with the artificial intelligence response output device 10010, thereby inputting location information such as a URL for specifying non-natural language information source data to the artificial intelligence response output device 10010. Alternatively, as explained in Figure 2C, location information such as a URL for specifying non-natural language information source data, file name information, etc., may be input by displaying an information-storing image such as a two-dimensional code on the display panel 20011 of the mobile information processing terminal 20010, performing image recognition processing on the image captured by the imaging unit 1180 of the artificial intelligence response output device 10010, and obtaining the results of the image recognition processing.

[0178] Furthermore, the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file already exists on a server or other location connected to a network such as the Internet. For example, if it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 1170 of the artificial intelligence response output device 10010 in an instruction statement, the artificial intelligence response output device 10010 may upload the non-natural language information source data to a second server 19002 via the Internet 19000 and include the Internet location information (so-called URL, etc.) and file name of the uploaded non-natural language information source data on the second server 19002 in the instruction statement. In this case, the second server 19002 functions as a so-called intermediate server.

[0179] Similarly, if it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in the instruction text, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the internet location information (so-called URL, etc.) and file name of the non-natural language information source data on the second server 19002 to the artificial intelligence response output device 10010, and the character operation program of the artificial intelligence response output device 10010 may include the acquired internet location information (so-called URL, etc.) and file name of the non-natural language information source data uploaded to the second server 19002 in the instruction text.

[0180] Furthermore, the character operation program of the artificial intelligence response output device 10010 may work in cooperation with the memory 1109 and the storage unit 1170 to construct a media server within the artificial intelligence response output device 10010 that can be accessed from other servers via the internet 19000. In this case, when the artificial intelligence response output device 10010 specifies non-natural language information source data in an instruction statement using the first method, it only needs to store in the instruction statement location information on the internet (such as a URL) indicating the media server constructed within the artificial intelligence response output device 10010 itself, and the file name of the corresponding non-natural language information source data.

[0181] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction is, for example, simply to store (attach) the non-natural language information source data itself in the instruction (prompt) and send it. Generally, non-natural language information source data such as images, videos, and audio are larger in data size than natural language text information. Therefore, in this case, the data size of the instruction (prompt) will be larger than that of the first method. The character operation program of the artificial intelligence response output device 10010 can store the non-natural language information source data that it wants to store (attach) in the instruction (prompt) in memory 1109, and when sending the instruction (prompt), it can store (attach) the data in the instruction (prompt) via the communication unit 1132 and output it to the large-scale language model server 20001. The non-natural language information source data that the character operation program of the artificial intelligence response output device 10010 stores in memory 1109 may be acquired by the communication unit 1132 via the internet 19000, acquired by the communication unit 1132 from the mobile information processing terminal 20010, or read from the storage unit 1170 and stored in memory 1109.

[0182] As described above, the artificial intelligence response output device 10010 can transmit or specify non-natural language information source data using instruction sentences.

[0183] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. As shown in the example in Figure 2D, through the first round of user instructions, it can acquire images of a swimming pool and poolside, which are non-natural language information sources 20061, and natural language text information. As a result of this inference, it can output natural language text information as shown in the figure, in response to the first round of user instructions.

[0184] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example in Figure 2D, in the response to the second round of user instructions, the large-scale language model server 20001 can include the non-natural language information source 20062 generated by the inference of the multimodal large-scale language model in its response and send it to the artificial intelligence response output device 10010. In Figure 2D, the non-natural language information source 20062 is an example of an image in which a circle image 20063 is attached to an image of a swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 2D, but may also be a video or audio.

[0185] When the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, the method can be the first method or a method similar to the second method used by the artificial intelligence response output device 10010 described above to transmit or specify non-natural language information source data in the instruction statement.

[0186] Specifically, in a method similar to the first method described above, the large-scale language model server 20001 may store information specifying the location and file name of the non-natural language information source file in the instruction statement in its response. The non-natural language information source 20062, such as images, videos, and audio, may be kept by the large-scale language model server 20001, or it may be transferred to and kept by the second server 19002, which functions as an intermediate server. In either case, the large-scale language model server 20001 may store information specifying the location and file name of the non-natural language information source file in the instruction statement in its response. The artificial intelligence response output device 10010, having received the response, may use the location and file name information of the non-natural language information source file described in the instruction statement to access the large-scale language model server 20001 or the second server 19002 to obtain the non-natural language information source 20062.

[0187] Furthermore, specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and send it to the artificial intelligence response output device 10010. The artificial intelligence response output device 10010 can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction and use it for various outputs to the user 230.

[0188] As described above using Figure 2D, the operation of the artificial intelligence response output device 10010 and artificial intelligence response output system of Embodiment 2 involves the transmission and reception of instruction sentences and responses between the character displayed on the artificial intelligence response output device 10010 and the user 230, enabling conversation using natural language information and / or non-natural language information such as images, videos, and audio. This makes it possible to achieve more sophisticated and natural conversations, as shown in each message in Figure 2D.

[0189] <Example of display> Next, an example of the display of the artificial intelligence response output device 10010 of Embodiment 2 of the present invention will be described using Figure 2E. The example in Figure 2E shows an example of displaying the response from the large-scale language model to the user instruction sentences described in Figures 2A to 2D on the display unit 10011 of the artificial intelligence response output device 10010. Specifically, this is an example of displaying the text 10063 of the natural language information source data, the image 10064 of the non-natural language information source data, and / or the video 10065 of the non-natural language information source data, which are the response from the large-scale language model, together with the video of the character 19051 on the display unit 10011. The text 10063, image 10064, and / or video 10065, which are the response from the large-scale language model, may be displayed superimposed in front of the video of the character 19051, as shown in Figure 2E.

[0190] Furthermore, the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on it. The display in Figure 2E is just one example, but for example, if user 230 adjusts the volume of the audio output of the audio output unit 1140 of the artificial intelligence response output device 10010 to the minimum or sets the audio output to OFF by operating via the touch operation input sensor of the operation input unit 1107 or the display unit 10011, user 230 will not be able to confirm the response from the large-scale language model by voice. In this case, the control unit 1110 may control the system to start a display mode in which the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in Figure 2E.

[0191] In this way, even when the user 230 wishes to reduce voice output, the artificial intelligence response output device 10010 can be used more favorably. Furthermore, the user 230 may manually switch ON / OFF the display mode that displays the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, together with the image of character 19051, via the operation input unit 1107 or the touch operation input sensor of the display unit 10011. As shown in the display example in Figure 2E, the multimodal artificial intelligence response output device 10010 can more favorably output responses from the large-scale language model.

[0192] As described above, the artificial intelligence response output device 10010 and artificial intelligence response output system according to Embodiment 2 can provide users with a more advanced conversational experience that includes not only natural language information but also non-natural language information by using a multimodal large-scale language model.

[0193] In the above description of Example 2, an example was described in which the large-scale language model possessed by the large-scale language model server 20001 is used as the large-scale language model. In contrast, the artificial intelligence response output device 10010 may be equipped with a local LLM processing unit 10028 as shown in Figure 1B, and the multimodal large-scale language model possessed by the local LLM processing unit 10028 may be used. In this case, the multimodal large-scale language model possessed by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model possessed by the large-scale language model server 20001. In this case, the operation or processing of the local LLM processing unit 10028 may be controlled by the local LLM application 8015 as described in Figure 1E.

[0194] In this case, in the above description of Example 2, the multimodal large-scale language model possessed by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model possessed by the local LLM processing unit 10028 of the artificial intelligence response output device 10010. In this case as well, by using the multimodal large-scale language model, it is possible to provide the user with a more advanced conversational experience that includes not only natural language information but also non-natural language information.

[0195] <Example 3> Next, Embodiment 3 of the present invention will be described. The artificial intelligence response output device 10010 of Embodiment 3 of the present invention is an artificial intelligence response output device 10010 that has the function of displaying a character or avatar as described in Embodiment 2, and further has the function of switching the character or avatar to be displayed. In this embodiment, the differences from Embodiment 2 will be explained, and the same configuration as in these embodiments will not be described repeatedly. Note that since the repeated notation of "character or avatar" is redundant, in the following description of this embodiment, it will simply be referred to as "character".

[0196] Here, using Figure 3A, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3A shows an example of operation in which the character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 can switch the displayed character based, for example, on operation input input to the operation input unit 1107 or operation detected by the touch operation input sensor of the display unit 10011.

[0197] In the example in Figure 3A, in addition to character 19051 (named "Koto") used in the explanations of Figures 2A to 2E, characters 19052 (named "Tom") and 19053 (named "Necco") are shown. Characters 19051 (named "Koto") and 19052 (named "Tom") are human characters, while character 19053 (named "Necco") is a cat character. Switching the display of characters on the display unit 10011 can be done by rendering different virtual 3D characters for each character and then switching the displayed image on the display unit 10011.

[0198] Furthermore, when the character operation program executed by the control unit 1110 switches the display of the characters shown on the display unit 10011, it is preferable that the synthesized voice used for each character's "utterance" is also changed. This can be done by pre-storing synthesized voice data with corresponding voice to each character in the storage unit 1170, and then performing the synthesized voice change process when switching the display of the characters.

[0199] In the example shown in Figure 3A, the system is configured so that user 230 can converse with any of the characters. The artificial intelligence response output device 10010 in Figure 3A assigns different roles, names, conversational characteristics, or personalities to each of these characters. Furthermore, the memories of each character based on their conversation history are managed separately for each character.

[0200] Therefore, the artificial intelligence response output device 10010 constructs the database shown in Figure 3B in the storage unit 1170, and uses this database to manage character settings and the character's conversation history.

[0201] Next, using Figure 3B, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3B is an explanatory diagram of the database 19200 for managing character settings and character conversation history for multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.

[0202] The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 constructs the database 19200 in the storage unit 1170, for example. The character ID is an identification number that identifies each of the multiple characters that can be displayed by the artificial intelligence response output device 10010, and may be a natural number or use the alphabet, etc. The name is data of the name of each of the multiple characters that can be displayed by the artificial intelligence response output device 10010.

[0203] The configuration instruction is natural language text information that describes the settings of each of the multiple characters that can be displayed by the artificial intelligence response output device 10010, such as their role, name, conversational characteristics, or personality. Since this configuration instruction is the main data of the configuration instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, it is desirable that the content be such that the large-scale language model of the artificial intelligence in the large-scale language model server 20001 can read it directly.

[0204] The conversation history entries, numbered 1, 2, ..., are records of conversations between each character and the user, and are recorded separately for each character. Since this conversation history will be included in the natural language text information, which is the main data of the configuration instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, it is desirable that the content be such that the large-scale language model of the artificial intelligence in the large-scale language model server 20001 can read it directly.

[0205] The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010, when the character displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched, uses the database 19200 in Figure 3B to select and switch the setting instructions and conversation history used for the natural language text information, which is the main data of the setting instructions sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, so that they correspond to the character displayed on the display unit 10011 of the artificial intelligence response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the history of that conversation in the conversation history area of ​​the database 19200 in Figure 3B that corresponds to the character displayed on the display unit 10011.

[0206] By using the database 19200, the character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 uses the same large-scale language model of the same artificial intelligence on the same large-scale language model server 20001 to establish a conversation between the user 230 and the character. However, from the user's perspective, the uniqueness of each character's personality and other settings is preserved, and the memory of different conversations continues for each character. From the user's perspective, this is preferable because it is perceived that the identity of the character's role, name, conversational characteristics, or personality settings and memories from previous conversations is more readily maintained for each character. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.

[0207] Therefore, even when the artificial intelligence response output device 10010 is configured to switch between displaying characters from among multiple character candidates on the display unit 10011, the operation using the database 19200 described above will result in a less jarring experience for the user in conversations with each character, and will allow them to share memories with each of the multiple characters, providing a more enjoyable character conversation experience.

[0208] Furthermore, if the user is prevented from editing the setting instructions for multiple characters, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained in a state close to the intentions of the provider of the artificial intelligence response output device 10010 or the creator of the character's content. Alternatively, the user may be allowed to edit the character setting instructions in response to input from the operation input unit 1107 or the like. In this case, the user can customize the character's role, name, conversational characteristics, or personality to their liking, and converse with a character they have set up themselves. In this case, the character's 3D model, its rendered image, and the type of synthesized voice of the character may also be replaced accordingly.

[0209] Next, an example of a response template database (DB) in the artificial intelligence response output device 10010, which can display multiple characters, as described in Figures 3A and 3B, will be explained using Figure 3C. In the example in Figure 3C, the response template DB 19300 stores the template responses that the artificial intelligence response output device 10010 outputs for each condition assigned a condition number. For these conditions, in the example in Figure 3C, individual response templates are set for each of the multiple characters. For example, for each of the three characters described in Figures 3A and 3B—character 1: Koto, character 2: Tom, and character 3: Necco—a response template for each condition is stored.

[0210] Here, a concrete example of a standard response for character 1: Koto, stored in the response template DB19300 in Figure 3C, and the operation of the artificial intelligence response output device 10010 using it, is as follows.

[0211] First, for example, if the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107 as described above, as in condition number 1, the response should be output using a standard response phrase such as "Good morning" or "Today is [Month] [Day]." The part marked with "[Month] [Day]" can be generated using information stored in the memory 1109 or other location of the artificial intelligence response output device 10010.

[0212] Furthermore, in the example of a standard response phrase in the database shown in Figure 3C, if multiple standard response phrases separated by / are stored, the control unit 1110 can be controlled to randomly select one of the standard response phrases using a random number or the like and output the response. This can resolve and improve the situation where the response becomes monotonous under the same conditions. The explanation for the examples of condition numbers 2, 3, and 4 is the same as for the example of condition number 1. The control unit 1110 should be controlled to output using the standard response phrases shown in Figure 3C for each example of the condition content shown in Figure 3C.

[0213] Next, we will explain an example of condition number 5 shown in Figure 3C. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of user input obtained via the touch panel, microphone 1139, or operation input unit 1107 as natural language, or when there is an obvious grammatical error in the user input, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I may not know about that." By responding in this way, the system can prompt the user to input again and wait for corrected user input.

[0214] Next, we will explain an example of condition number 6 shown in Figure 3C. Condition number 6 is an example where the control unit 1110 detects an error (abnormal state) in any of the parts that make up the artificial intelligence response output device 10010 shown in Figure 1B, and user input is received via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 controls the system to output a response using the standard response phrase "It seems to be malfunctioning." By responding in this way, the system can explain to the user that the artificial intelligence response output device 10010 is malfunctioning and prompt the user to take action to address the error.

[0215] The above describes a specific example of a standard response for character 1: Koto and the operation of the artificial intelligence response output device 10010 using it. However, in the example of the response standard phrase DB 19300 in Figure 3C, standard phrases for other characters such as character 2: Tom or character 3: Necco are also stored.

[0216] In the example shown in Figure 3C, the control unit 1110 selects a corresponding predefined response from the predefined response database 19300 based on the character displayed in the artificial intelligence response output device 10010 and the current conditions, and uses it to control the output as a response from the character. For example, in the example of the predefined response database 19300 in Figure 3C, even under the same conditions, the predefined response is changed to an expression or content that corresponds to the personality of each character. As a result, the artificial intelligence response output device 10010 can provide the user with conversations that correspond to the personality of the displayed character. The user can feel that each character is a being with a more consistent personality. This makes it possible to realize an artificial intelligence response output device 10010 that gives multiple characters a greater sense of reality.

[0217] The artificial intelligence response output device 10010 may output a response using the response template database described in Figure 3C, instead of a response from a large-scale language model such as the local large-scale language model provided by the artificial intelligence response output device 10010, the large-scale language model provided by the large-scale language model server 19001, or the multimodal large-scale language model provided by the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the response template database.

[0218] The response template DB 19300 shown in Figure 3C, as described above, is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 can use it. However, the response template DB 19300 shown in Figure 3C may also be provided on the large-scale language model server 20001 side. In this case, for example, in the configuration of Figure 1D, the response template DB 19300 shown in Figure 3C is stored in the various DBs 22021 of the storage unit 22020, and the control unit 22030 of the large-scale language model server 20001 can generate responses using the response template DB 19300. The control unit 22030 of the large-scale language model server 20001 can then send the response generated using the response template DB 19300 stored in the various DBs 22021 of the storage unit 22020 to the artificial intelligence response output device 10010, instead of the response generated by the large-scale language model. In this way, even if the artificial intelligence response output device 10010 is not equipped with a response template database, it becomes possible to generate responses using a response template database.

[0219] Here, using Figures 3D, 3E, and 3F, we will describe an example of control that has been further improved to allow the user to perceive each character as having a more consistent personality, in the artificial intelligence response output device 10010 and artificial intelligence response output system according to Embodiment 3 of the present invention. Figures 3D, 3E, and 3F describe an example of control regarding the characteristics of character conversations in the artificial intelligence response output system when a client application and an LLM application cooperate to produce a series of artificial intelligence response outputs. Specifically, this is an example in which the response generation process of the artificial intelligence response output device 10010 more preferably combines response generation processing using a large-scale language model on the network or response generation processing using a local large-scale language model (such as the local LLM processing unit 10028) that the artificial intelligence response output device 10010 has internally, and response generation processing using a response template database to generate response output.

[0220] First, an example of the operation of the artificial intelligence response output device and artificial intelligence response output system will be explained using Figure 3D. The configuration of the artificial intelligence response output system in Figure 3D is an improvement over the artificial intelligence response output system shown in Figure 2C. Components with the same reference numerals as in Figure 2C are the same as those in Figure 2C, so repeated explanations will be omitted.

[0221] In the example shown in Figure 3D, as described in Explanation 7010, the artificial intelligence response output device 10010 includes a client application as software that is deployed in, for example, the memory 1109 shown in Figure 1B and executed by the control unit 1110. This corresponds to the client application 8010 in Figure 1E.

[0222] The client application can control the input of user instructions as described in the figures of the embodiments described above. The client application also uses user instructions to send and receive information with the large-scale language model. The client application obtains a response from the large-scale language model. Based on this, the client application can control the parts shown in Figure 1B, including the display unit 10011 and the audio output unit 1140, to output a response to the user. The user instructions can be generated by the client application in response to user input. As an example of a user input interface that accepts user input, similar to the explanation in Figure 2C, an example of the operation input unit 1107 in Figure 1B is a mouse, keyboard, touch panel, etc. The microphone 1139 that picks up the user's voice can also be considered a user input interface. The communication unit 1132 that communicates with the mobile information processing terminal 20010 used by the user, which is a smartphone or tablet information processing terminal, can also be considered a user input interface. Furthermore, the output interface for the client application to output a response to the user includes a display unit 10011 that outputs the response from the large-scale language model in media such as text, images, or video, and an audio output unit 1140 that outputs the response from the large-scale language model in audio.

[0223] Here, the client application can control the insertion of predefined phrases into the responses output by the artificial intelligence response output device 10010. This includes the output control of predefined phrases as explained in Figure 3C. Examples of such predefined phrase insertions include the insertion of predefined greeting phrases as explained in Figure 3C, as well as the insertion of predefined acknowledgment phrases. These predefined phrases can be stored in the storage unit 1170 in Figure 1B of the artificial intelligence response output device 10010, and then loaded into memory 1109 for use by the client application.

[0224] Furthermore, in the example of Figure 3D, as shown in explanation 7020, there is an LLM application on the large-scale language model side that controls the input and output of information to the large-scale language model. As in the example of Figure 3D, when the LLM application is on the large-scale language model server 20001 side, the LLM application is loaded into the memory 22031 of the large-scale language model server 20001 and executed by the control unit 22030 of the large-scale language model server 20001. The LLM application can accept preset instructions to adjust the output of the large-scale language model to specific specifications. While user instructions are instructions whose content changes each time a user inputs, the preset instructions for adjusting the output of the large-scale language model to specific specifications are instructions that are given on a regular basis and whose content does not change with each user input, so these preset instructions may also be called regular instructions. Alternatively, these may be called instructions for customizing the output of the large-scale language model. For example, a preset instruction can be made to set the endings of natural language response sentences generated by the large-scale language model to a predetermined catchphrase. Alternatively, a preset instruction can be made to insert a predetermined interjection into the natural language response sentences generated by the large-scale language model. The preset instructions for the LLM application may be set by sending instructions from the client application of the artificial intelligence response output device 10010 to the LLM application. Alternatively, the preset instructions for the LLM application may be set by communicating with the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user.

[0225] In the example shown in Figure 3D, an LLM application exists on the large-scale language model side that controls the input and output of information to the large-scale language model. This corresponds to the server LLM application 8020 in Figure 1E. In contrast, as another modification, the artificial intelligence response output device 10010 may be configured to have an LLM application that controls the input and output of information to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This corresponds to the local LLM application 8015 in Figure 1E. In this case, the LLM application may be software that is deployed in memory 1109 and executed by the control unit 1110. In this case, the LLM application that controls the input and output of information to the local LLM processing unit 10028 and the client application described above may be different applications, but they may also be the same application.

[0226] In the example shown in Figure 3D, the artificial intelligence response output device 10010 can send not only instruction text but also control information to the large-scale language model server 20001. This control information may be sent together with the instruction text. The transmission of the control information may occur before the transmission of the instruction text. The client application can control the transmission of this control information and instruction text.

[0227] The following describes an example of the details of the control information. First, the control information may contain authentication information for logging into the LLM application. For example, a user may create an account in the LLM application in advance and generate authentication information including identification information and a password. The creation of this account may be performed by communicating with the LLM application via the communication unit 1132 of the artificial intelligence response output device 10010. Alternatively, the creation of this account may be performed by communicating with the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user. The communication unit 1132 of the artificial intelligence response output device 10010 sends the authentication information to the LLM application and performs login through the authentication process. If authentication is successful, the artificial intelligence response output device 10010 and the LLM application establish communication as the user corresponding to the authentication information. As a result, the user can use the information available to their account from the information stored in the memory area of ​​the LLM application. In the case of an LLM application on the large-scale language model server 20001 (server LLM application 8020), the memory area for the LLM application can be provided in the memory 22031 or storage unit 22020 of the large-scale language model server 20001. Furthermore, if the LLM application is a local LLM application (local LLM application 8015) that controls the input and output of information to and from the local LLM processing unit 10028 of the artificial intelligence response output device 10010, the memory area for the LLM application can be provided in the storage unit 1170 or memory 1109 shown in Figure 1B.

[0228] The information available to the user's account includes the preset instructions mentioned above. Furthermore, as shown in Figures 3A, 3B, or 3C, if the artificial intelligence response output device 10010 can switch between and display multiple characters, the LLM application may store information corresponding to each of the multiple characters in its memory area. In this case, if the control information transmitted from the artificial intelligence response output device 10010 to the LLM application includes a character ID to identify the character, the LLM application can determine which character's preset instruction the user wants to apply. In other words, this involves including information on switching preset instructions for each character in the control information, and enabling the LLM application to switch preset instructions according to that switching information. Additionally, if the user wants to set character-specific preset instructions for the LLM application, they can store and transmit the setting information for those preset instructions in the control information transmitted from the artificial intelligence response output device 10010 to the LLM application. Alternatively, the user may transmit the setting information for those preset instructions to the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user. The LLM application, upon receiving the preset instruction setting information, stores the preset instruction information based on that setting information in its memory. As described above, the preset instruction information may be stored for each character. Alternatively, it may be stored for each user account. Once the preset instruction settings are complete, the user transmits a character ID identifying the desired character from the artificial intelligence response output device 10010 to the LLM application, allowing the LLM application to select the preset instruction to apply to that character and apply it to the large-scale language model. In this case, it is not necessary to transmit information equivalent to the preset instruction to the LLM application each time an instruction is sent, thus reducing the amount of communication data.

[0229] Another example is that, each time an instruction is sent, the control information transmitted from the artificial intelligence response output device 10010 to the LLM application may contain information equivalent to a preset instruction and transmit it. In this case, although the transmission frequency of information equivalent to a preset instruction will increase, it will be possible to omit preparations such as prior user account registration and prior setting of preset instructions. Alternatively, each time an instruction is sent, the information equivalent to the preset instruction may be stored in the setting instruction area of ​​the instruction rather than the user instruction area and transmitted. In this case as well, it will be possible to omit preparations such as prior user account registration and prior setting of preset instructions.

[0230] Next, using Figure 3E, we will explain a specific example of how the client application and the LLM application work together to generate response output by more favorably combining the response generation process using a large-scale language model and the response generation process using a response template database.

[0231] Figure 3E is an explanatory diagram illustrating a control example when the artificial intelligence response output system is a character conversation system or an AI assistant system. Furthermore, the example in Figure 3E is an example where the artificial intelligence response output system can switch between and display multiple characters or AI assistants, as shown in Figures 3A, 3B, or 3C. Here, the example in Figure 3E uses two character examples for explanation. Using only two character examples simplifies the explanation; it is possible to switch between and display three or more characters, including the other characters described in previous embodiments. These two characters include a character with character ID 3 and name Necco. This character is the same character described in the previously described embodiments. Additionally, the above two characters include a new character with character ID 4 and name Airia.

[0232] Here, in the table of FIG. 3E, the character IDs, names, and display examples of these two characters are shown. Further, in the table of FIG. 3E, for each character, examples of the fixed phrases for greetings and responses that the client application inserts are shown. The process of the client application inserting the fixed phrases for greetings and responses has been described in FIG. 3C, so repeated explanation is omitted.

[0233] In the table of FIG. 3E, in order to give these characters personalities, characteristic expressions (keywords) are used in both these fixed phrases for greetings and fixed phrases for responses. For example, since Necco is a character modeled after a cat, "nya" is attached to the end of the fixed phrase. Also, Necco's first-person pronoun is "boku". For example, Airia has a tone where the end of the fixed phrase is "desuwa" or "masuwa". That is, it is set so that the end becomes "wa". Also, Airia's first-person pronoun is "watakushi". These keywords may be called keywords indicating the personalities of the characters. These keywords can express the personalities of the characters by being commonly used in the conversations of each character. When performing the character switching process as described in FIG. 3A, the client application may switch the keywords according to the setting information such as the table of FIG. 3E.

[0234] Here, in the example of Figure 3E, in addition to the settings of the fixed phrases inserted by the client application, there are settings for speech habits and responses in the preset instructions of the LLM application. The preset instructions of the LLM application can be set by sending control information from the client application to the LLM application. Since this process is as described in Figure 3D, repeated explanations are omitted. Here, in the example of Figure 3E, in the settings of the speech habits and responses in the preset instructions of the LLM application, instructions are included to also output common characteristic expressions (keywords) in the greeting fixed phrases and response fixed phrases inserted by the client application in the responses of the large language model. This allows the personality of the character to be reflected in the response output of the large language model. Specifically, since Necco is a character modeled after a cat, similar to the settings of the greeting fixed phrases and response fixed phrases inserted by the client application, in the settings of the speech habits and responses in the preset instructions of the LLM application, instructions are set so that "nya" is added to the end of the words. Also, similar to the settings of the greeting fixed phrases and response fixed phrases inserted by the client application, in the settings of the speech habits in the preset instructions of the LLM application, instructions are set so that Necco's first person is "boku". Similarly, for the character Airia, similar to the settings of the greeting fixed phrases and response fixed phrases inserted by the client application, in the settings of the speech habits and responses in the preset instructions of the LLM application, instructions are set so that "wa" is added to the end of the words. Also, similar to the greeting fixed phrases and response fixed phrases inserted by the client application, in the settings of the speech habits in the preset instructions of the LLM application, instructions are set so that Airia's first person is "watakushi". That is, in the example of the table information shown in Figure 3E, characteristic expressions (keywords) that overlap with the characteristic expressions (keywords) indicating the personality of each character included in the settings of the greeting fixed phrases and response fixed phrases inserted by the client application are also included in the settings of the speech habits and responses in the preset instructions of the LLM application.Furthermore, if a character switching process is performed as described in Figure 3A, the client application should switch the settings related to the character's conversational characteristics, as shown in Figure 3E, including the settings for standard greetings, standard interjections, and the preset instructions for catchphrases and interjections in the LLM application, according to the character switching process.

[0235] The information shown in the table in Figure 3E can be stored as table information in the storage unit 1170 or memory 1109 of the artificial intelligence response output device 10010 shown in Figure 1B. Client applications can then utilize this information. The information shown in the table in Figure 3E may also be referred to as setting information regarding the characteristics of the character's conversation.

[0236] Conversation examples that apply the control of standard phrases and preset instructions for the table information shown as a table in Figure 3E, which incorporates these improved settings, will be explained using Figures 3F and 3G. For clarity, the conversation examples in Figures 3F and 3G are shown as examples of conversations that apply the control of standard phrases and preset instructions for the table information shown as a table in Figure 3E.

[0237] First, in Figures 3F and 3G, the right side shows examples of user instructions that the user can input. The details of how to input user instructions have already been explained, so a repeated explanation will be omitted. On the left side of Figures 3F and 3G, examples of output from the artificial intelligence response output device 10010 of the artificial intelligence response output system are shown. Both the user instruction input on the right and the output from the artificial intelligence response output device 10010 on the left are shown in chronological order from top to bottom.

[0238] Here, the example output from the artificial intelligence response output device 10010 on the left first shows a standard phrase response 1 (greeting). This is an example of outputting a standard phrase greeting from the artificial intelligence response output device 10010 using the standard phrase output control explained in Figure 3C, etc. This standard phrase output control can be performed by the client application 8010. Next, when a user instruction sentence 1 is input to the artificial intelligence response output device 10010, the client application 8010 acquires it and sends the user instruction sentence 1 and control information to the large-scale language model. The large-scale language model that receives the user instruction sentence 1 performs inference at the timing of the star mark 9001 and generates LLM response 1. This large-scale language model may also be the large-scale language model of the large-scale language model server 20001. In this case, the input and output of information to and from this large-scale language model is controlled by the server LLM application 8020. Alternatively, this large-scale language model may also be the large-scale language model of the local LLM processing unit 10028. In this case, the input and output of information to the large-scale language model is controlled by the local LLM application 8015. The LLM response 1 generated by the large-scale language model is sent to the client application 8010 by these LLM applications and output as an artificial intelligence response from the artificial intelligence response output device 10010. If the artificial intelligence response output device 10010 is a character conversation device, the artificial intelligence response output is recognized by the user as a response from the displayed character.

[0239] Next, an example is shown in which user instruction 2 is input to the artificial intelligence response output device 10010 in response to LLM response 1. Here, an example is shown in which the client application 8010, which has acquired user instruction 2, outputs a standard response 2 (acknowledgment). This is an example in which the artificial intelligence response output device 10010 outputs a standard acknowledgment using the standard output control described in Figure 3C, etc. Multiple types of standard acknowledgments can be prepared in advance, and the output can be controlled to be randomly spaced at intervals of a predetermined period or longer to avoid unnaturally high frequency. Meanwhile, while outputting the standard response, the client application 8010 transmits the user instruction 2 and control information to the LLM application (server LLM application 8020 or local LLM application 8015) mentioned above. The LLM application that has acquired user instruction 2 controls the large-scale language model (the large-scale language model of the large-scale language model server 20001 or the large-scale language model of the local LLM processing unit 10028) to perform inference at the timing of the star mark 9002. The large-scale language model inference generates LLM response 2. This LLM response 2 generated by the large-scale language model is sent to the client application 8010 by the LLM application and output as an artificial intelligence response from the artificial intelligence response output device 10010. If the artificial intelligence response output device 10010 is a character conversation device, this artificial intelligence response output is perceived by the user as the response of the displayed character.

[0240] Through the above series of processes, a conversation takes place between the artificial intelligence response output system and the user. Figures 3F and 3G show an example of a series of conversations between the user and the AI ​​regarding Haneda Airport in Japan, with greetings and acknowledgments mixed in. The series of AI response outputs from the AI ​​response output device 10010 are perceived by the user as a series of responses with a certain degree of consistency. However, these series of responses are composed of response sentences generated by a large-scale language model whose information input and output are controlled by the LLM application and whose output is controlled by the client application's control of fixed phrases. In other words, the collaboration between the client application and the LLM application results in AI response outputs that are more suitable for the user.

[0241] Here, we will explain the details of the control using the table information shown in Figure 3E for each of the conversation examples in Figure 3F and Figure 3G.

[0242] First, Figure 3F shows an example of a conversation in which the control of the standard phrases and preset instructions from the table information shown in Figure 3E is applied. Figure 3F shows an example of a conversation when the client application is controlling the output of an artificial intelligence response from the artificial intelligence response output device 10010 as a response from the character Necco. The client application reads the information of the character Necco's standard greeting phrases and standard interjection phrases from the table information shown in Figure 3E and uses it to output standard phrases for the artificial intelligence response. As shown in the conversation example in Figure 3F, standard phrase response 1 uses the character Necco's standard greeting phrase, with the first-person pronoun being "boku" and ending with "nyaa". Also, standard phrase response 2 uses the character Necco's standard interjection phrase, ending with "nyaa". Here, in the table in Figure 3E, the settings for catchphrases and interjections based on the LLM application preset instructions include characteristic expressions (keywords) that overlap with characteristic expressions (keywords) that represent the personality of character Necco, namely "meow" and the first-person pronoun "boku" which are included in the settings for sentence endings and interjections. As a result, in the conversation example in Figure 3F, LLM response 1 ends with "meow". In LLM response 2, the sentence ending is "meow" and the first-person pronoun is "boku". In addition, LLM response 2 includes the interjection "Hmm meow" which was generated by the large-scale language model according to the LLM application preset instructions. The ending of this interjection is also "meow". Thus, in a series of conversations, whether the output from the artificial intelligence response output device 10010 as the response of character Necco is a fixed sentence output by the client application or a response output by the large-scale language model controlled by the LLM application, a common characteristic expression (keyword) is consistently used. This allows users to perceive a more consistent impression of the character's personality in the content output from the artificial intelligence response output device 10010 as the character Necco's response.

[0243] Similarly, Figure 3G is another example of a conversation in which the control of the standard phrases and preset instructions from the table information shown in Figure 3E is applied. Figure 3G shows an example of a conversation in which the client application controls the output of an artificial intelligence response from the artificial intelligence response output device 10010 as a response from the character Airia. The client application reads the information of the character Airia's standard greeting phrases and standard interjection phrases from the table information shown in Figure 3E and uses it for outputting standard phrases for the artificial intelligence response. Then, as shown in the conversation example in Figure 3G, standard phrase response 1 uses the character Airia's standard greeting phrase, with the first-person pronoun being "watakushi" and the ending being "wa". Also, standard phrase response 2 uses the character Airia's standard interjection phrase, with the ending being "wa". Here, in the table in Figure 3E, the settings for catchphrases and interjections based on the LLM application preset instructions include characteristic expressions (keywords) that overlap with characteristic expressions (keywords) that represent the personality of character Airia, namely "wa" and the first-person pronoun "watakushi" which are included in the settings for sentence endings and interjections. As a result, in the conversation example in Figure 3G, LLM response 1 ends with "wa". Also, LLM response 2 ends with "wa" and uses "watakushi" as the first-person pronoun. Furthermore, LLM response 2 includes the interjection "Uuun desu wa" which was generated by the large-scale language model according to the LLM application preset instructions. The ending of this interjection is also "wa". Thus, in a series of conversations, whether the output from the artificial intelligence response output device 10010 as the response of character Airia is a fixed phrase output by the client application or a response output by the large-scale language model controlled by the LLM application, a common characteristic expression (keyword) is consistently used. As a result, users can get a more consistent impression of the character's personality from the content output from the artificial intelligence response output device 10010 as the character Airia's response.

[0244] As described above, in the artificial intelligence response output system explained using Figures 3D to 3G, the client application and the LLM application work together to perform a series of artificial intelligence response outputs. Furthermore, in the artificial intelligence response output system explained using Figures 3D to 3G, the system controls the use of consistently common characteristic expressions (keywords) for each character in the standardized text output by the client application and the response output of the large-scale language model controlled by the LLM application. This allows the user to get a more consistent impression of the character's personality in the content output from the artificial intelligence response output system. In addition, in this system, the artificial intelligence response output device 10010 can control the use of consistently common characteristic expressions (keywords) for each character in the standardized text output by the client application and the response output of the large-scale language model controlled by the LLM application by sending and receiving control information to and from the client application it is executing. This control by the client application executed by the artificial intelligence response output device 10010 allows the user to get a more consistent impression of the character's personality in the content output from the artificial intelligence response output system.

[0245] In the conversation examples shown in Figures 3F and 3G, the output of the standardized response 2, which is a set phrase for acknowledgment, begins after receiving the user instruction sentence 2, but before the LLM response output 2, which is the response output of the large-scale language model controlled by the LLM application. Since the inference of the large-scale language model takes a certain amount of time, the client application is controlled to insert the standardized acknowledgment at this timing in order to prevent an unnatural gap in the response output from the artificial intelligence response output device 10010 to the user. In other words, the time immediately before the start of the inference of the large-scale language model is one of the suitable timings for outputting a standardized response.

[0246] Furthermore, the term "character" in each embodiment of the present invention described above includes the concepts of AI assistants and avatars.

[0247] The table information described in Figure 3E above contains various settings related to the characteristics of character conversations, but the output medium for character conversations is mainly natural language. Therefore, if the language output by the character conversations is different, the various settings information related to the characteristics of these character conversations also needs to be different. Accordingly, if the artificial intelligence response output device 10010 can switch between multiple languages ​​as the language output by the character conversations, it is sufficient to store various settings information related to the characteristics of character conversations, such as the table information described in Figure 3E, corresponding to each of the multiple languages, in the storage unit 1170 or memory 1109. When the artificial intelligence response output device 10010 switches the language output by the character conversations, it is desirable to also switch to the table information corresponding to that language.

[0248] <Example 4> Next, Example 4 of the present invention will be described.

[0249] In this embodiment, we will explain the differences from the embodiments described so far, and will omit repeated explanations of configurations similar to those in those embodiments. In other words, in each figure of this embodiment, configurations that are denoted by the same reference numerals as those in the figures of the embodiments described so far have the same function and configuration. Such configurations will not be explained repeatedly in order to simplify the explanation.

[0250] Example 4 uses artificial intelligence (AI) to recognize gesture operations. Here, the AI ​​used in Example 4 is a multimodal language model that learns primarily from natural language text data, as well as from non-natural language sources such as images, videos, and audio. A so-called base model may also be used. Depending on the size of the parameters, multimodal language models can be large-scale language models (LLM), medium-scale language models (MLM), small-scale language models (SLM), etc. A multimodal language model of any size may be used, but in this example, an example using a large-scale language model (LLM) will be described. The AI ​​response output device in Example 4 receives and detects the user's gesture operation input at the interface, obtains the gesture recognition result as a response through interaction with the AI ​​(LLM), and executes processing corresponding to the result. In this example, the artificial intelligence response output device is referred to as the AI ​​response output device. This has the same meaning as the artificial intelligence response output device described in previous examples.

[0251] Here, the AI ​​response output device and the AI ​​response output system including it in Example 4 perform the following in the process from detecting a user's gesture input to executing processing on said gesture input: The AI ​​response output system inputs an instruction (prompt) to an LLM (Large-Scale Language Model) corresponding to the gesture operation, performs inference and generates a response by the LLM, and executes processing on said response.

[0252] Furthermore, the response output device and system of Example 4, which determine gesture input from the user, can also be called a gesture determination device and gesture determination system. Additionally, the response output device and system of Example 4 outputs a response to the gesture input from the user. This can also simply be called an output device and output system.

[0253] [Mechanism: A system that uses a multimodal language model for gesture recognition] First, an example of the operation of the AI ​​response output device and AI response output system of Example 4 will be described using Figure 4A. Figure 4A shows an overview of the mechanism and overall configuration example of using AI (LLM) for gesture recognition in Example 4. The configuration of the AI ​​response output system in Figure 4A is an improvement on the AI ​​response output system in Figure 2C. Components with the same reference numerals as in Figure 2C are the same as those in Figure 2C, so repeated explanations will be omitted.

[0254] This AI response output system is a system in which the AI ​​response output device 10010 is connected to the server LLM application 8020 (or local LLM application 8015) described above (Figure 1E). The AI ​​response output device 10010 is a device used by user 230. In addition, a mobile terminal (mobile information processing terminal) 20010 used by user 230 may be connected to the AI ​​response output device 10010.

[0255] In Figure 4A, as shown in (1) of Explanation 4010, user 230 performs operation input (sometimes referred to as gesture operation input) to the AI ​​response output device 10010 (particularly the user interface) using gestures with their finger, hand, or physical object (e.g., a pen). An example of such a gesture is a gesture of drawing characters on a two-dimensional surface or in three-dimensional space. This may also be called a gesture that mimics characters. The illustrated example of gesture operation input schematically shows the case of drawing the English alphabet character "JU" on the two-dimensional screen of the display unit 10011.

[0256] The AI ​​response output device 10010 detects gesture input from the user 230 using sensors, etc. Specifically, for example, it detects touch operations (in other words, touch gestures) by the user to the touch operation input sensor of the touch panel of the display unit 10011. Another example is the detection of gestures based on the imaging unit 1180 capturing images of the user 230's fingers, etc. In another example, if the response output device 10010 is an aerial floating image display device 185 (Figure 1K), it detects aerial operations by the user 230's fingers, etc., using the aerial operation detection sensor 1351.

[0257] Through the detection means such as the sensors described above, the AI ​​response output device 10010 can detect the gesture input of the user 230 (gesture detection) and obtain gesture transmission data 4A01 corresponding to the detection result. Gesture transmission data 4A01 is data that represents the detection result of the gesture input by the user 230 and is data for transmitting the gesture to the AI ​​(LLM). Specifically, gesture transmission data 4A01 is image data that represents the trajectory corresponding to the gesture (for example, finger movement) and line segment data that represents line segments that constitute that trajectory.

[0258] Alternatively, user 230 may input a gesture operation from a mobile information processing terminal 20010 used by user 230. The mobile information processing terminal 20010 (Figure 2B) detects the touch operation (touch gesture) by user 230 via a touch operation input sensor on, for example, the touch panel 20011 display panel. The mobile information processing terminal 20010 transmits a detection signal / control signal (or gesture transmission data created by the mobile information processing terminal 20010) corresponding to the detection result of the gesture operation to the AI ​​response output device 10010 via communication. Based on the control signal, the AI ​​response output device 10010 detects the gesture operation input by user 230 and obtains the gesture transmission data 4A01.

[0259] Note that the detection of the above gesture operation input (gesture detection) corresponds to the detection of a trajectory corresponding to a gesture in two or three dimensions. For the technical means of this gesture detection, existing technologies may be used, and it is not particularly limited to a specific technology. The format of the gesture transmission data 4A01 and the like are not limited to a specific format, and an example thereof will be described later.

[0260] Next, as shown in (2) of the explanation 4020, the AI response output device 10010 generates gesture transmission data 4A01 based on the detected gesture operation input (gesture detection result) by the user 230. Note that the sensor, processor, etc. in (1) may generate the gesture transmission data 4A01 during detection. The AI response output device 10010 combines the gesture transmission data 4A01 and natural language to generate an instruction text (prompt) 4A02. Here, the natural language combined with the gesture transmission data 4A01 is natural language characters for enabling a multimodal LLM to understand the instruction. This is different from a comment that inserts characters into a part that is ignored as a program for human understanding in the description of a programming language. That is, in this embodiment, the natural language combined with the gesture transmission data 4A01 in the instruction text is the natural language described in a part that is not ignored by the multimodal LLM in the instruction text. Then, the AI response output device 10010 transmits the generated instruction text 4A02 to the LLM application 8020.

[0261] The processing in (2) may be performed by the control unit 1110 of the AI response output device 10010, particularly the client application 8010 in FIG. 1E. Here, the LLM application 8020 may be the server LLM application 8020 in FIG. 1E or the local LLM application 8015 in FIG. 1E. In any case, the LLM application is a multimodal LLM that can process images and the like in addition to text.

[0262] In the example of the description 4020 in FIG. 4A, the gesture transmission data 4A01 is, as an example, image data generated based on a gesture operation input (detection result) by the user 230. However, the gesture transmission data 4A01 is not limited to this, and it may be any data that represents a gesture, such as line segment data generated based on a gesture operation input (detection result). In the illustrated example, specifically, in the image data that is the gesture transmission data 4A01, a line such as the characters "JU" of the English alphabet is drawn. The illustrated example of the gesture transmission data 4A01 is image data in which the trajectory of the characters "JU" is detected as the presence or absence of touch by a touch operation detection sensor of a touch panel having a two-dimensional screen (FIG. 4O described later). Note that the gesture transmission data 4A01 is not character data (e.g., character code).

[0263] Here, as an example of the natural language included in the instruction 4A02 together with the gesture transmission data 4A01, there is "This image has English alphabet characters written on it. Please return the alphabet characters of the first two characters from the left as character data." That is, the natural language that the AI response output device 10010 includes in the instruction 4A02 is a natural language that, through the inference of a multimodal LLM, recognizes and discriminates the characters (the characters represented by the gesture) included in the gesture transmission data 4A01 and instructs to return the character data of the said characters. Details and examples of this natural language will be described later.

[0264] The natural language of the said instruction 4A02 includes (1) that English alphabet characters are written, and (2) to return the alphabet characters of the first two characters from the left as character data. These correspond to examples of rules / premises determined in the user interface of this system. The said rules can also be selected by the user (described later).

[0265] Next, the server LLM application 8020 (or local LLM application 8015) that received the instruction 4A02 performs LLM inference in accordance with the instruction 4A02, and as a result of the inference, generates a response 4A03 as shown in (3) of explanation 4030. Response 4A03 is the character (character data) of the gesture recognition result. Response 4A03 is, for example, two alphabetic character data (e.g., character code) such as "JU".

[0266] Next, the AI ​​response output device 10010, particularly the client application 8010, which received the response 4A03, performs processing according to the content of the response 4A03. In the example of (4) in Description 4040, the client application 8010 uses a correspondence table (described later) between registered characters and processing held by the AI ​​response output device 10010 to select a registered character based on predetermined selection conditions using the character data which is the response 4A03, and then determines and executes the processing associated with that registered character.

[0267] In the example shown in Figure 4A, the character "JU" from the character data of response 4A03 is used with the correspondence table described above to select the registered character "JUMP" from the character "JU", and the process jump(), which is associated with the registered character "JUMP", is determined. The jump() process is, for example, a process that makes the character 19051 (character image) on the screen jump. This process can also be called a command or function. The AI ​​response output device 10010 executes the jump() process. As a result, an image of character 19051 jumping is displayed. Details of the process using the correspondence table will be described later.

[0268] Through the series of actions shown in Figure 4A, the AI ​​response output device 10010 and the system recognize, through interaction with the LLM, that user 230 has made a gesture of drawing the letters "JU" in natural language with their finger, and execute a process determined according to the recognized letters. As a result, for example, a video of character 19051 jumping is output. In other words, from user 230's perspective, it feels as if they have made an operation to make character 19051 jump by imagining "JUMP" as the action of character 19051 and drawing the letters "JU". To put it another way, from user 230's perspective, it feels as if they have given an instruction to character 19051 with a gesture corresponding to natural language letters, and have received a response of action from character 19051 in response.

[0269] Therefore, according to the AI ​​response output device 10010 and system shown in Figure 4A described above, when performing gesture judgment on a gesture operation input by user 230, an instruction sentence 4A02 including gesture transmission data 4A01 and natural language is generated, and an LLM inference is performed based on the instruction sentence 4A02 to obtain a response 4A03 including character data. Furthermore, according to this system, processing can be selected, decided, and executed based on a correspondence table or the like from the characters in the response 4A03.

[0270] [Processing flow] Next, the series of operations in the AI ​​response output device, method, and system of Example 4, from the user 230's gesture input to the response processing, will be explained using the flowchart in Figure 4B.

[0271] The flow in Figure 4B includes, in order, the steps S4401 for gesture operation input, S4402 for instruction text generation, S4403 for LLM inference and response generation, and S4404 for response processing.

[0272] An example of the process in step S4401 is the process shown in explanation 4010 of Figure 4A. An example of the process in step S4402 is the process shown in explanation 4020. An example of the process in step S4403 is the process shown in explanation 4030. An example of the process in step S4404 is the process shown in explanation 4040.

[0273] In step S4401, user 230 makes a gesture input to the user interface (consisting of a display unit 10011, etc.) of the AI ​​response output device 10010, and the AI ​​response output device 10010 detects that gesture input.

[0274] In step S4402, the AI ​​response output device 10010 generates gesture transmission data 4A01 based on the detection result of the gesture operation input in step S4401, and generates an instruction sentence (prompt) 4A02 that includes the gesture transmission data 4A01 and natural language. The AI ​​response output device 10010 sends the generated instruction sentence 4A02 to the LLM (LLM application 8020 or LLM application 8015).

[0275] Alternatively, the gesture transmission data 4A01 may be generated in step S4401 or in step S4402. The gesture transmission data 4A01 may be the aforementioned image data or line segment data.

[0276] Instruction 4A02 is generated by combining (1) the generated gesture transmission data 4A01 (image data or line segment data) with (2) a predetermined natural language string. The predetermined natural language is a natural language for instructions that corresponds to the rules of the system's user interface (e.g., the language to be used, character type, number of characters, etc.). The predetermined natural language can be created by the client application 8010 based on setting information in the system settings or user settings (which may be selections made by the user 230 on the UI of the screen described later).

[0277] In step S4403, based on the instruction statement 4A02 generated in step S4402, the LLM performs inference and generates a response 4A03 containing character data as an inference result, which is then transmitted to the AI ​​response output device 10010.

[0278] In step S4404, based on the characters of response 4A03 (the characters of the gesture recognition result) generated in step S4403, the client application 8010 of the AI ​​response output device 10010 determines the registered characters and processing (response processing) to be associated with the characters of response 4A03 from the correspondence table and selection conditions, and executes the processing. The selection conditions include, for example, a search by prefix matching or random number selection.

[0279] In this embodiment, the processing performed in response to the recognition of gesture input (in other words, response processing, target processing, etc.) is, in the example shown in Figure 4A, a process that instructs the character on the screen to perform a predetermined action (in other words, a process that outputs a video of the character performing a predetermined action). For example, in the case of the jump() process, the response output device 10010 (video control unit 1160, etc.) generates a video of the character jumping through rendering processing based on the character's 3D model.

[0280] This response processing is not limited to this example; it can be any processing that can be executed by the response output device 10010. This processing can be, for example, various operations / interfaces of a typical PC or various applications. Examples include copy, paste, undo, and save. For example, the character for gesture operation input could be "CO", the registered character could be "COPY", and the response processing could be copy().

[0281] Furthermore, the response processing is not limited to being implemented solely by the response output device 10010; it may also be implemented through cooperation between the response output device 10010 and an external server, or in other words, through a server service. For example, the response output device 10010, acting as a client, sends a request for response processing to an external server. The server executes the response processing (e.g., rendering character images) and sends the processing result data as a response to the response output device 10010. The response output device 10010 receives and outputs that response.

[0282] The following sections, starting with Figure 4C, will explain specific examples of the operation of the AI ​​response output device and system in these steps.

[0283] [User Interface] Using Figure 4C, we will explain the details of step S4401 in Figure 4B. Figure 4C shows an example of a display screen 4101 as an example of a user interface (UI) that accepts gesture operation input in the AI ​​response output device 10010. The display screen 4101 is displayed on the display unit 10011.

[0284] When the AI response output device 10010 is the display device 181 in (1) or the information processing terminal 182 in (2) of FIG. 1C, the screen such as the display panel they have is the display screen 4101 in FIG. 4C. When the AI response output device 10010 is the HMD 183 in (3) of FIG. 1C, the screen displayed on the optical image of the virtual image generated by the eyepiece optical system is the display screen 4101 in FIG. 4C. When the AI response output device 10010 is the HUD 184 in (4) of FIG. 1C, the screen displayed on the optical image of the virtual image generated by the HUD is the display screen 4101 in FIG. 4C. When the AI response output device 10010 is the airborne floating image display device 185 in (5) of FIG. 1C, the screen displayed on the optical image of the real image generated by the airborne floating image display device 185 is the display screen 4101 in FIG. 4C. When the AI response output device 10010 is the projection-type image display device 186 in (6) of FIG. 1C, the screen displayed on the projection image projected by the projection-type image display device 186 is the display screen 4101 in FIG. 4C.

[0285] (1) in FIG. 4C is an example in which a standby screen is displayed as display content on the display screen 4101 of the AI response output device 10010. (2) in FIG. 4C is an example in which a gesture operation input screen is displayed as display content on the display screen 4101 of the AI response output device 10010.

[0286] In the standby screen (and the corresponding standby mode) of (1), as the object image, for example, a character image 19051 in the standby state (or the idling state) is displayed. On the standby screen, buttons such as the button 4111 are also displayed as objects. The button 4111 is, for example, a gesture operation input button. This button 4111 is a general icon button that can be touched and operated by the user 230 with a finger 4201 or the like. The button 4111 is a button that can switch the screen and the corresponding mode to gesture operation input, that is, a mode switch button. When the AI response output device 10010 detects a touch operation on the button 4111, it changes the content of the display screen 4101 to the gesture operation input screen (and the corresponding gesture operation input mode) shown in (2).

[0287] In the gesture input screen of (2), several UI images are displayed in front of or around the character image 19051. In this example, the UI images displayed include the input language selection UI 4113, the rule explanation UI 4112, the Clear button 4115, and the Done button 4116.

[0288] The input language selection UI 4113 allows user 230 to select the language (e.g., English, Japanese) to be used for gesture input by touching it. For example, if English is selected as the language, the user can use English alphabet characters for the characters / strings drawn with gestures, and if Japanese is selected, the user can use Japanese katakana / hiragana / kanji characters. When using English alphabet characters, the user can input characters that represent natural language (words) such as "JUMP" and "DASH" with gestures. When using Japanese katakana characters, the user can input characters that represent natural language (words) such as "ジャン" and "ダッシュ" with gestures. User 230 can select the language that is easiest for them to use from the options. Note that the input language selection UI 4113 may also allow selection of character types in addition to language. The example shown illustrates the case where "Input Language: English" is selected and uppercase alphabet characters are used.

[0289] Rule explanation UI4112 displays messages such as explanations / guides regarding the basic rules of gesture input in the system's user interface. In this example, Rule explanation UI4112 displays "Please Input Touch Gesture by 2 characters" in English. In this example, the rule is that user 230 should draw the first two characters of the natural language (word) they imagine for the action they want character 19051 to perform, using a touch gesture. In this example, the number of characters used for gesture input is 2.

[0290] Depending on the design of the system and user interface, including the response output device 10010, one rule is to pre-set the number of characters used in a single gesture input. Since it is preferable to minimize the effort required for gesture input, this number of characters is set to a relatively small number. In this example, the number of characters used is shown as 2 characters (2 English alphabet characters). However, it is not limited to this, and it is possible to set it to 1 character, 3 characters, 4 characters, etc. Alternatively, a maximum number of characters may be set, such as 4 characters or less, and input may be allowed freely within that maximum number of characters. It is also possible to allow user 230 to input any number of characters as desired without restricting the number of characters used (although it is restricted in internal information processing).

[0291] Note that while the registered characters in the correspondence table described later represent regular words (natural language), the characters used for gesture input (recognition result character data) may be, for example, two characters and do not necessarily represent regular words (natural language). For example, the characters "DA" do not identify a unique word. These gesture input characters can also be described as the first character, abbreviation, etc.

[0292] User 230 performs a gesture operation input with the desired characters according to the UI of screen 4101, for example, by making a touch gesture to the touch panel. For example, if User 230 wants to make character 19051 perform a dashing action, User 230 imagines "DASH" as the natural language (word) that represents that action, and inputs the first two letters, "DA," with a gesture. User 230 makes a touch gesture with their finger 4201 to draw "DA" on screen 4101. The AI ​​response output device 10010 detects this gesture operation input and obtains gesture transmission data 4A01 (e.g., image data) that includes the trajectory of "DA."

[0293] Furthermore, at this time, the AI ​​response output device 10010 may display the status and result of the detected gesture operation input on the screen 4101 in near real time. In this example, the trajectory of the character "DA" is displayed as an image (gesture detection image) 4114.

[0294] After user 230 has made a desired gesture input, for example, if they look at image 4114 and wish to clear the content of that gesture input, they operate the Clear button 4115. As a result, the AI ​​response output device 10010 clears the content of the detected gesture input and also erases image 4114.

[0295] After the user 230 makes a desired gesture input, for example, by looking at image 4114, if they decide to proceed based on the content of the gesture input, they operate the Done button 4116. This causes the AI ​​response output device 10010 to use the detected gesture input content (gesture transmission data 4A01) to generate an instruction statement 4A02, etc. As a variation, this may be performed automatically based on the passage of time, even without the operation of the Done button 4116.

[0296] Figure 4D shows another configuration example for the screen in Figure 4C. (1) is the case where the language used is set to Japanese and the rule is to input gesture operations using a single katakana character. For example, user 230 imagines a dash and inputs the character "ダ" with a gesture. It is also possible to input using hiragana or kanji. Furthermore, it is possible not to restrict the type of characters used, such as katakana, hiragana, or kanji, when inputting in Japanese. In addition, the display language of the UI on screen 4101 may be switched to match the selected language used. In the illustrated example, the display language of the UI is also shown to be Japanese.

[0297] The example in (2) is when there is no limit on the number of characters that can be used, and the user 230 inputs a gesture with the desired number of characters. For example, when the user inputs "JUMP" using four letters of the alphabet. On screen 4101, there is no instruction on the number of characters, and the explanation / guide is simply displayed as "Please Input Touch Gesture". The AI ​​response output device 10010 may then create an instruction sentence 4A02, including natural language, based on the gesture transmission data 4A01, to use a predetermined number of characters (for example, 2 characters) as an internal rule of the AI ​​response output device 10010 that is not presented to the user 230, and perform gesture recognition.

[0298] Alternatively, as a variation, the AI ​​response output device 10010 may create an instruction sentence 4A02 that includes natural language, without specifying the number of characters to be used in the natural language added to the gesture transmission data 4A01, so as to cause the AI ​​(LLM) to perform gesture recognition. As a result, the response 4A03 may be, for example, "JUMP" (4 characters).

[0299] Alternatively, in a modified example, the AI ​​response output device 10010 may, if it can detect the number of characters input via gesture (e.g., 4 characters), create an instruction sentence 4A02 that includes natural language and uses the detected number of characters. As a result, the response 4A03 may contain characters such as "JUMP" (4 characters).

[0300] [Instructions (prompts)] Figure 4E shows an example of the structure of instruction 4A02. In Figure 4E, the format, data structure, specific examples, and example responses of instruction 4A02 are summarized in a table. The example response corresponds to the gesture judgment result (character data of response 4A03). An example is shown for each row. The instruction structure is gesture transmission data 4A01 + natural language, as described above (Figure 4A).

[0301] In Example 1, the instruction format is gesture image data + natural language. This is the case where the gesture transmission data 4A01 is image data. An example of natural language is in Japanese: "This image has English alphabet letters written on it. Please return the N letters of the alphabet from the left as character data." This is the case when the instruction 4A02 in Japanese is given to the LLM. In this format, the parts "English alphabet letters" and "N letters" can be set (changed) according to the specific design rules of the AI ​​response output device 10010. In the specific instruction example, the "N letters" part of the natural language is set to "2 letters". The gesture image data 4E01 is an image with the letters "DASH" drawn on it. The LLM recognizes and determines the first two letters (from the left) from this image and returns the character data "DA". The response examples all show the case where the answer is correct. Note that in this example, uppercase letters are used, but it is not limited to this, and lowercase letters can also be used.

[0302] In Example 2, the format is gesture image data + natural language, and the example of natural language is in English: "This image includes an alphabet character. Please return N characters from the left side in the image as character data." This is the case when giving the instruction 4A02 in English to the LLM. In this format, the "alphabet" and "N characters" parts can be set (changed) according to the specific design rules of the AI ​​response output device 10010. In the specific instruction example, the "N characters" part of the natural language is set to "2 characters". The gesture image data 4E02 is an image with the characters "DA" drawn on it. The LLM recognizes and determines the first two characters (from the left) from this image and returns the character data "DA".

[0303] In Example 3, the format is gesture image data + natural language, and the example of natural language is in Japanese: "This image has Japanese katakana written on it. Please return the katakana characters from the left as character data." This is the case when giving the instruction sentence 4A02 in Japanese to the LLM. In this format, the parts "Japanese katakana" and "M characters" can be set (changed) according to the specific design rules of the AI ​​response output device 10010. In the specific instruction sentence, the part "M characters" in natural language is set to "1 character". The gesture image data 4E03 is an image with the character "ジャ" drawn on it. The LLM recognizes and determines the first character (from the left) from this image and returns the character data "ジ".

[0304] In Example 4, the format is gesture line segment data + natural language. An example of natural language is in Japanese: "This line segment data contains Japanese katakana characters. Please return the katakana characters starting from the left as character data." In the example instruction, the "M characters" in the natural language is set to "1 character". Gesture line segment data 4E04 is line segment data in which the line segment constituting the character "ジャ" has been detected. LLM recognizes and determines the first character (from the left) from this line segment data and returns the character data "ジ".

[0305] Furthermore, the following are other possible configurations for instruction 4A02.

[0306] Instruction 4A02 may also include information specifying the character encoding to be used (such as the corresponding character types). For example, it could state in natural language, "Please output the character data in UTF-8."

[0307] Information regarding rule settings such as the language used, character type, character code, and number of characters, as described above, may be included in the user instruction statement mentioned above, or it may be included in the system instruction statement (the setting instruction statement mentioned above).

[0308] [Error handling method 1] Furthermore, in response to instruction 4A02, the character encoding of the text data in response 4A03 from the AI(LLM) may fall outside the specified character encoding range for the language set as the input language. In other words, the AI(LLM) does not necessarily return a judgment result that conforms to the instruction in response to instruction 4A02. With conventional AI technology, even if the instruction specifies the alphabet, the AI ​​may return Greek letters in its response. Alternatively, even if the input language is set to the English alphabet, the user may mistakenly input Japanese. To handle such cases, error countermeasures such as those described below may be implemented.

[0309] If the character code of the character data in response 4A03 from the AI ​​(LLM) falls outside the predetermined character code range for the language set as the input language, the AI ​​response output device 10010 may determine this to be an error after checking response 4A03 (character data) and display an error message on the display screen for the user 230. For example, in the case of the first error, the response output device 10010 may display "Please enter again." Furthermore, in the case of a second error, it may display "Please review your input language settings." By retrying in this manner, the character code of response 4A03 (character data) from the LLM may become an appropriate result within the specified character code range.

[0310] For example, suppose user 230 intends to input the letter "V" using a touch gesture. The natural language instruction 4A2 to the LLM, corresponding to the gesture image data obtained, might say, "This image contains English letters. Please return the first letter from the left as character data. Output the character data in UTF-8." In response, suppose the LLM outputs the Greek letter "ν" (nu) as response 4A03 (character data). In this case, the UTF-8 character code for "ν" (nu) is "CEBD," which falls outside the range of "41" to "7A," the range of UTF-8 character codes for so-called alphabets (basic Latin characters). In such a case, the response output device 10010 should determine this is an error and display the error message described above.

[0311] The correspondence table may specify the range of character codes that will be accepted, and if the input characters (gesture recognition results) fall outside that range, an error message or instructions on how to handle the situation may be output.

[0312] [Error handling method 2] As another example of error handling, the AI ​​response output device 10010 may maintain a correspondence table in advance that lists the character codes of similar characters across multiple character types (corresponding character code ranges). If the character code of the character data in response 4A03 is outside the predetermined character code range set in the input language, the AI ​​response output device 10010 will determine it to be an error and perform the following countermeasures: The response output device 10010 may determine and output as an alternative response a character that is recorded / registered in the above correspondence table as being similar to the character code of the character included in response 4A03, and that falls within the character code range of the character type in the "input language setting".

[0313] For example, you can pre-register the Greek letter "ν" (nu), whose UTF-8 character code is "CEBD", the uppercase letter "V" (vi), whose UTF-8 character code is "56", and the lowercase letter "v" (vi), whose UTF-8 character code is "76", as similar characters in the above correspondence table.

[0314] For example, suppose user 230 intended to input the letter "V" using a touch gesture, and the natural language instruction 4A02 for the gesture image data was "This image contains English letters. Please return the first letter from the left as character data. Please output the character data in UTF-8." Suppose the Greek letter "ν" is output as response 4A03. At this point, the AI ​​response output device 10010 determines that the UTF-8 character code for "ν" is "CEBD," which falls outside the range of "41" to "7A," which is the range of UTF-8 character codes for so-called alphabets (basic Latin letters). The AI ​​response output device 10010 then checks the above correspondence table and decides to output the uppercase letter "V," which has a UTF-8 character code of "56" and is registered as a "similar character" to the Greek letter "ν," as an alternative response. Based on this determined "V," a registered character is selected, and the processing associated with the registered character is determined, as explained in Figure 4A.

[0315] The above countermeasure using "similar characters" aims to make the LLM accept as wide a range of recognized characters as possible.

[0316] Furthermore, regarding the similar characters mentioned above, if there are multiple options such as uppercase and lowercase letters of the alphabet, a predetermined priority order, such as "prioritize uppercase" or "prioritize lowercase," may be set in a correspondence table, and the system may be configured to make a decision according to that priority order.

[0317] As shown in the example above, the response output of the LLM may not be fully controllable by instruction alone, so the error detection and countermeasures for the error detection results on the AI ​​response output device 10010 side are effective.

[0318] [Correspondence Table] Figure 4F shows an example of the configuration of the correspondence table 4F00 held by the AI ​​response output device 10010. The correspondence table 4F00 may be stored, for example, in the storage unit 1170 of the AI ​​response output device 10010. Alternatively, the correspondence table 4F00 may be held, for example, in the memory 1109 of the AI ​​response output device 10010. The client application 8010 can refer to the correspondence table 4F00 and use it for control. The correspondence table 4F00 in Figure 4F has the following column items: ID, Registered character (English alphabet), Registered character (Japanese katakana), Response processing (function), and Meaning of the Response processing. One piece of correspondence information is registered / set for each row (ID). In this example, "Registered character" is shown as having two character types, alphabet and katakana, but it is not limited to this, and it is sufficient that at least one character type of registered character is set, depending on the system and UI design. The "response processing (function)" is the process that the response output device 10010 determines and executes in response to response 4A03 (as shown in (4) of Figure 4A above).

[0319] In the example with ID=1, the registered characters are the English alphabet "JUMP" and the Japanese katakana "ジャンプ". The process associated with these registered characters is "jump()", and its meaning (process) is for the character to jump. In other words, this process makes the character on the screen perform a jumping motion.

[0320] In the example with ID=2, the registered characters are the English alphabet "DANCE" and the Japanese katakana "ダンス". The process associated with these registered characters is "dance()", and its meaning (what it does) is for the character to dance.

[0321] In the example with ID=3, the registered characters are the English alphabet "DASH" and the Japanese katakana "ダッシュ". The process associated with these registered characters is "run()", and its meaning (the process performed) is to make the character run.

[0322] In the example with ID=4, the registered characters are the English alphabet word "SING" and the Japanese katakana word "うたう". The process associated with these registered characters is "sing()", and its meaning (the process performed) is for the character to sing.

[0323] In the example with ID=5, the registered characters are the English alphabet "RUN" and the Japanese katakana "Hashiru". The process associated with these registered characters is "run()", just like with ID=3.

[0324] In the example with ID=6, the registered characters are the English alphabet "TURN" and the Japanese katakana "タン". The process associated with these registered characters is "turn()", and its meaning (process) is that the character rotates (turns).

[0325] In the example with ID=7, the registered characters are the English alphabet "IDLE" and the Japanese katakana "Taiki". The process associated with these registered characters is "idle()", and its meaning (process) is to return the character to an idle state.

[0326] In the example in Figure 4F, for ID=3 and ID=5, one identical process (run()) is associated with two different registered characters, "DASH" and "RUN". Such settings are permitted. In this way, the same single response process can be associated with multiple different registered characters. Other registered characters not shown in the diagram, such as "run," "hashiru," and "suru," can also be associated. In other words, this accepts a wide range of natural language (words) to represent a given response process, allowing for ambiguity and redundancy, thereby enabling ambiguous input via gestures.

[0327] Furthermore, in the example in Figure 4F, when attempting to determine between two different registered characters, "DANCE" and "DASH," based on the first two characters, ID=2 and ID=3 will result in the same "DA." If the character in response 4A03 is "DA," then determining from the first two characters of the illustrated registered characters (for example, a prefix match search) will result in the two registered characters "DANCE" and "DASH" being matched. Thus, it is acceptable for multiple registered characters to match the character in response 4A03. The AI ​​response output device 10010 can then separately perform a process to select and determine one registered character from the multiple matching registered characters based on predetermined selection conditions (described later).

[0328] The processing of the AI ​​response output device 10010 using the correspondence table 4F00 as shown in Figure 4F generally follows the flow shown in (4) of Figure 4A above. This processing should be executed by the client application 8010 under the control of the control unit 1110. In other words, the registered character in the correspondence table 4F00 is searched for and selected from the characters of response 4A03, and the processing (response correspondence processing) associated with that registered character is determined.

[0329] In this system, the correspondence table may be updated as needed. For example, the original data of the correspondence table is registered and stored on a designated server on the network. A business operator or other entity updates the content of this original data of the correspondence table (i.e., information on the correspondence between registered characters and processes) as needed. Each response output device 10010 connected to the network accesses the designated server as needed, downloads and obtains the latest correspondence table data, and updates the correspondence table it holds on its own device.

[0330] Conventionally, the correspondence between gestures and processes in a gesture UI is based on the vendor's own rules. Target gestures may have predetermined shapes and directions. Users need to memorize the predetermined shapes and directions of gestures in advance by reading user manuals, etc. Furthermore, in conventional cases, updating the gesture UI involves updating the vendor's own rules, and the gesture recognition program, including image and line segment recognition processing, also needs to be updated. Users also need to relearn the correspondence between gestures and processes by reading user manuals, etc., which are distributed by some means. In contrast, in this embodiment, the target gestures of the gesture UI are based on natural language and correspond to registered natural language characters held in the processing correspondence table. This means that users only need to input gestures by associating the shape of natural language characters obtained from thinking about the desired process. In this case, users do not need to read user manuals, etc., distributed by the vendor in advance. Also, in this embodiment, when updating the gesture UI, it is sufficient to update the correspondence table 4F00 to update the information on the correspondence between registered characters and processes. Image and line segment recognition processing can directly utilize a multimodal language model (LLM). Therefore, there is no need to update gesture recognition programs or other components for image and line segment recognition processing. Even if gesture inputs are updated and new gesture inputs are added, since they are natural language (character gesture inputs), users do not need to memorize new gesture shapes. As before the update, users can simply perform gesture inputs that mimic the shape of natural language characters that they associate with the processing they desire.

[0331] [Processing flow using a correspondence table] Figure 4G shows an example of a processing flow using a correspondence table. This processing flow can be executed by the client application 8010 under the control of the control unit 1110 of the AI ​​response output device 10010. In step 4G01, the AI ​​response output device 10010 obtains the characters (e.g., two characters) of the character data of response 4A03 from the LLM. In step 4G02, the response output device 10010 searches the character data (e.g., two characters) for registered characters in the correspondence table 4F00, for example, using a prefix match condition. As a result, in step 4G03, it is confirmed whether there is one or more matching registered characters. If one or more matching registered characters are not found (NO), in step 4G04, for example, the error handling process described above is performed.

[0332] If one or more matching registered characters are obtained (YES), step 4G05 verifies whether there is one matching registered character or two or more. If there is one, the process proceeds to step 4G07, where that single registered character is determined. If there are two or more, in step 4G06, the response output device 10010 performs a process to select one registered character from the two or more registered characters. For example, one method is for the response output device 10010 to perform a random number selection process and select one registered character using a random number.

[0333] In step 4G07, one registered character is determined, and the response output device 10010 determines the response processing associated with that registered character from the correspondence table 4F00. In step 4G08, the response output device 10010 executes the determined response processing.

[0334] Another example of error handling in step 4G04 is described below. Even if user 230 inputs a gesture operation with characters they imagine, and the resulting characters recognized by the AI ​​(LLM) have the correct character code, there may be cases where there is no registered character corresponding to it in the correspondence table 4F00. In this case, the response handling process cannot be determined, so if no countermeasures are taken, there will be no output or response to user 230, or only error messages will be output. Such a situation is not desirable. The following countermeasures may be taken: If the input characters are not registered in the correspondence table 4F00, the response output device 10010 may select and determine a predetermined response handling process. For example, one may be randomly selected from several predetermined response handling processes. An example of a response handling process in this case may be a predetermined action or utterance by a character. For example, multiple types of standby motions and multiple types of standard standby phrases may be prepared, and one may be selected from these. The character may perform a predetermined idle pose motion and also utter standard phrases such as "Hello" or "It's a nice day today" to deflect the question. This allows the user 230 to receive at least some output or response. Furthermore, even if the characters recognized by the AI ​​(LLM) are in the correct character code, situations where they are not registered in the correspondence table 4F00 can be avoided by increasing the amount of information in the correspondence table 4F00. In other words, for all combinations of input characters corresponding to the input rules (language used, character type, number of characters) displayed on the gesture operation input screen shown in Figure 4C(2) and Figure 4D(1), at least one registered character and the corresponding process should be stored in the correspondence table 4F00. For example, the combinations of input characters corresponding to the input rules should consider all characters included in the range of the predetermined character type for the language used in the predetermined character code. For example, if the input rules are Japanese, katakana, and 1 character, and the full-width code of Shift JIS is used as the character code, there are 86 katakana characters.The correspondence table 4F00 should store at least one registered character that corresponds to each of the 86 inputs, along with the corresponding process. As another example, if the input rules are Japanese, katakana, and 2 characters, and the full-width code of Shift JIS is used as the character code, there are 86 katakana characters, and since there are 2 characters, the number of combinations is 86 squared, or 7396. In this case, the correspondence table 4F00 should store at least one registered character that corresponds to each of the 7396 inputs, along with the corresponding process. In this way, even if the characters recognized by AI (LLM) have the correct character code, it is possible to avoid the error state where there is no corresponding registered character in the correspondence table 4F00.

[0335] An example of the processing in step 4G06 is described below. When user 230 inputs a gesture, and the AI ​​(LLM) recognizes a character (e.g., "DA"), there may be multiple registered characters that match that character in correspondence table 4F00 (e.g., "DANCE", "DASH"). In this case, the response output device 10010 selects one registered character in a predetermined way and determines one processing action to be associated with that registered character. The following are examples of methods for doing this.

[0336] The first method involves using a random number to select a single registered character and process. For example, when the input is the character "DA", the response output device 10010 randomly selects one of two candidate registered characters, "DANCE" and "DASH". For example, on the first input, "DASH" is selected, and the process run() associated with the selected "DASH" is determined. Similarly, when the same character is input a second time or later, the decision can be made using a random number.

[0337] As a variation of the first method, for example, if "DASH" is selected the first time, the system may be controlled to select a different registered character, "DANCE," the second time. In other words, in the variation, while using random numbers, the system may be controlled so that, in the case of the same character input, different registered characters and processes are selected sequentially as much as possible each time.

[0338] The second method involves selecting a single registered character and process in a predetermined priority order. If there are multiple registered characters corresponding to the input character in the correspondence table, a priority order is set among these multiple registered characters. For example, for the two registered characters "DANCE" and "DASH", priority = 1 is set for "DANCE" and priority = 2 for "DASH". In this case, the response output device 10010, upon the first input of the character "DA", selects "DANCE" with priority = 1 and determines the process to dance(). Upon the second input of the same character, it selects "DASH" with priority = 2 and determines the process to run(). For subsequent inputs, the selection is again made in order of highest priority.

[0339] As a variation, a selection probability may be set for each registered character among multiple registered characters that correspond to the same character. For example, "DANCE" could be set to 70%, "DASH" to 30%, and so on.

[0340] Another possible approach is as follows: If there are multiple registered characters (e.g., "DASH", "DANCE") for a given input character (e.g., "DA"), the user interface can display options to prompt the user 230 to determine which registered character they intend. For example, it could display, "Is it 'DASH' or 'DANCE'?" By selecting from the options in response to the prompt, the user 230 can decide and execute a single process.

[0341] Furthermore, when using the history function (a function that allows you to select and execute characters from the history) described later (Figure 4J), if there are multiple registered characters and processes corresponding to the same input character as described above, the following method may be used.

[0342] The first method is to select and execute the same registered characters and process that were used in the past. For example, suppose the first time, when the input character is "DA", the registered character "DASH" is randomly selected and the process run() is executed. The character "DA" (meaning the registered character "DASH") is registered in the history. User 230 selects this character "DA" from the history. Then, for the second time, for the same character "DA", instead of random selection, the process run() associated with the same registered character "DASH" as the first time is selected and executed again.

[0343] The second method involves applying a similar random number selection method when selecting characters from the history and executing a process. For example, suppose the first time, when the input character is "DA", the registered character "DASH" is randomly selected and the process run() is executed. The character "DA" is registered in the history. User 230 selects this character "DA" from the history. Then, for the second time, random number selection is performed for the same character "DA", and if, for example, the registered character "DANCE" is selected, the process dance() is executed.

[0344] [Example of output display] Figure 4H shows an example of output display (character movement example as a specific example) in the user interface of the AI ​​response output device 10010 in response to the execution of a response processing. Figure 4H shows the case of English alphabet characters. The table in Figure 4H shows several display examples related to character video 19051 (Koto), corresponding to the examples of ID=7,1,3,4,6 in Figure 4F. Display example 4H00 is a display of the idle state in response to the idle() process. Display example 4H01 is a display of the jumping action in response to the jump() process. Display example 4H02 is a display of the running action in response to the run() process. Display example 4H03 is a display of the singing action in response to the sing() process. Display example 4H04 is a display of the turning action in response to the turn() process.

[0345] In the AI ​​response output device 10010, which serves as a character display device, the motion of the character corresponding to the response processing is executed, and the rendered image (character image 19051) is displayed on the screen of the display unit 10011. In this embodiment, the AI ​​response output device 10010 has a function to generate a character image by rendering processing based on a dataset including a pre-held 3D model. This rendering processing may be performed, for example, by the video control unit 1160. The AI ​​response output device 10010 generates a character image by rendering processing when the response processing is executed. Alternatively, it may store character image data that has been rendered in advance, read it, and play it back. This playback processing of character image data may be performed, for example, by the video control unit 1160. Furthermore, rendering processing may be performed on an external server, and the AI ​​response output device 10010 may acquire the rendering result video data via the communication unit 1132, and the video control unit 1160 may play it back.

[0346] The output of a character in response processing is typically a video / animation that depicts the character's actions, but it is not limited to this; it may also be a still image or other type of output.

[0347] Furthermore, the character's output in response to the processing is not limited to screen display; it may also be text display or audio output. For example, the character's speech / conversation text (string) may be displayed on the screen, or the speech / conversation audio may be output from a speaker.

[0348] Here, in the output of the user interface, not only may a character image corresponding to the response processing be displayed on the screen, but character data or registered characters associated with that response processing may also be displayed on the screen simultaneously. Displaying them simultaneously in association with each other has the advantage of making it easier for the user 230 to remember the relationship between the characters of the gesture input and the character movements that are output.

[0349] In Figure 4H, the gesture input character display (in other words, the gesture recognition character display) 4H11 is the case where the character "JU" is displayed in a frame and in a button-like manner simultaneously with the character video display in display example 4H01. The registered character display 4H12 is the case where the registered character "JUMP" from correspondence table 4F00 is displayed simultaneously with the character video display in display example 4H01. Only the gesture input character display 4H11 may be displayed, only the registered character display 4H12 may be displayed, or both may be displayed. The registered character display may be displayed in the form of a button or similar. Alternatively, when the user 230's fingers approach the gesture input character display 4H11, the corresponding registered character display or other detailed information (for example, the meaning of the response processing, for example, a character movement explanation, etc.) may be displayed.

[0350] FIG. 4I similarly shows an example of output display according to the execution of response handling, with an example in Japanese katakana characters. Each display example 4I00 - 4I04 (character video content) in FIG. 4I is the same as in FIG. 4H. Also, for the gesture input character display (or gesture recognition character display) 4I11 in FIG. 4I, when the character "ジ" is displayed surrounded by a frame in a button-like manner simultaneously with the character video display of display example 4I01. The registered character display 4I12 is the case where the registered character "ジャンプ" in correspondence table 4F00 is displayed simultaneously with the character video display of display example 4I01.

[0351] <所提供の内容には、このタグに対応する翻訳可能なテキストがありません。 As in the above example, not limited to simply displaying several characters of natural language (character data or registered characters) corresponding to the gesture determination result as character images, it may also be displayed as a UI object such as an icon or a mark or a button, like in the example of gesture input character display 4H11.

[0352] As will be described later, a character button such as the above gesture input character display (or gesture recognition character display) can be touched by user 230 and functions to cause the response handling to be re-executed.

[0353] Since FIG. 4H is an example of English input, the display of registered characters associated with the response handling is also displayed in English (alphabet) such as "JUMP". Since FIG. 4I is an example of Japanese input, the display of registered characters associated with the response handling is also displayed in Japanese (katakana) such as "ジャンプ".

[0354] As in the above example, by also displaying the characters associated with the response handling together, user 230 who performed the gesture operation input can understand that their gesture input indicated by the displayed characters / icons is recognized as a gesture determination result by the character display device, and the response handling of the result is output as a character action associated with the displayed characters.

[0355] In the example shown in Figure 4I, the character "タ" (gesture input character display 4I13) indicating the standby state ("Taiki") and the character "タ" (gesture input character display 4I14) indicating the turn are the same character. In other words, the gesture input character is the same, and there are two or more registered characters and corresponding response processes associated with it. In this case, the display manner of the gesture input character display (in other words, the gesture recognition character display) may be made different. In the illustrated example, the buttons and characters are displayed in different colors in gesture input character display 4I13 and gesture input character display 4I14. This makes it easier for the user 230 to recognize that there are different response processes (character actions) for the same character, and makes it easier for the user 230 to remember their correspondences.

[0356] The manner in which the above gesture input characters and registered characters are displayed may be expressed not only by color, but also by shape, size, font, etc.

[0357] [UI Examples (1)] Figure 4J shows a specific example of output in the user interface (another configuration example). The gesture operation input described above becomes more cumbersome for user 230 if it is performed repeatedly. Therefore, as an example of a user interface, Figure 4J includes a gesture operation input history function to make it easier for user 230 to input a gesture operation input that has been input and recognized / determined once.

[0358] In Figure 4J, screen 4101 shows, for example, a case where the character "DA", corresponding to the registered character "DASH", is entered via gesture input, resulting in the display example 4H02 of character 19051 showing a running motion. In addition to the display example 4H02, the bottom also shows the gesture input character display 4J11 ("DA") and the registered character display 4J12 ("DASH"), as explained in Figure 4H.

[0359] On this screen 4101, there is first a "1more" icon 4112. This is a button that makes the character 19051 perform the same action (for example, running) that was previously performed in response to a gesture input. When user 230 touches the "1more" icon 4112, the AI ​​response output device 10010 re-executes the response processing (for example, run()) that was executed immediately before. Alternatively, it may re-read and play the previously executed video data related to the response processing.

[0360] Furthermore, the lower part of screen 4101 has a gesture input history display (history area) 4J13. In the gesture input history display 4J13, gesture input history icons 4121 are displayed as a history of past input characters. The gesture input history icons 4121 display one or more past gesture input characters (character data of gesture recognition result / judgment result) as a history of gesture operation input (corresponding gesture recognition result), in the form of icons / marks / buttons, etc. For example, going backward from the immediately preceding character "DA", there are "TU", "JU", and "SI".

[0361] The gesture input history icon 4121, including the button for the previous gesture input character display ("DA") 4J11, can be touch-operated by the user 230. When the user 230 touches a desired gesture input history icon 4121, the AI ​​response output device 10010 determines the response processing associated with the registered character from the character indicated by the gesture input history icon 4121 (the character of the gesture operation input) and executes the response processing again. Alternatively, it may read and play back the video data that has already been executed related to the response processing. This makes it possible to make the character 19051 associated with that character perform the action again.

[0362] Furthermore, in this case, if there are multiple registered characters based on the character indicated by the gesture input history icon 4121, a single response processing may be determined by performing a similar random number selection, as described above. Alternatively, by keeping information of response processing previously determined and executed for that character in the history, the same response processing that was previously determined and executed may be selected based on that character, even if there are multiple registered characters.

[0363] In the illustrated example, the gesture input history display 4J13 shows older history on the far left and newer history on the right. This history can be rotated in direction 4122 by touch operation. For example, a swipe operation with a finger 4202 can rotate the gesture input history icon 4121 in direction 4122 to display older history.

[0364] Furthermore, the gesture input history display 4J13 may also include not only a gesture input character display (e.g., "TU") but also a registered character display (e.g., "TURN"). Alternatively, only one of the gesture input character display or registered character display may be provided. At least one of them accepts touch operations, etc.

[0365] Furthermore, as shown in this screen 4101, a favorite gesture registration display (favorite area) 4131 may be provided. The user 230 can select a gesture input history icon 4121 that they want to register as a favorite from the gesture input history display 4J13 and drop it onto the favorite gesture registration display 4131 by dragging with their finger 4202. As a result, the AI ​​response output device 10010 registers the character corresponding to the selected gesture input history icon 4121 (for example, "JU") as a favorite gesture input icon 4J14 and displays it fixedly on the favorite gesture registration display 4131.

[0366] User 230 can touch the favorite gesture input icon 4J14 registered in the favorite gesture registration display 4131. When User 230 touches the desired favorite gesture input icon 4J14, the AI ​​response output device 10010 determines the registered character and corresponding response processing from the character indicated by that favorite gesture input icon 4J14, and then executes the corresponding response processing again (or re-reads and plays back the previously executed video data, etc.). This allows the character 19051 corresponding to that character to perform the action again.

[0367] In addition, in the case of the favorite gesture registration display 4131, similar to the gesture input history display 4J13, if there are multiple registered characters for the character indicated by the favorite gesture input icon 4J14, it may be decided again by random number selection, or the same response processing that was previously decided and executed may be decided.

[0368] Furthermore, in the favorite gesture registration display area 4131, the registered text (favorite registration text) may be displayed along with, or instead of, the text of the favorite gesture input icon 4J14.

[0369] As shown in the example above, by using various UIs, it is possible to enable user 230 to perform response processing (character movement) with only simple touch operations on icons, without having to repeatedly perform the same gesture input.

[0370] [UI Examples (2)] Figure 4K shows a specific example of output in the user interface (another configuration example). Screen 4101 in Figure 4K shows a case where, for example, the character "SI" corresponding to the registered character "SING" is entered via gesture operation in English input, and as a result, the singing motion of character 19051 is displayed as display example 4H03. In addition to the motion of character 19051, lyrics text, audio, or music may also be output. Furthermore, along with the display example 4H03, there are also gesture input character display 4K11 ("SI") and registered character display 4K12 ("SING") at the bottom.

[0371] In Figure 4K, screen 4101 includes, in addition to the various UI elements mentioned above, a preset button display area 4141. The preset button display area 4141 displays preset buttons. For example, there are preset button 4K13 ("CC" display) and preset button 4K14 ("DC" display). These preset buttons are UI elements such as icons, marks, and buttons that perform predetermined actions / processes different from the response processing based on gesture input.

[0372] The preset button 4K13 (indicated as "CC") is an acronym for Character Change and is used to change or switch the character displayed on the screen. For example, several characters are available as changeable characters, as explained in Figure 3A, and the user can change or switch to a character selected from these. When user 230 touches the preset button 4K13 (indicated as "CC"), the AI ​​response output device 10010 controls the character displayed on the screen 4101 to toggle (cycle) between, for example, character 19051 (Koto), character 19052 (Tom), and character 19053 (Necco) in Figure 3A, with each touch operation.

[0373] The preset button 4K14 ("DC" displayed) is an acronym for Dress Change and is used to change or switch the costume (appearance) of the character displayed on the screen. For example, multiple costumes are available for each character that can be changed, and the user can change or switch to the costume selected from these. For example, when character 19051 (Koto) is displayed, if user 230 touches the preset button 4K14 ("DC" displayed), the AI ​​response output device 10010 controls the costume of the character displayed on screen 4101 to toggle (cycle) between multiple costumes with each touch operation. The specifics of Dress Change will be described later.

[0374] For example, in addition to the preset button 4K13 ("CC" displayed), explanations such as "Character Change" or "You can change characters" may be displayed nearby, or such explanations may be displayed when the finger 4201 approaches the preset button 4K13.

[0375] [UI Examples (3)] Figure 4L shows a specific example of output in the user interface (another configuration example). In Figure 4L, the gesture input is in Japanese. Screen 4101 in Figure 4L shows a case where both gesture input that imagines the aforementioned characters (natural language) (referred to as "character gesture input") and gesture input that is different (referred to as "non-character gesture input") can be accepted in combination. In other words, character gesture input is natural language gesture input, and non-character gesture input is non-natural language gesture input.

[0376] Non-character gesture input is a gesture input method that uses a predetermined, simple trajectory, different from the trajectory that represents characters (registered characters in the correspondence table). Non-character gesture input can be applied using simple touch gestures with the fingers 4201, such as horizontal or vertical swipes.

[0377] Non-character gesture inputs do not require recognition and judgment by AI (LLM). The AI ​​response output device 10010 detects non-character gesture inputs and executes predetermined processing (rule-based processing, not AI processing) associated with the detected non-character gesture inputs.

[0378] Furthermore, non-character gesture inputs may have different functions / processes executed depending on the starting point (touch) position (location or area within the screen 4101) when the non-character gesture input is performed. For example, even with the same non-character gesture input (e.g., a horizontal swipe), the first process may be executed in the case of the first position / area, and the second process may be executed in the case of the second position / area.

[0379] Non-text gesture input does not require the display of icon buttons (explanatory displays are acceptable). This reduces the amount of content and background clipping on screen 4101.

[0380] The acceptance and processing of two types of gesture inputs—character gesture inputs and non-character gesture inputs—may be handled separately by time / mode, or by separating areas within the same screen. In the example in Figure 4L, character gesture input is accepted and response processing is executed in the gesture input mode corresponding to the aforementioned button 4111. Then, non-character gesture input is accepted in areas 4151 and 4152, which partially overlap with the character 19051 on screen 4101.

[0381] The AI ​​response output device 10010 accepts and detects a swipe operation 4153 (e.g., a horizontal swipe) as a non-character gesture input in region 4151 (the upper region around the character). When the AI ​​response output device 10010 detects a swipe operation 4153, it changes the displayed character according to the details of the operation. This is the same process as the one performed by the preset button 4K13 ("CC" display) described above (Figures 3A and 4K), and can be achieved, for example, by toggling.

[0382] The AI ​​response output device 10010 accepts and detects a swipe operation 4154 (e.g., a horizontal swipe) as a non-character gesture input in region 4152 (the lower region around the character). When the AI ​​response output device 10010 detects a swipe operation 4154, it changes the costume of the displayed character according to the details of the operation. This is the same process as the one performed by the preset button 4K14 ("DC" display) described above (Figure 4K), and can be achieved, for example, by toggling.

[0383] Another example of using two non-character gesture inputs, "CC" and "DC," together is to associate horizontal swipes with "CC" processing and vertical swipes with "DC" processing.

[0384] FIG. 4M is a modification of FIG. 4L, and shows an example in which, within the same screen, an area for receiving character gesture operation inputs and an area for receiving non-character gesture operation inputs are provided separately. The screen 4101 of FIG. 4M is divided into two parts: a lower area 4M01 and an upper area 4M02. In the lower area 4M01, character gesture operation inputs are received. In the upper area 4M02, non-character gesture operation inputs are received. Note that an image or description representing the area may be displayed so that the distinction between the lower area 4M01 and the upper area 4M02 is easy for the user 230 to understand. In the lower area 4M01, for example, when a character gesture operation input 4M03 in which the character "SI” is drawn is detected, as a result of the response 4A03 from the LLM, the action of the character 19051 singing (display example 4H03) is displayed. In the upper area 4M02, when a horizontal swipe operation 4M04, for example, is detected as a non-character gesture operation input, a process corresponding to "CC” (character change) is performed.

[0385] In the case of an operation with a simple trajectory such as a horizontal swipe operation, the trajectory is similar to a simple character such as the character "I” or "—”, so misrecognition thereof is not desirable. Therefore, as in the above example, by defining character gesture operation inputs and non-character gesture operation inputs separately by time (mode) or area, combined use while preventing misrecognition becomes possible. As a method of providing the above two areas, they may be divided into an area that overlaps with the character video and an area that does not overlap.

[0386] FIG. 4N shows a specific display example of the above dress change in the output of the user interface. The AI response output device 10010 holds, for example, character video data with different appearances such as clothing for each same character. When the AI response output device 10010 detects an operation for character change in FIG. 4K or FIG. 4L, it uses the held character video data to switch the display of the clothing (appearance) of the character on the screen. Further, the AI response output device 10010 may display characters or the like indicating the name of the clothing together with the switching of the clothing display.

[0387] In the example in Figure 4N, the display example (character video) 4N01 with costume ID=1 is the costume named "Classic". The display example (character video) 4N02 with costume ID=2 is the costume named "Party". The display example (character video) 4N03 with costume ID=3 is the costume named "Sports". Text display 4N11 shows the costume name.

[0388] [Sensor] Figure 4O shows an example of a sensor configuration capable of detecting gesture input. (1) is a touch panel 4O01 equipped with a touch sensor 4O02. The touch panel 4O01 can be configured on the surface of the display unit 10011 in Figure 1B. The touch sensor 4O02 detects whether or not a finger or the like has touched a pixel on the screen, and can obtain image data or line segment data as the gesture transmission data 4A01. If the gesture transmission data 4A01 is image data, for example, in pixels corresponding to a two-dimensional screen, pixels that have been touched by a finger are set to 1, and pixels that have not been touched by a finger are set to 0. For example, the pixels that make up the trajectory of the letter "J" are the pixels that have been touched by a finger. If the gesture transmission data 4A01 is line segment data, for example, in pixels corresponding to a two-dimensional screen, the coordinates of the positions of pixels that have been touched by a finger are recorded over time and discretely, and the line segment data is obtained by continuously connecting the coordinates of the pixels that have been touched by a finger. For example, the trajectory of the coordinate system that makes up the letter "J" over time is a discrete record of the position of the pixel where a finger touches.

[0389] (2) is an aerial operation detection sensor 1351 provided in, for example, the aerial floating image display device 185 (Figure 1C). This shows an example of the aerial operation detection sensor 1351 in Figure 1K. As one example configuration, the aerial operation detection sensor 1351 emits light in a direction parallel to the spatial region including the display range (screen) of the aerial image 195 (indicated by a dashed arrow). When there is an aerial operation (gesture input) by fingers, etc. on the aerial image 195, the light is reflected by the fingers, etc., and the aerial operation detection sensor 1351 detects the reflected light. From such sensing results, the aerial operation detection sensor 1351 can detect the trajectory of the gesture by fingers, etc. on the display range (screen) of the aerial image 195. The configuration of the image data and line segment data as gesture transmission data 4A01 is the same as in (1) above.

[0390] (3) is the imaging unit 1180 (camera). This is an example of the imaging unit 1180 shown in Figures 1B and 1K. The imaging unit 1180 (camera) is positioned, for example, on the far side from the screen 4O03 of any display device (for example, it may be an aerial image 195, a virtual image of the HUD 184, or a projected image of the projector 186. It may also simply be a screen to be captured set in empty space). In the case of the HMD 183, the imaging unit 1180 (camera) is positioned on the near side from the screen 4O03. The imaging unit 1180 (camera) captures images of the screen 4O03 and the user's fingers, etc. The user, who is on the near side of the screen 4O03, performs gesture inputs to the screen 4O03 using their fingers, etc. Based on the video captured by the imaging unit 1180 (camera), the trajectory of the gestures in time can be captured. The imaging unit 1180 may also consist of multiple cameras or distance sensors. The structure of the image data and line segment data as gesture transmission data 4A01 is the same as in (1) above.

[0391] In the example shown in Figure 4A, the trajectories of the two characters "JU" are separated in the gesture transmission data 4A01. However, depending on the sensor or other method, the trajectories may not be separated, appearing as a single continuous line. Even in the case of such gesture transmission data, the method can be applied if character recognition by AI (LLM) is possible.

[0392] As described above, Embodiment 4 provides a more suitable response output technology. Furthermore, Embodiment 4 provides a suitable interface for operation input using gestures.

[0393] In this response output system, user 230 simply needs to imagine natural language characters related to the actions they want character 19051 (response output device 10010) to perform, and then perform a gesture operation input (character gesture operation input) that represents those natural language characters. User 230 does not need to memorize or needs to memorize a large amount of rules that differ for each device or interface in the conventional technology examples, such as unique gestures.

[0394] In conventional interfaces using command buttons and icons, the correspondence between each command button or icon to a specific process is often defined by the interface manufacturer or vendor. Therefore, users typically need to read manuals or other resources to learn the correspondences for each specific interface. To address this issue, Example 4 uses natural language (words, etc.) that is generally recognizable to humans as the input (gesture input) that corresponds to the desired process or action. This leverages the fact that humans possess a vast amount of tacit knowledge of natural language (words, etc.) beforehand. Therefore, users do not need to learn the correspondences from manuals or other resources beforehand, or the effort required to learn them is significantly reduced. Furthermore, Example 4 incorporates features such as correspondence tables and UI displays to facilitate the correspondence between inputs and processes.

[0395] In Example 4, User 230 simply needs to imagine a desired natural language (e.g., a word) to represent a target process or action (e.g., character movement) and input a character gesture corresponding to that natural language. The system uses AI (LLM) to recognize this natural language (characters / strings) and associates it with the target process or action. The system is designed to output a response that associates at least some kind of process with User 230's gesture input. Even if User 230 inputs a random character gesture, some kind of response will be output (including the error handling mentioned above). Furthermore, when User 230 inputs a character gesture, a process or action that User 230 did not anticipate may be associated with that character through a mechanism such as random number selection. This can be particularly interesting in applications involving character movement.

[0396] Since each of the response output devices and response output systems described in the above embodiments performs information processing, they may also be referred to as information processing devices and information processing systems, respectively.

[0397] The technology described in this embodiment makes it possible to provide a more suitable artificial intelligence response output technology. Such artificial intelligence response output technology is expected to be introduced into higher quality and more reliable infrastructure. As this technology is introduced into infrastructure, it can contribute to economic development and support for human welfare that focuses on affordable and equitable access for all. This will contribute to "Goal 9 of the United Nations Sustainable Development Goals (SDGs)," namely "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."

[0398] Furthermore, the technology according to this embodiment makes it possible to provide a more suitable artificial intelligence response output technology. Such artificial intelligence response output technology is expected to be introduced into public transport facilities to improve access to transportation systems for vulnerable people. As this technology is introduced into public transport, it can contribute to improving transportation safety through the expansion of public transport, and to realizing access to sustainable transport systems that are safe, inexpensive, and easily accessible to all. This will contribute to "SDG 11: Make cities and human settlements inclusive, safe, resilient and sustainable," as advocated by the United Nations.

[0399] Although various embodiments have been described in detail above, the present invention is not limited to the embodiments described above, but includes various modifications. For example, the embodiments described above are detailed explanations of the entire system in order to explain the present invention in an easy-to-understand manner, and are not necessarily limited to those having all the described configurations. Furthermore, it is possible to replace parts of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add configurations from other embodiments to the configuration of one embodiment. In addition, it is possible to add, delete, or replace parts of the configuration of each embodiment with other configurations. [Explanation of Symbols]

[0400] 10010...Artificial intelligence response output device, 10011...Display unit, 10028...Local LLM processing unit, 1107...Operation input unit, 1110...Control unit, 1132...Communication unit, 1140...Audio output unit, 1139...Microphone, 1160...Video control unit, 1170...Storage unit, 1180...Imaging unit.< / audio> < / video>

Claims

1. A response output device, A sensor that accepts gesture input from the user that mimics natural language characters, Control unit and Output section, Equipped with, The control unit generates an instruction sentence including natural language and gesture transmission data generated based on the operation input to the sensor, transmits the instruction sentence to a multimodal language model, acquires character data of the gesture recognition result as a response generated by the inference performed by the multimodal language model based on the instruction sentence, determines the processing to be performed for the gesture operation input by the user based on the acquired character data and the correspondence table held by the response output device, and performs control by outputting an output from the output unit based on the determined processing. Response output device.

2. In the response output device according to claim 1, The natural language included in the instruction statement includes information specifying the language, character type, or number of characters for the characters represented in the gesture communication data. Response output device.

3. In the response output device according to claim 1, The gesture transmission data is image data or line segment data in which the characters are drawn. Response output device.

4. In the response output device according to claim 1, The control unit performs the following control functions: it selects registered characters from the correspondence table held by the response output device using the characters of the character data as a response obtained from the multimodal language model, determines a response response process which is a process associated with the selected registered characters, executes the determined response response process as a process for the gesture operation input by the user, and outputs the execution result from the output unit. Response output device.

5. In the response output device according to claim 4, The aforementioned correspondence table contains multiple registered combinations of the response processing and the registered characters that represent the response processing. Response output device.

6. In the response output device according to claim 4, In the aforementioned correspondence table, multiple registered characters representing the same single response processing are registered in association with that response processing. Response output device.

7. In the response output device according to claim 4, The control unit's control of the response processing includes searching for registered characters in the correspondence table using the characters in the character data, and if multiple registered characters match, selecting one registered character from the multiple registered characters according to predetermined selection conditions. Response output device.

8. In the response output device according to claim 7, The aforementioned selection condition is a random number. Response output device.

9. In the response output device according to claim 4, Along with the execution of the response processing, at least one of the characters from the character data and the selected registered characters is output to the output unit. Response output device.

10. In the response output device according to claim 9, The output unit is capable of outputting video, The output unit outputs an object representing at least one of the characters in the character data and the selected registered character as an object that can be manipulated by the user. The control unit, upon detecting an operation on the object, performs the response processing associated with the registered character corresponding to the object. Response output device.

11. In the response output device according to claim 9, The output unit is capable of outputting video, The output unit outputs an object representing at least one of the characters in the character data and the selected registered character, as an object that indicates the history of user operations and is operable by the user. When the control unit detects an operation on the object indicating the user's operation history, it performs the response processing associated with the character of the character data indicated by the object or the registered character. Response output device.

12. In the response output device according to claim 9, The control unit performs control to register an object representing at least one of the characters of the output character data and the selected registered character as a favorite, in response to the user's operation, so that it can be operated by the user. When the control unit detects an operation on the object registered as a favorite, it performs the response processing associated with the characters of the character data indicated by the object or the registered characters. Response output device.

13. In the response output device according to claim 1, In addition to the gesture input that mimics natural language characters, the sensor can also receive non-character gesture input. When the control unit detects a non-character gesture input, it performs control to execute a predetermined process associated with the non-character gesture without using the multimodal language model. Response output device.

14. In the response output device according to claim 4, The output unit is capable of outputting video, The response processing described above is a process that causes the character image displayed on the output unit's screen to perform an action represented by the registered character. Response output device.

15. A response output method in a response output device, The steps performed by the response output device include: A gesture input step that accepts gesture input from the user that mimics natural language characters, An instruction sentence generation step that generates an instruction sentence including natural language and gesture transmission data generated based on the gesture operation input entered in the gesture operation input step, A transmission step of sending the instruction sentence, which includes the natural language and the gesture transmission data generated in the instruction sentence generation step, to a multimodal language model; The acquisition step involves acquiring character data of the characters of the gesture recognition result as a response, which are generated by the inference performed by the multimodal language model based on the instruction sentence. Based on the acquired character data and the correspondence table held by the response output device, the output step determines the processing for the gesture operation input by the user and outputs based on the determined processing, A response output method having a response output.

16. In the response output method according to claim 15, In the output step, based on the characters of the character data as a response obtained from the multimodal language model, the response output device selects registered characters recorded in the correspondence table it holds, determines a response response process which is a process associated with the selected registered characters, executes the determined response response process as a process for the user's gesture input, and outputs the execution result. Response output method.

17. An information processing device, A sensor that accepts gesture input from the user, Control unit and Equipped with, The control unit generates an instruction sentence including natural language and gesture transmission data generated based on the operation input to the sensor, transmits the instruction sentence to a multimodal language model, and obtains character data of the gesture recognition result as a response generated by inference performed by the multimodal language model based on the instruction sentence. Information processing device.

18. In the information processing apparatus according to claim 17, The natural language included in the instruction statement includes information specifying the language, character type, or number of characters for the characters represented in the gesture communication data. Information processing device.

19. In the information processing apparatus according to claim 17, The gesture transmission data is image data or line segment data in which the characters are drawn. Information processing device.

Citation Information

Patent Citations

  • Structural unit for tank construction

    JP1977008512A