Response output device, system, and method
The response output device optimizes user interaction by incorporating user-specified input and AI-generated responses, leveraging shared large-scale language models to enhance interaction and response quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MAXELL LTD
- Filing Date
- 2025-10-20
- Publication Date
- 2026-07-23
AI Technical Summary
Existing response output technologies using artificial intelligence lack sufficient configuration for optimal user interaction and response generation.
A response output device equipped with an interface for user input to specify character generation characteristics, a prompt creation mechanism, and a system to receive and output AI-generated responses, utilizing local or external large-scale language models for improved interaction.
Enhances the suitability and effectiveness of AI response output by allowing tailored user inputs and efficient resource utilization through shared large-scale language models, improving user interaction and response quality.
Smart Images

Figure JP2025036857_23072026_PF_FP_ABST
Abstract
Description
Response Output Device, System, and Method
[0001] The present disclosure relates to response output technology using artificial intelligence such as a language model.
[0002] Regarding response output technology using artificial intelligence such as a language model, for example, it is disclosed in Patent Document 1.
[0003] Japanese Patent Application Laid-Open No. 2019-528512
[0004] However, in the disclosure of Patent Document 1, the consideration regarding the configuration for more suitably providing the response output technology using artificial intelligence to the user was not sufficient.
[0005] An object of the present invention is to provide a more suitable response output technology.
[0006] In order to solve the above problems, for example, the configuration described in the claims is adopted. This application includes a plurality of means for solving the above problems. If an example is given, it may be configured as follows. A response output device that outputs a response from artificial intelligence, comprising an interface for artificial intelligence, in the interface, receiving an input for specifying the characteristics of a character to be generated by the artificial intelligence by the user, creating a prompt for generating a character reflecting the characteristics based on the input information, transmitting the prompt to the artificial intelligence, receiving character data reflecting the characteristics generated by the artificial intelligence, and outputting the character in the interface based on the received data.
[0007] According to the present invention, a more suitable response output technology can be provided. Other problems, configurations, and effects will be clarified in the description of the following embodiments.
[0008] This is a diagram showing an example of an artificial intelligence response output device and system according to one embodiment of the present invention. This is a diagram showing an example of an artificial intelligence response output device according to one embodiment of the present invention. This is a diagram showing an example of an artificial intelligence response output device according to one embodiment of the present invention. This is a diagram showing an example of a server according to one embodiment of the present invention. This is an explanatory diagram of an example of an artificial intelligence response output system according to one embodiment of the present invention. This is a diagram showing an example of an optical system and retroreflector for an aerial floating image display device according to one embodiment of the present invention. This is a diagram showing an example of an aerial floating image display device according to one embodiment of the present invention. This is a diagram showing an example of an optical system for an aerial floating image display device according to one embodiment of the present invention. This is a diagram showing an example of a retroreflector for an aerial floating image display device according to one embodiment of the present invention. This is a diagram showing an example of an aerial floating image display device according to one embodiment of the present invention. This is a diagram showing an example of an aerial floating image display device according to one embodiment of the present invention. This is an explanatory diagram of an example of an artificial intelligence response output device and system according to one embodiment of the present invention. This is an explanatory diagram of an example of a mobile information processing terminal according to one embodiment of the present invention. This is an explanatory diagram of an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. This is an explanatory diagram of an example of the display example of an artificial intelligence response output device according to one embodiment of the present invention. This is an explanatory diagram of an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. This is an explanatory diagram of an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. This is an explanatory diagram illustrating an example of the operation of an artificial intelligence response output device and system according to one embodiment of the present invention. This is an explanatory diagram illustrating an example of table information according to one embodiment of the present invention. This is an explanatory diagram illustrating an example of a conversation example used in describing one embodiment of the present invention. This is an explanatory diagram illustrating an example of a conversation example used in describing one embodiment of the present invention. This is a diagram showing the configuration and overview of a system including an artificial intelligence response output device according to one embodiment of the present invention. This is a diagram showing an example of the interface screen (list of feature items) of a response output device according to one embodiment of the present invention. This is a diagram showing an example of the interface screen (text input) of a response output device according to one embodiment of the present invention.This figure shows an example of the interface screen (image input) of a response output device according to one embodiment of the present invention. This figure shows an example of a display character in a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (character display) of a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (before feature modification) of a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (after feature modification) of a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (feature maintenance specification) of a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (character display of feature modification) of a response output device according to one embodiment of the present invention. This is an explanatory diagram regarding the exclusion function in a response output device according to one embodiment of the present invention. This is an explanatory diagram regarding reloading in a response output device according to one embodiment of the present invention. This is an explanatory diagram regarding multiple outputs in a response output device according to one embodiment of the present invention. This is an explanatory diagram regarding display control in an aerial floating image display device, which is a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (addition of feature items) of a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (history of instruction statements) of a response output device according to one embodiment of the present invention. This figure shows an example of the configuration of a system including an aerial floating image display device, which is a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (output modification, modification location specification) of a response output device according to one embodiment of the present invention. This figure shows an example of the interface screen (output modification, maintenance location specification) of a response output device according to one embodiment of the present invention. This figure shows the configuration and overview of a system including an artificial intelligence response output device according to one embodiment of the present invention. This figure shows an example of the input interface (image capture screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (input (exclusion / change) screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (input (select from list) screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (input (select from list) screen) of a response output device according to one embodiment of the present invention.This figure shows an example of the input interface (input (select from list) screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (input (select from list) screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the configuration of the feature list creation process of a response output device according to one embodiment of the present invention. This figure shows an example of the configuration of the feature list creation process of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (input (select from image) screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (detailed settings screen) of a response output device according to one embodiment of the present invention. This figure shows an example of the input interface (exclusion / change instruction GUI) of a response output device according to one embodiment of the present invention. This figure shows an example of the default exclusion / change process of a response output device according to one embodiment of the present invention. This figure shows an example of avatar output of a response output device according to one embodiment of the present invention. This figure shows an example of the avatar output screen of a response output device according to one embodiment of the present invention. This figure shows an example of the screen when a response output device is re-instructed according to one embodiment of the present invention. This figure shows an example of the screen before and after a re-instruction of a response output device according to one embodiment of the present invention. This figure shows an example of the screen when a response output device is reloaded according to one embodiment of the present invention. This figure shows an example of the screen before and after a reload of a response output device according to one embodiment of the present invention. This figure shows an example of multiple avatar output screens of a response output device according to one embodiment of the present invention. This figure shows an example of the screen during deformation of the response output device according to one embodiment of the present invention. This figure shows an example of the screen before and after deformation of the response output device according to one embodiment of the present invention. This figure shows an example of the screen for selecting a deformed image of the response output device according to one embodiment of the present invention. This is an explanatory diagram regarding the QR code generation function in a system including a response output device according to one embodiment of the present invention.
[0009] Embodiments of the present invention will be described in detail below with reference to the drawings. However, the present invention is not limited to the examples described herein, and various modifications and alterations are possible by those skilled in the art within the scope of the technical ideas disclosed herein. Furthermore, in all the figures used to illustrate the present invention, components having the same function are given the same reference numerals, and repeated descriptions may be omitted.
[0010] Furthermore, if the artificial intelligence response output device according to each embodiment of the present invention has a display screen, it may be called a display device. If the artificial intelligence response output device has a voice output function, it may be called a voice output device. The artificial intelligence response output device may simply be called an information processing device. A system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence response output system. Also, if the artificial intelligence response output device provides a response service of a large-scale language model, which is artificial intelligence, to the user and assists the user, the artificial intelligence response output device or the display output of the artificial intelligence response output device can become an artificial intelligence (AI) assistant for the user. Therefore, in this case, the artificial intelligence response output device may be called an AI assistant device or an AI assistant display device. Similarly, in this case, a system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an AI assistant system or an AI assistant display system. Also, in this case, since the artificial intelligence response output device becomes an interface between the user and artificial intelligence, it may be called an artificial intelligence interface device. In this case, a system including the artificial intelligence response output device and a large-scale language model server that holds a large-scale language model may be called an artificial intelligence interface system.
[0011] <Example 1> As Example 1 of the present invention, an artificial intelligence response output device and system that outputs a response from a large-scale language model artificial intelligence will be described.
[0012] An example of the artificial intelligence response output device 10010 of the present invention will be described using Figure 1A. Furthermore, an example of a system including the artificial intelligence response output device 10010 and the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 will be described in the case where the artificial intelligence response output device 10010 cooperates with the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 through communication or other means.
[0013] In the example shown in Figure 1A, the artificial intelligence response output device 10010 has a display unit 10011. In the example shown in Figure 1A, the display unit 10011 may be a flat panel display, a screen that projects images from the back, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which pixels emit light themselves. The display unit 10011 may also be equipped with a touch operation input sensor and configured as a touch panel.
[0014] In the example shown in Figure 1A, the voice output unit 1140 of the artificial intelligence response output device 10010 is composed of a speaker. The artificial intelligence response output device 10010 is also equipped with a microphone 1139 that can pick up the user's voice. Through voice input from the microphone 1139 and user operation input via the operation input unit described later, the artificial intelligence response output device 10010 can acquire user input that forms the basis of instructions (prompts) for the large-scale language model, which is an artificial intelligence.
[0015] The artificial intelligence response output device 10010 may be equipped with a local large-scale language model. In this case, the response of the large-scale language model may be output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.
[0016] Furthermore, the artificial intelligence response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001 and / or a multimodal large-scale language model server 20001, and output the response received from the large-scale language model server 19001 and / or the multimodal large-scale language model server 20001 as the display output of the display unit 10011 and / or the audio output of the audio output unit 1140.
[0017] Alternatively, the artificial intelligence response output device 10010 may also include a local large-scale language model and be configured to communicate with an external large-scale language model server 19001 having a large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model. In this case, the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be switched and output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140. Alternatively, the response generated based on both the response of the local large-scale language model and the response received from the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be output as the display output of the display unit 10011 and / or as the audio output of the audio output unit 1140.
[0018] The configuration when the artificial intelligence response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via a communication unit 1132. In the example in Figure 1A, the communication between the communication unit 1132 and the communication device 19011 is shown as a wireless example, but wired communication is also acceptable. The communication path from the communication unit 1132 to the communication device 19011 may have both wired and wireless sections, or it may pass through routers or repeaters. Similarly, the communication path from the communication unit 1132 to the Internet 19000 may also have both wired and wireless sections, or it may pass through routers or repeaters. The artificial intelligence response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via the communication device 19011 and the internet 19000. The configuration including the artificial intelligence response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.
[0019] In the following explanation, unless otherwise specified, the term "large-scale language model" should be understood as encompassing the local large-scale language model provided by the artificial intelligence response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
[0020] In the example shown in Figure 1A, the display unit 10011 displays elements in two display areas: an instruction display area 10051 where the user inputs instructions (prompts) to a large-scale language model which is artificial intelligence, and an artificial intelligence response display area 10061 which displays the response from the large-scale language model. In the example shown in Figure 1A, the instruction display area 10051 displays an icon 10052 representing the user, text 10053 such as natural language or software code as components of the instruction, an image 10054 as a component of the instruction, a video 10055 as a component of the instruction, and so on. In the example shown in Figure 1A, the artificial intelligence response display area 10061 displays an icon 10062 representing artificial intelligence or an artificial intelligence assistant, text 10063 such as natural language or software code as components of the response from artificial intelligence, an image 10064 as a component of the response from artificial intelligence, a video 10065 as a component of the response from artificial intelligence, and so on. Note that the display example of the display unit 10011 of the artificial intelligence response output device 10010 shown in Figure 1A is merely an example. Depending on the implementation example in which the artificial intelligence response output device 10010 is used, a different display from the example shown in Figure 1A may be used.
[0021] Here, we will explain large-scale language models. Large-scale language models are also referred to as LLMs (Large Language Models). Specifically, various models such as GPT-1, GPT-2, GPT-3, InstructGPT, and ChatGPT have been made publicly available. These technologies can be used in this embodiment as well. These large-scale language models are artificial intelligence models generated through extensive pre-training on natural language contained in a large number of documents and texts that exist in the human world. The number of parameters of these artificial intelligence models exceeds hundreds of millions. Furthermore, in addition to this, there are also models that incorporate reinforcement learning based on human feedback. An example of a base model is a model called Transformer. As an example of training these models, for example, Reference 1 is publicly available.
[0022] [Reference 1] Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https: / / arxiv.org / pdf / 2203.02155.pdf
[0023] These large-scale language models are capable of natural language translation, natural language proofreading, and natural language text summarization. More advanced models can even perform natural language question answering (also called dialogue or conversation), natural language suggestion generation, and programming code generation. Because these artificial intelligence models have a very large number of parameters, training requires enormous amounts of data and computing resources. Therefore, training this level of artificial intelligence for a specific application is extremely resource-inefficient. To address this, foundation models that can be applied to various uses are generated through large-scale pre-training. For example, the large-scale language model server 19001 and / or multimodal large-scale language model server 20001 shown in Figure 1A may be equipped with such large-scale language models and configured to be usable by various terminals via an API (Application Programming Interface). Alternatively, the artificial intelligence response output device 10010 shown in Figure 1A may be equipped with a local large-scale language model and configured to use it itself. These large-scale language models can be generated by performing large-scale pre-training separately, and the generated large-scale language models can be duplicated and provided to the large-scale language model server 19001, the multimodal large-scale language model server 20001, and the artificial intelligence response output device 10010. In this way, instead of performing pre-training for each application or terminal, duplicating the large-scale language model, which is the base model generated by performing large-scale pre-training, and using it on individual servers and terminals allows for the sharing of resource consumption used for training, resulting in better resource efficiency.
[0024] Furthermore, even if a large-scale language model is generated as a foundational model through extensive pre-training, it may be configured to perform additional training, such as transfer learning, on individual servers or devices, depending on the application and purpose.
[0025] Furthermore, large-scale language models can pre-train on natural language and perform input / output processing targeting natural language. In addition, multimodal large-scale language model artificial intelligence capable of processing not only natural language text information but also other types of information is also applicable to the embodiments of the present invention. Figure 1A shows a large-scale language model server 20001 having a multimodal large-scale language model. For example, specific examples of multimodal large-scale language model artificial intelligence include GPT-4 (see Reference 2) and Gato (see Reference 3), which have been made publicly available. These technologies may also be used in this embodiment. These multimodal large-scale language models are artificial intelligence models generated by performing large-scale pre-training on natural language and other types of information (e.g., images, videos, audio, etc.) contained in numerous documents and texts existing in the human world. Furthermore, there are also models that incorporate reinforcement learning based on human feedback. Hereinafter, information other than natural language text information, such as images, videos, and audio, may be referred to as non-natural language information sources.
[0026] [Reference 2] Open AI “GPT-4 Technical Report”, https: / / cdn.openai.com / papers / gpt-4.pdf [Reference 3] Scott Reed, et. al. “A Generalist Agent”, https: / / arxiv.org / pdf / 2205.06175.pdf
[0027] Next, using Figure 1B, we will describe an example configuration of an artificial intelligence response output device 10010 that receives user input to artificial intelligence such as these large-scale language models and outputs a response from the artificial intelligence such as the large-scale language model to the user input.
[0028] The artificial intelligence response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a position sensor 1114, a local LLM processing unit 10028, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, a clock 1196, and the like. The artificial intelligence response output device 10010 may have a large screen, such as a so-called monitor or television.
[0029] The display unit 10011 may be a flat panel display, a screen that projects images from the back, or a display that projects an optical image into the air to show a floating image. If the display unit 10011 is a flat panel display, it may be a liquid crystal display having a liquid crystal panel and a backlight. Alternatively, the display unit 10011 may be a plasma display. The display unit 10011 may be an organic EL display in which pixels emit light themselves. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be equipped with a touch operation input sensor and configured to accept touch operation input from the user 230's finger. In this case, the display unit 10011 may be configured as a touch panel. Through the user's operation input via the touch panel, the artificial intelligence response output device 10010 can acquire user input that forms the basis of instructions (prompts) to the large-scale language model, which is artificial intelligence.
[0030] The communication unit 1132 may be configured with a Wi-Fi® communication interface, a Bluetooth® communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 1132 of the artificial intelligence response output device 10010 can communicate with the communication device 19011 connected to the Internet 19000. The communication path from the communication unit 1132 to the communication device 19011 may include both wired and wireless sections, and may also pass through routers or repeaters. In the case of a wired connection, the communication unit 1132 may have an Ethernet® connection interface as hardware and communicate using a LAN communication method. This allows the artificial intelligence response output device 10010 to communicate with various servers connected to the Internet 19000.
[0031] The artificial intelligence response output device 10010 is equipped with a control unit 1110 such as a CPU and a memory 1109, and the control unit 1110 controls the display unit 10011 and the communication unit 1132, etc.
[0032] The power supply 1106 converts the AC current input from the outside via the external power input interface 1111 into DC current and supplies the necessary DC current to each part of the artificial intelligence response output device 10010. The secondary battery 1112 stores the power supplied from the power supply 1106. In addition, the secondary battery 1112 supplies power to each part that requires power via the external power input interface 1111 when power is not supplied from the outside.
[0033] The operation input unit 1107 is, for example, an operation button, a signal receiving unit such as a remote controller, or an infrared light receiving unit, and inputs signals for operations other than touch operations by the user to the touch operation input sensor of the display unit 10011. Separately from the user who touches the touch operation input sensor of the display unit 10011, the operation input unit 1107 may be used, for example, by an administrator to operate the artificial intelligence response output device 10010. Through the user's operation input via the operation input unit 1107, the artificial intelligence response output device 10010 can obtain user input that forms the basis of instructions (prompts) to the large-scale language model, which is artificial intelligence. Note that there may also be a modified configuration in which the touch operation input sensor of the display unit 10011 is included as part of the operation input unit 1107.
[0034] The video signal input unit 1131 receives video data by connecting an external video output device. Various digital video input interfaces are possible for the video signal input unit 1131. For example, it may be configured as a video input interface conforming to the HDMI (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 1131 may also be a USB interface or the like.
[0035] The audio signal input unit 1133 receives audio data by connecting an external audio output device. The audio signal input unit 1133 may be configured as an HDMI audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface. The audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI interface, the video signal input unit 1131 and the audio signal input unit 1133 may be configured as an interface with integrated terminals and cables.
[0036] The audio output unit 1140 is capable of outputting audio based on audio data input to the audio signal input unit 1133. The audio output unit 1140 is also capable of outputting audio based on audio data stored in the storage unit 1170. The audio output unit 1140 may be configured as a speaker. The audio output unit 1140 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 1140 may be configured to output audio signals as digital signals to external devices, such as the Audio Return Channel function defined in the HDMI standard. Alternatively, the audio output unit 1140 may be configured to output audio signals as analog signals to external devices such as headphones.
[0037] Microphone 1139 is a microphone that picks up sounds from the vicinity of the artificial intelligence response output device 10010, converts them into signals, and generates an audio signal. The microphone may be configured to record a person's voice, such as the user's voice, and the control unit 1110, described later, may perform speech recognition processing on the generated audio signal to obtain textual information from the audio signal. Through the audio input from microphone 1139, the artificial intelligence response output device 10010 can obtain user input that forms the basis of instructions (prompts) for the large-scale language model, which is an artificial intelligence.
[0038] The imaging unit 1180 is a camera having an image sensor. The camera may be provided on the front of the display unit 10011 side of the artificial intelligence response output device 10010, or on the back of the display unit 10011 side. Both a front camera and a rear camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a rear camera.
[0039] The storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 1170 may be composed of a magnetic recording medium such as a hard disk drive (HDD) or a semiconductor memory such as a solid-state drive (SSD). For example, the storage unit 1170 may have various types of information, such as video data, image data, and audio data, pre-recorded in it at the time of product shipment. The storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from external devices or external servers via the communication unit 1132. The video data, image data, etc., recorded in the storage unit 1170 are output to the display unit 10011. The video data, image data, etc., recorded in the storage unit 1170 may also be output to external devices or external servers via the communication unit 1132.
[0040] The video control unit 1160 performs various controls related to the video signal input to the display unit 10011. The video control unit 1160 may also be called a video processing circuit and may be composed of hardware such as an ASIC, FPGA, or video processor. The video control unit 1160 may also be called a video processing unit or image processing unit. For example, the video control unit 1160 performs control such as switching between video signals, such as which video signal to input to the display unit 10011 from among the video signals to be stored in the memory 1109 and the video signals (video data) input to the video signal input unit 1131. The video control unit 1160 may also perform control to perform image processing on the video signals input from the video signal input unit 1131 and the video signals to be stored in the memory 1109. Examples of image processing include scaling processing to enlarge, reduce, and transform images, brightness adjustment processing to change brightness, contrast adjustment processing to change the contrast curve of an image, and retinex processing to decompose an image into its light components and change the weighting of each component.
[0041] The attitude sensor 1113 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the attitude of the artificial intelligence response output device 10010. Based on the attitude detection result of the attitude sensor 1113, the control unit 1110 may control the operation of each connected part.
[0042] The position sensor 1114 is a sensor that measures the position of the artificial intelligence response output device 10010 using a GNSS (Global Navigation Satellite System) such as GPS (Global Positioning System). If the position of the artificial intelligence response output device 10010 can be estimated based on information obtained through communication by the communication unit 1132, the position estimated by that method may be used as a substitute for the measurement result of the position sensor 1114. The position sensor 1114 may also receive UTC (coordinated universal time) time information transmitted by the GNSS.
[0043] The non-volatile memory 1108 stores various data used by the artificial intelligence response output device 10010. The data stored in the non-volatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the artificial intelligence response output device 10010, display icons, data and layout information for objects that the user can operate. The memory 1109 stores video data to be displayed on the display unit 10011 and control data for the device. The control unit 1110 may read various software from the storage unit 1170, expand it into the memory 1109, and store it.
[0044] The local LLM processing unit 10028 has a memory that can hold a large language model (LLM) and can execute inferences of the large language model based on the control of the control unit 1110. As hardware, it may be composed of, for example, a so-called GPU (Graphics Processing Unit). As another example of the hardware of the local LLM processing unit 10028, it may be composed of a so-called NPU (Neural network Processing Unit). The local LLM processing unit 10028 may not only perform inferences but also perform learning. Note that the local LLM processing unit 10028 is not necessarily required when the execution of inferences of the large language model in the local environment of the artificial intelligence response output device 10010 is unnecessary, etc.
[0045] The control unit 1110 controls the operations of each connected unit. Also, the control unit 1110 may perform arithmetic processing based on the information acquired from each unit within the artificial intelligence response output device 10010 in cooperation with the program stored in the memory 1109. The control states by the control unit 1110 include, for example, a state of outputting the response from the large language model of the local LLM processing unit 10028 or the response from the large language model of the large language model server 19001 or the multimodal large language model of the multimodal large language model server 20001 acquired via the communication unit 1132 via the display unit 10011 or the audio output unit 1140 such as a speaker. The specific configuration of the control unit 1110 is a CPU or the like, and may also be referred to as a processor or a control circuit.
[0046] Clock 1196 is a clock that counts time or time, which can be used in the control of the control unit 1110 of the artificial intelligence response output device 10010. Clock 1196 may function as an internal clock, rotating the value of a counter. Alternatively, clock 1196 may be a clock that counts regional times such as UTC (Coordinated Universal Time), which is a time used worldwide, or JST (Japan Standard Time). Clock 1196 may operate autonomously as an internal clock, or it may calibrate its time using external information at predetermined timings. For example, if clock 1196 is a clock that counts UTC time, it may acquire time information such as NTP information for UTC from an NTP (Network Time Protocol) server or external device via the communication unit 1132, and calibrate the time using the NTP information acquired by clock 1196. Alternatively, the position sensor 1114 may acquire UTC time information received from GNSS, and calibrate the time using the UTC time information acquired by clock 1196. For example, if clock 1196 is a clock that counts regional time such as JST, it may acquire regional time information such as TOT (Time Offset Table) from a server or external device via the communication unit 1132, and calibrate the time using the regional time information acquired by clock 1196. Alternatively, the position sensor 1114 may acquire location information including the latitude and longitude of the artificial intelligence response output device 10010, which is received from the GNSS, along with UTC time information. The clock 1196 may then use the acquired UTC time information and location information to calculate time information for the region where the artificial intelligence response output device 10010 is located, and use this information to calibrate the time.
[0047] In addition, when there is an input from the user via the above touch panel, microphone 1139, or operation input unit 1107, an instruction sentence is generated based on the input, and it is transmitted to the local large language model of the local LLM processing unit 10028 included in the artificial intelligence response output device 10010, the large language model included in the large language model server 19001, or the multimodal large language model included in the large language model server 20001, and the control for obtaining a response from these large language models may be performed by the control unit 1110 in any case.
[0048] In the examples of FIGS. 1A and 1B, an example in which the artificial intelligence response output device 10010 includes the display unit 10011 has been described. However, the artificial intelligence response output device 10010 according to the embodiment of the present invention does not necessarily have to include the display unit 10011. For example, even if it does not include the display unit 10011, it may be configured to receive an input from the user for the artificial intelligence via the voice signal input unit 1133 or the microphone 1139, and output a response from an artificial intelligence such as a large language model to the input from the user via the voice output unit 1140.
[0049] According to the artificial intelligence response output device and the artificial intelligence response output system according to Embodiment 1 of the present invention described above, it is possible to receive an input from the user for an artificial intelligence such as a large language model, and output a response to the input from the user generated by inference of an artificial intelligence such as the large language model possessed by the server device on the network or the local large language model possessed by the artificial intelligence response output device.
[0050] The artificial intelligence response output device 10010 according to the embodiment of the present invention can be implemented as a device in various forms. An example is shown in Figure 1C. For example, the artificial intelligence response output device 10010 may be a display device 181 such as a television, monitor, or display device as shown in Figure 1C(1). The display device 181 displays images on a display screen such as a display panel. The artificial intelligence response output device 10010 may be an information processing terminal 182 such as a smartphone, tablet terminal, or PC (personal computer) as shown in Figure 1C(2). The information processing terminal 182 displays images on a display screen such as a display panel. The artificial intelligence response output device 10010 may be a head-mounted display (HMD) 183 as shown in Figure 1C(3). In this case, the artificial intelligence response output device 10010 includes an eyepiece optical system that projects the image displayed by the display unit 10011 onto the user's eyes and generates a virtual image of the image. The artificial intelligence response output device 10010 may be a head-up display (HUD) 184, as shown in Figure 1C(4), which generates a virtual image 192 by reflecting projected light onto a transparent member 191 such as glass, and displays an image on the virtual image. In this case, the artificial intelligence response output device 10010 includes an optical system for projecting light emitted from the display unit 10011 onto the transparent member 191. The artificial intelligence response output device 10010 may also be an aerial floating image display device 185, as shown in Figure 1C(5), which generates an optical image 195, which is a real image, in the air by reflecting light emitted from a display device 193 with an imaging optical plate 194 such as a retroreflective member or a corner reflector array, and displays an image on the optical image. The artificial intelligence response output device 10010 may be a projection-type image display device 186, such as a projector that projects a projected image 196 onto a screen or wall, as shown in Figure 1C(6). In this case, the artificial intelligence response output device 10010 comprises a display panel which is a display unit 10011, a light source, an illumination optical system that guides light from the light source to the display unit 10011, and a projection optical system that projects the light transmitted or reflected by the display panel onto a screen or wall.
[0051] Next, an example of the configuration of the large-scale language model server 20001 according to an embodiment of the present invention will be described with reference to Figure 1D. Note that although the configuration in this figure is described as an example of the configuration of the large-scale language model server 20001, the large-scale language model server 19001 and the second server 19002 may have similar configurations.
[0052] The large-scale language model server 20001 may include, for example, a user interface 22010, an input interface 22011, an output interface 22012, an external power input interface 22014, a power supply 22015, a storage unit 22020, a learning data database (DB) 22021, various databases (DB) 22022, a control unit 22030, a memory 22031, a non-volatile memory 22032, a communication unit 22035, an AI calculation unit 22040, an LLM processing unit 22041, various AI processing units 22042, a cooling unit 22050, and so on.
[0053] Specifically, the user interface 22010 is a user interface for a server user, such as an administrator or maintenance worker of the large-scale language model server 20001, to operate the server. The user interface 22010 includes, for example, an input interface 22011. The input interface 22011 is an interface for inputting operation inputs and control information for the user to operate or control the server. For example, the input interface 22011 may be a physical operation button, an operation input panel such as a touch panel, or a remote controller. Alternatively, or in addition to the above, an interface may be provided that connects to an information terminal owned by the user and inputs control information from that information terminal. Specifically, a communication interface such as a USB interface using a Universal Serial Bus or an Ethernet interface may be provided. The user interface 22010 also includes an output interface 22012. The output interface 22012 is an interface for outputting information from the server to a user, such as an administrator. For example, the information to be output may include information on the operating status such as the server load, whether or not there are errors, and the type of error if there are errors. Alternatively, the information to be output may also include information related to the AI calculation unit 22040. The information related to the AI calculation unit 22040 may include information on the type of large-scale language model performing inference on the LLM processing unit 22041, or information on the execution status such as the execution progress rate of the inference. For example, the output interface 22012 may be an interface that connects to an information terminal owned by the user and outputs information to that information terminal. Specifically, it may be a communication interface such as a USB interface or an Ethernet interface. Alternatively, or in addition to the above, the output interface 22012 may be provided with a display unit. Information can be transmitted to the user via a displayed image or video. Note that if the user interface 22010 is a communication interface, it may share hardware with the communication unit 22035 described later.
[0054] The power supply 22015 converts the AC current input from the outside via the external power input interface 22014 into DC current and supplies the necessary DC current to each part of the large-scale language model server 20001.
[0055] The storage unit 22020 is a storage device that records various types of information, such as image data, video data, audio data, and text data. This information may be stored in a database (DB) structure. It may also be composed of magnetic recording media such as hard disk drives (HDDs) or semiconductor memory such as solid-state drives (SSDs). In the example shown in this figure, the storage unit 22020 is shown as being located within the large-scale language model server 20001. However, the storage unit 22020 may also be located outside the large-scale language model server 20001 and connected to the large-scale language model server 20001 via a communication interface. In the example shown in this figure, the storage unit 22020 includes various DBs 22021. The information stored in the various DBs 22021 may be output to, for example, the AI calculation unit 22040, which will be described later. In this case, the AI calculation unit 22040 can use the information stored in the various DBs 22021 for inference. Furthermore, the storage unit 22020 may record data such as image data, video data, audio data, and text data, which are the results of the inference performed by the AI calculation unit 22040, in various DBs 22021. The information stored in the various DBs 22021 may also be output as a response to access from external devices connected via the communication unit 22035 and the Internet 19000, as described later. Additionally, the storage unit 22020 may record data such as image data, video data, audio data, and text data transmitted from external devices connected via the communication unit 22035 and the Internet 19000 in various DBs 22021.
[0056] Furthermore, when training an AI model, such as a large-scale language model, to be used in the AI calculation unit 22040 described later, the storage unit 22020 may store a training data DB 22022, which is a database of training data, such as image data, video data, audio data, and text data, that will be the target data for training.
[0057] The control unit 22030 controls the operation of each part within the large-scale language model server 20001. The control unit 22030 may also work in cooperation with a program stored in memory 22031 to perform calculations based on information obtained from each part within the large-scale language model server 20001. Control states by the control unit 22030 include receiving various control information and instructions from external devices via the communication unit 22035 and the internet 19000, and using this control information and instructions to cause the AI calculation unit 22040 to perform inference of the large-scale language model. Furthermore, control states by the control unit 22030 include, for example, outputting the output from the large-scale language model of the AI calculation unit 22040 (described later) to external devices via the communication unit 22035 and the internet 19000. The specific configuration of the control unit 22030 is a CPU, and may also be referred to as a processor or control circuit.
[0058] The non-volatile memory 22032 stores various data used by the large-scale language model server 20001. The data stored in the non-volatile memory 22032 includes, for example, various data output from the output interface 22012 of the large-scale language model server 20001.
[0059] Memory 22031 stores control data and the like used by the control unit 22030. The control unit 22030 may read various software / applications from the storage unit 22020, load them into memory 22031, and store them there.
[0060] The communication unit 22035 is a communication interface such as Ethernet and connects to the internet 19000 via a router, switch, gateway device, etc. The communication path from the communication unit 22035 to the internet 19000 may include both wired and wireless sections, and may also pass through repeaters. If the communication unit 22035 is an Ethernet interface, communication may be performed using a LAN communication method up to a predetermined router, switch, or gateway device. This allows the communication unit 22035 to communicate with the communication unit 1132 of the artificial intelligence response output device 10010 and other devices such as servers via the internet 19000. The communication unit 22035 communicates with the communication unit 1132 of the artificial intelligence response output device 10010 and can send and receive instructions, control information, and artificial intelligence responses between the AI calculation unit 22040 (described later) and the artificial intelligence response output device 10010 using an API.
[0061] Next, the AI computation unit 22040 is a computation unit that performs inference on AI models such as neural networks. Specifically, it is a processor or circuit such as a GPU (Graphics Processing Unit). It may also be configured to have multiple cores, with the GPU and memory for the GPU as cores. Another example of the hardware of the AI computation unit 22040 is that it may be configured as a so-called NPU (Neural Network Processing Unit). The AI model on which the AI computation unit 22040 performs inference is, for example, a neural network of various deep learning types. Specifically, it may be an LLM (Large Language Model) model, or an image recognition processing model, or an audio recognition processing model. The model can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Transformer, an Autoencoder, a Generative Adversarial Network (GAN), or a Diffusion Model.
[0062] In the example shown in Figure 1D, the AI calculation unit 22040 includes an AI model of an LLM and an LLM processing unit 22041 that performs inference on the LLM. For example, the LLM on which the LLM processing unit 22041 performs inference is an AI model that has learned at least text information. Alternatively, the LLM on which the LLM processing unit 22041 performs inference may be a multimodal LLM that has learned not only text data but also media other than text data. Such media other than text data may be, for example, image data, video data, or audio data. The LLM processing unit 22041 obtains instruction statements and control information from external devices such as an artificial intelligence response output device 10010 connected to the Internet 19000 via communication through a communication unit 22035 using an API, and performs inference on the LLM based on this information. The result of the inference is output to the external device as an artificial intelligence response via communication through the communication unit 22035 using an API. These processes can be controlled, for example, by the control unit 22030.
[0063] Furthermore, the AI calculation unit 22040 may include various AI models other than LLM, and may also include various AI processing units 22042 that perform inference on these various AI models. Examples of various AI models other than LLM include clustering AI models that have undergone unsupervised learning, regression prediction AI models that have undergone supervised learning, and classification AI models that have undergone supervised learning. The various AI processing units 22042 acquire control information from external devices such as an artificial intelligence response output device 10010 connected to the Internet 19000 via communication through a communication unit 22035 using an API, and perform inference on the various AI models based on this information. The results of the inference are output to the external devices as an artificial intelligence response via communication through the communication unit 22035 using an API. These series of processes can be controlled, for example, by the control unit 22030.
[0064] Furthermore, when training AI models such as large-scale language models on the large-scale language model server 20001, the AI calculation unit 22040 can use image data, video data, audio data, text data, etc., stored in the training data DB 22022 of the storage unit 22020 as training data to train the AI model.
[0065] The cooling unit 22050 is primarily configured to cool the AI calculation unit 22040. Learning and inference of AI models require a lot of power and generate heat. Therefore, a configuration is needed to forcibly cool the AI calculation unit 22040 in addition to natural convection air cooling. Specifically, the cooling unit 22050 may be a forced convection device for air cooling such as a fan, or a forced convection device for liquid cooling such as a pump. Alternatively, the cooling unit 22050 may be a heat exchanger such as a heat pipe.
[0066] As described above using Figure 1D, the server configuration allows the AI calculation unit 22040 to perform inference on an AI model based on control information received via communication through the communication unit 22035, and the result to be output as an artificial intelligence response via communication through the communication unit 22035.
[0067] Note that the server in Figure 1D is described as a single physical server as an example. However, this server may also be a virtualized server using multiple physical servers. This server may also be configured as a cloud server where various functions are distributed and virtualized on the cloud (cloud computing system).
[0068] Furthermore, the control unit 22030 may monitor the operating status of the power supply 22015, the load of the AI calculation unit 22040, and the load of the forced convection device such as the fan or pump of the cooling unit 22050. This operating status information monitored by the control unit 22030 may be output through an output interface 22012, such as the display unit of the user interface 22010. This operating status information monitored by the control unit 22030 may also be output to a network such as the Internet 19000 via the communication unit 22035, and the outputted operating status information may be configured to be obtainable by an external device such as an artificial intelligence response output device 10010 connected to the network.
[0069] Next, using Figure 1E, we will describe an example of deploying and executing a client application and an LLM application in an artificial intelligence response output system. Figure 1E is a schematic diagram that focuses on the deployment, execution, and communication of the client application and the LLM application in an artificial intelligence response output system configuration that includes an artificial intelligence response output device 10010 and a large-scale language model server 20001. Other components besides the client application and the LLM application are omitted in notation and description for the sake of simplicity.
[0070] In the example shown in Figure 1E, the client application 8010 is executed in the artificial intelligence response output device 10010. Specifically, the client application 8010 is loaded into the memory 1109 shown in Figure 1B and executed by the control unit 1110. In addition, the server LLM application 8020 is executed in the large-scale language model server 20001. Specifically, the server LLM application 8020 is loaded into the memory 22031 shown in Figure 1D and executed by the control unit 22030. The server LLM application 8020 may also be an application that controls the entire large-scale language model possessed by the large-scale language model server 20001. For example, the server LLM application 8020 controls the execution of inference of the AI model of the LLM provided in the LLM processing unit 22041 of the AI calculation unit 22040 shown in Figure 1D. The server LLM application 8020 may also be an application that controls the input and output of information to and from the large-scale language model possessed by the large-scale language model server 20001. The client application 8010 and the server LLM application 8020 communicate via, for example, the communication unit 1132 shown in Figure 1B, the communication device 19011 shown in Figure 1A, a network such as the Internet 19000, and the communication unit 22035 shown in Figure 1D. For example, the client application 8010 sends an instruction to the server LLM application 8020. Upon receiving the instruction, the server LLM application 8020 performs control to execute inference of the LLM AI model provided by the LLM processing unit 22041 of the AI calculation unit 22040 shown in Figure 1D. The server LLM application 8020 then sends the artificial intelligence response, which is the result of this inference, to the client application 8010. The client application 8010 can output to the user based on this artificial intelligence response. In this way, in the artificial intelligence response output system of the example in Figure 1E, the client application 8010 and the server LLM application 8020 cooperate to output a more suitable artificial intelligence response to the user. Here, the large-scale language model possessed by the large-scale language model server 20001 may also be referred to as the server large-scale language model.The server LLM application 8020 may also be called a server-wide large-scale language model application.
[0071] In the example shown in Figure 1E, the local LLM application 8015 is executed in the artificial intelligence response output device 10010. Specifically, the local LLM application 8015 is loaded into the memory 1109 shown in Figure 1B in the artificial intelligence response output device 10010 and executed by the control unit 1110. The local LLM application 8015 may also be an application that controls the entire large-scale language model of the local LLM processing unit 10028 shown in Figure 1B. For example, the local LLM application 8015 controls the execution of inference for the large-scale language model of the local LLM processing unit 10028 shown in Figure 1B. Alternatively, the local LLM application 8015 may also be an application that controls the input and output of information to and from the large-scale language model of the local LLM processing unit 10028. The client application 8010 and the local LLM application 8015 communicate, for example, via a communication path such as a bus within the artificial intelligence response output device 10010. For example, the client application 8010 sends an instruction to the local LLM application 8015. Upon receiving the instruction, the local LLM application 8015 controls the local LLM processing unit 10028, shown in Figure 1B, to perform inference on its large-scale language model. The local LLM application 8015 then sends the artificial intelligence response, which is the result of this inference, to the client application 8010. The client application 8010 can then output an output based on this artificial intelligence response to the user. In this way, in the artificial intelligence response output system shown in the example in Figure 1E, the client application 8010 and the local LLM application 8015 work together to output a more suitable artificial intelligence response to the user. Here, the large-scale language model of the local LLM processing unit 10028 may also be called a local large-scale language model. The local LLM application 8015 may also be called a local large-scale language model application.
[0072] Furthermore, in the artificial intelligence response output system shown in the example of Figure 1E, the client application 8010, the local LLM application 8015, and the server LLM application 8020 work together to output a more suitable artificial intelligence response to the user.
[0073] In Figures 1A, 1D, and 1E, an example was shown in which the artificial intelligence response output device 10010 and the large-scale language model server 20001 are connected via the Internet 19000. That is, in Figure 1E, the server LLM application 8020 was connected to the client application 8010 via the Internet 19000. However, the network connecting the artificial intelligence response output device 10010 and the large-scale language model server 20001 is not limited to the Internet 19000, but may also be a so-called LAN (Local Area Network). The network connecting the artificial intelligence response output device 10010 and the large-scale language model server 20001 may also be a so-called local 5G network. In these cases, the artificial intelligence response output device 10010, which has the client application 8010 and the local LLM application 8015, may be on the same subnet mask as the server LLM application 8020 and the large-scale language model server 20001. Thus, an artificial intelligence response system, including an artificial intelligence response output device 10010 and a large-scale language model server 20001, built without using the internet 19000, can be used in local settings such as homes, factories, and hospitals. In this case, even if communication with the internet 19000 is interrupted for any reason, the client application 8010 of the artificial intelligence response output device 10010 can communicate with the server LLM application 8020 of the large-scale language model server 20001 via a LAN or local 5G network, enabling collaboration.
[0074] Here, a more specific configuration example of the artificial intelligence response output device 10010 according to an embodiment of the present invention being an aerial floating image display device will be described using Figures 1F to 1K.
[0075] The following examples in Figures 1F to 1K relate to an image display device capable of transmitting an image generated by image light from an image light source through a transparent component that partitions a space, such as glass, and displaying it as an aerial floating image outside the transparent component. In the description of Figures 1F to 1K, the image floating in the air is referred to as an "aerial floating image." Instead of this term, it is also acceptable to use terms such as "aerial image," "spatial image," "spatial floating image," "aerial floating optical image of a displayed image," or "spatial floating optical image of a displayed image." The term "aerial floating image," which is mainly used in the description of the embodiment, is used as a representative example of these terms.
[0076] <Example Configuration 1 of an Aerial Floating Image Display Device> First, using Figure 1F, we will explain an example of an optical system used in the artificial intelligence response output device 10010, which is an aerial floating image display device.
[0077] In the optical system shown in Figure 1F, the display device 1 comprises a liquid crystal display panel 11 and a light source device 13. The display device 1 outputs image light of a specific polarization. The surface of the display device 1 may be equipped with an absorbing polarizing plate 12 that transmits the image light of the specific polarization and absorbs the other polarization. By providing the absorbing polarizing plate 12, unwanted reflected light can be reduced, thereby reducing stray light and other unwanted light. The image light of the specific polarization output from the display device 1 is input to the polarization separation member 101B. The polarization separation member 101B is a member that selectively transmits the image light of the specific polarization. The polarization separation member 101B is not integrated with the transparent member 100, but has an independent plate-like shape. Therefore, the polarization separation member 101B may also be described as a polarization separation plate. The polarization separation member 101B may be configured, for example, as a reflective polarizing plate formed by attaching a polarization separation sheet to a transparent member. Alternatively, the transparent member may be formed with a metal multilayer film that selectively transmits specific polarizations and reflects polarizations of other specific polarizations. In Figure 1F, the polarization separation member 101B is configured to transmit image light of a specific polarization output from the display device 1.
[0078] The image light that has passed through the polarization separation member 101B is incident on the retroreflector 2. A λ / 4 plate 21 is provided on the image light incident surface of the retroreflector. The image light is polarized from one polarization to the other by passing through the λ / 4 plate 21 twice, once when it is incident on the retroreflector and once when it is emitted. Here, the polarization separation member 101B has the property of reflecting the polarization of the other polarization that has been polarized by the λ / 4 plate 21, so the image light after polarization conversion is reflected by the polarization separation member 101B. The image light reflected by the polarization separation member 101B passes through the transparent member 100, forming a real image of a floating image 3 on the outside of the transparent member 100.
[0079] Here, we will explain a first example of polarization design in the optical system shown in Figure 1F. For example, the display device 1 may be configured to emit P-polarized (P stands for parallel; polarization in which the electric field oscillates within the incident plane) image light to the polarization separation member 101B, and the polarization separation member 101B may be configured to reflect S-polarized (S stands for senkrecht; polarization in which the electric field oscillates perpendicular to the incident plane) and transmit P-polarized light. In this case, the P-polarized image light that reaches the polarization separation member 101B from the display device 1 passes through the polarization separation member 101B and heads towards the retroreflector 2. When the image light is reflected by the retroreflector 2, it passes through the λ / 4 plate 21 provided on the incident surface of the retroreflector 2 twice, so the image light is converted from P-polarized to S-polarized. The image light converted to S-polarized then heads towards the polarization separation member 101B again. Here, since the polarization separation member 101B has the characteristic of reflecting S-polarized light and transmitting P-polarized light, the S-polarized image light is reflected by the polarization separation member 101 and transmitted through the transparent member 100. Since the image light transmitted through the transparent member 100 is light generated by the retroreflector 2, it forms an aerial levitation image 3, which is the optical image of the display image of the display device 1, at a position that is mirror-image to the display image of the display device 1 with respect to the polarization separation member 101B. With such a polarization design, the aerial levitation image 3 can be suitably formed.
[0080] Next, a second example of polarization design in the optical system shown in Figure 1F will be described. For example, the display device 1 may be configured to emit S-polarized image light to the polarization separation member 101B, and the polarization separation member 101B may be configured to reflect P-polarized light and transmit S-polarized light. In this case, the S-polarized image light that reaches the polarization separation member 101B from the display device 1 passes through the polarization separation member 101B and heads towards the retroreflector 2. When the image light is reflected by the retroreflector 2, it passes through the λ / 4 plate 21 provided on the incident surface of the retroreflector 2 twice, so the image light is converted from S-polarized to P-polarized light. The image light converted to P-polarized light heads towards the polarization separation member 101B again. Here, since the polarization separation member 101B has the characteristic of reflecting P-polarized light and transmitting S-polarized light, the P-polarized image light is reflected by the polarization separation member 101 and transmitted through the transparent member 100. Since the image light transmitted through the transparent member 100 is light generated by the retroreflector 2, it forms a floating image 3, which is an optical image of the display image of the display device 1, at a position that is mirror-like to the display image of the display device 1 with respect to the polarization separation member 101B. With this polarization design, the floating image 3 can be suitably formed.
[0081] In Figure 1F, the image display surface of the display device 1 and the surface of the retroreflector 2 are arranged parallel to each other. The polarization separation member 101B is positioned at an angle α (for example, 45°) relative to the image display surface of the display device 1 and the surface of the retroreflector 2. As a result, in the reflection by the polarization separation member 101B, the direction of propagation of the image light reflected by the polarization separation member 101B (the direction of the principal ray of the image light) differs from the direction of propagation of the image light incident from the retroreflector 2 (the direction of the principal ray of the image light) by an angle β (for example, 90°). With this configuration, in the optical system of Figure 1F, image light is output at a predetermined angle shown toward the outside of the transparent member 100, forming the floating image 3, which is a real image. In the configuration of Figure 1F, when a user views from the direction of arrow A, the floating image 3 is visible as a bright image. However, when another person views from the direction of arrow B, the floating image 3 cannot be seen as an image at all. This characteristic makes it ideal for systems that display video requiring high security, or highly confidential video that should be kept hidden from the user.
[0082] Next, Figure 1F(2) shows an example of the surface shape of a typical retroreflector 2. The retroreflector 2 has a prism body in which regularly arranged triangular pyramidal recesses serve as reflective surfaces. Light rays incident on the arranged triangular pyramidal recesses are reflected by multiple reflective surfaces of the triangular pyramidal recesses and emitted as retroreflected light in the direction corresponding to the incident light, and the display device 1 displays a real image of a floating object in the air based on the image displayed on the display device 1.
[0083] The surface shape of the retroreflector in this embodiment is not limited to the examples described above. It may have various surface shapes that realize retroreflection. Specifically, retroreflective elements formed by periodically arranging triangular pyramidal prisms, hexagonal pyramidal prisms, other polygonal prisms, multi-vertex prisms, or combinations thereof may be provided on the surface of the retroreflector in this embodiment. Alternatively, retroreflective elements forming cube corners by periodically arranging these prisms may be provided on the surface of the retroreflector in this embodiment. These can also be expressed as corner reflector arrays or polyhedron reflector arrays. Alternatively, capsule lens type retroreflective elements formed by periodically arranging glass beads may be provided on the surface of the retroreflector in this embodiment. Since the detailed configuration of these retroreflective elements can be described using existing technology, a detailed explanation is omitted. Specifically, the technology disclosed in Japanese Patent Publication No. 2001-33609, Japanese Patent Publication No. 2001-264525, Japanese Patent Publication No. 2005-181555, Japanese Patent Publication No. 2008-70898, Japanese Patent Publication No. 2009-229942, etc., can be used.
[0084] As explained above, the optical system in Figure 1F can form a suitable aerial levitation image.
[0085] Next, Figure 1G shows an example of the configuration of an aerial levitation image display device 1000, which is one embodiment of the artificial intelligence response output device 10010. The aerial levitation image display device 1000 shown in Figure 1G is equipped with an optical system corresponding to the optical system in Figure 1F. The aerial levitation image display device 1000 shown in Figure 1G is installed vertically, for example, so that the side on which the aerial levitation image 3 is formed faces the front of the aerial levitation image display device 1000 (towards the user 230). That is, in Figure 1G, the aerial levitation image display device 1000 has a transparent member 100 installed on the front of the device (towards the user 230). The aerial levitation image 3 is formed on the user 230 side relative to the surface of the transparent member 100 of the aerial levitation image display device 1000. The light of the aerial levitation image 3 travels in the direction toward the user (the -y direction in the figure). If the aerial operation detection sensor 1351, which will be described later, is provided as shown, it is possible to detect operation of the aerial levitation image 3 by the user 230's finger.
[0086] <Example 2 of the Configuration of the Aerial Floating Image Display Device> Another example of the configuration of the optical system of the aerial floating image display device will be explained using Figure 1H. The optical system in Figure 1H is an optical system that uses a retroreflector 5, which is different from the retroreflector 2 used in Figure 1F. Components in Figure 1H that are given the same reference numerals as in Figure 1F have the same function and configuration as those in Figure 1F. Such components will not be explained again in order to simplify the explanation.
[0087] Figure 1H shows an example of the main components and retroreflective components of an aerial floating image display device according to one embodiment of the present invention. A display device 1 that emits image light is provided obliquely to a transparent member 100 such as glass. The display device 1 comprises a liquid crystal display panel 11 and a light source device 13 that generates light.
[0088] The principal ray 9020, which represents the light beam emitted from the display device 1, travels toward the retroreflector 5 and is incident on the retroreflector 5 at an incident angle α (here defined as the angle with respect to the plane direction of the retroreflector 5). The incident angle α can be, for example, 45°. However, the incident angle α is not limited to 45°; for example, 45° ± 15° can also be used.
[0089] The retroreflector 5 is an optical component having optical properties that retroreflect light rays in at least some directions. Furthermore, since the reflected light rays have optical properties that form an image, the retroreflector 5 may also be described as an imaging optical component or imaging optical plate.
[0090] The specific configuration of the retroreflector 5 will be described in detail using Figure 1I, but the retroreflector 5 causes the principal ray 9020 to propagate in the z direction while being retroreflected in the x and y directions. As a result, the reflected ray 9021 travels in a direction away from the retroreflector 5, following an optical path that is mirror-symmetric with respect to the principal ray 9020 with respect to the retroreflector 5, passing through the transparent member 100, and forming a floating image 3 as a real image at the imaging plane.
[0091] The light beam forming the floating image 3 is a collection of light rays converging from the retroreflector 5 to the optical image of the floating image 3, and these light rays continue to travel in a straight line even after passing through the optical image of the floating image 3. Therefore, unlike the diffused image formed on a screen by a typical projector, the floating image 3 is an image with high directivity. Thus, in the configuration of Figure 1H, when a user views from the direction of arrow A, the floating image 3 is visible as a bright image. However, when another person views from the direction of arrow B, the floating image 3 cannot be seen as an image at all. This characteristic is suitable for use in systems that display images requiring high security or highly confidential images that should be hidden from people directly facing the user.
[0092] Next, an example of the configuration of the retroreflector 5 will be described using Figure 1I. The retroreflector 5 has a configuration in which multiple corner reflectors 9040 are arranged in an array on the surface of a transparent material. This may also be called a corner reflector array or a multifaceted reflector array. Light rays 9111, 9112, 9113, and 9114 emitted from the light source 9110 are reflected twice by the two mirror surfaces 9041 and 9042 of the corner reflectors 9040, becoming reflected light rays 9121, 9122, 9123, and 9124. This double reflection is retroreflection in the x and y directions, where the light is reflected back in the same direction as the incident direction (moving in a direction rotated 180°), and in the z direction, it is specular reflection in which the angle of incidence and the angle of reflection coincide due to total internal reflection.
[0093] In other words, the light rays 9111 to 9114 produce reflected light rays 9121 to 9124 on a straight line symmetrical in the z direction with respect to the corner reflector 9040, forming an aerial real image 9120. The light rays 9111 to 9114 emitted from the light source 9110 are four representative rays of diffused light from the light source 9110, and depending on the diffusion characteristics of the light source 9110, the light rays incident on the retroreflector 5 are not limited to these, but any incident light ray will cause similar reflection and form an aerial real image 9120. For the sake of clarity in the drawing, the position of the light source 9110 and the position of the aerial real image 9120 in the x direction are shown offset, but in reality, the position of the light source 9110 and the position of the aerial real image 9120 in the x direction are at the same position, and when viewed from the z direction, they are in overlapping positions.
[0094] In the optical system shown in Figure 1F, the retroreflector 2 has retroreflective properties in three axes. As a result, when a diffusive incident light beam is incident on the retroreflector 2, a convergent reflected light beam travels toward the side of the incident light beam where the light source is located relative to the retroreflector 2. This convergent reflected light beam forms an image in the air, creating a floating image 3. The direction of propagation of the principal ray of the convergent reflected light beam reflected from the retroreflector 2 is opposite to the direction of propagation of the principal ray of the diffusive incident light beam incident on the retroreflector 2.
[0095] In contrast, in the optical system shown in Figure 1H, the retroreflector 5 has retroreflective properties in two axes and specular reflection in the other axis. As a result, when a diffuse incident light beam is incident on the retroreflector 5, the convergent reflected light beam reflected by the corner reflector array travels toward the retroreflector 5 toward the side of the incident light beam away from the light source. This convergent reflected light beam forms an image in the air, creating a floating image 3.
[0096] The direction of propagation of the principal ray of the convergent reflected light beam reflected by the corner reflector array of the retroreflector 5 is not in the opposite direction to the direction of propagation of the principal ray of the diffuse incident light beam incident on the retroreflector 5. The component of the direction of propagation of the principal ray of the diffuse incident light beam incident on the retroreflector 5 in the direction of the plate-shaped surface of the retroreflector 5, and the component of the direction of propagation of the principal ray after it has been reflected by the retroreflector 5 and become a convergent reflected light beam, remain in a straight line before and after reflection by the corner reflector array.
[0097] In other words, the diffusive incident light beam is converted into a convergent reflected light beam by reflection at the retroreflector 5, but in the direction normal to the plate-shaped surface of the retroreflector 5, the light beam will travel through the retroreflector 5. Here, the diffusive incident light beam that enters the retroreflector 5 and the convergent reflected light beam that exits the retroreflector 5 are geometrically symmetrical with respect to the plate-shaped surface of the retroreflector 5.
[0098] The shape of the retroreflector (imaging optical plate) in the optical system shown in Figure 1H is not limited to the example described above. It may have various shapes to achieve retroreflection. Specifically, it may be various cubic corner bodies, corner reflector arrays, slit mirror arrays, two-sided corner reflector arrays, multi-sided reflector arrays, or a shape in which combinations of their reflective surfaces are arranged periodically. Alternatively, a capsule lens type retroreflector element with glass beads arranged periodically may be provided on the surface of the retroreflector in this embodiment. The detailed configuration of these retroreflector elements can be described using existing technology, so a detailed explanation is omitted. Specifically, the technology disclosed in Japanese Patent Publication No. 2017-33005, Japanese Patent Publication No. 2019-133110, Japanese Patent Publication No. 2017-67933, WO2009 / 131128, etc., can be used.
[0099] Next, Figure 1J shows an example of another configuration of the aerial floating image display device 1000, which is one embodiment of the artificial intelligence response output device 10010. The aerial floating image display device 1000 shown in Figure 1J is equipped with an optical system corresponding to the optical system in Figure 1H. In the aerial floating image display device 1000 shown in Figure 1J, it is installed horizontally so that the side on which the aerial floating image 3 is formed faces upward.
[0100] In other words, in Figure 1J, the aerial levitation image display device 1000 has a transparent member 100 installed on the upper surface of the device. The aerial levitation image 3 is formed above the surface of the transparent member 100 of the aerial levitation image display device 1000. The light of the aerial levitation image 3 travels diagonally upward. If the aerial operation detection sensor 1351, which will be described later, is provided as shown, it is possible to detect operation of the aerial levitation image 3 by the user 230's finger.
[0101] In Figure 1G, the display device 1 and the floating image 3 are symmetrical with respect to the plane of the polarization separation member 101. In contrast, in Figure 1J, the display device 1 and the floating image 3 are symmetrical with respect to the plane of the retroreflector 5. Also, the configuration in Figure 1G includes a retroreflector 2 and a λ / 4 plate 21, but these are not present in Figure 1J. Furthermore, in Figure 1G, the presence of an absorptive polarizer 12 is preferable, but in Figure 1J, the absorptive polarizer 12 is not particularly necessary.
[0102] <<Block Diagram of the Internal Configuration of the Aerial Floating Image Display Device>> Next, a block diagram of the internal configuration of the aerial floating image display device 1000, as described in Figure 1G or Figure 1J, will be explained. Figure 1K is a block diagram showing an example of the configuration of the aerial floating image display device 1000. In the embodiment of the present invention, the aerial floating image display device 1000 is one aspect of the artificial intelligence response output device 10010, so the configuration example of the aerial floating image display device 1000 in Figure 1K is a configuration example in which some configurations have been changed or added compared to the configuration example of the artificial intelligence response output device 10010 in Figure 1B. In Figure 1K, components that are denoted by the same reference numerals as in Figure 1B are the same as in Figure 1B, so repeated explanations will be omitted.
[0103] The aerial levitation image display device 1000 includes, in place of or in addition to, the display unit 10011 in Figure 1B, a retroreflective unit 1101, an image display unit 1102, a light guide 1104, and a light source 1105. The aerial levitation image display device 1000 includes, in place of or in addition to, the operation input unit 1107 in Figure 1B, an aerial operation detection sensor 1351 and an aerial operation detection unit 1350. The aerial levitation image display device 1000 also includes a power supply 1106, an external power input interface 1111, a control unit 1110, an imaging unit 1180, etc. Furthermore, it may also include a secondary battery 1112, etc. The various processing units 9999 have the configuration already described in Figure 1B, and a detailed, repeated explanation is omitted in this figure. Specifically, these include memory 1109, non-volatile memory 1108, storage unit 1170, video control unit 1160, attitude sensor 1113, position sensor 1114, local LLM processing unit 10028, communication unit 1132, audio output unit 1140, microphone 1139, video signal input unit 1131, audio signal input unit 1133, etc.
[0104] Each component of the aerial floating image display device 1000 is arranged in the housing 1190. Note that the imaging unit 1180 and the aerial operation detection sensor 1351 shown in Figure 1K may be provided on the outside of the housing 1190.
[0105] The retroreflective section 1101 in Figure 1K corresponds to the retroreflective plate 2 in Figure 1F. The retroreflective section 1101 retroreflectively reflects light modulated by the image display section 1102. Of the reflected light from the retroreflective section 1101, the light output to the outside of the floating image display device 1000 forms the floating image 3. When the optical system in Figure 1H is applied, the retroreflective section 1101 corresponds to the retroreflective plate 5 in Figure 1H.
[0106] The image display unit 1102 in Figure 1K corresponds to the liquid crystal display panel 11 in Figures 1F, 1G, 1H, and 1J. The light source 1105 in Figure 1K corresponds to the light source device 13 in Figures 1F, 1G, 1H, and 1J. The image display unit 1102, light guide 1104, and light source 1105 in Figure 1K correspond to the display device 1 in Figures 1F, 1G, 1H, and 1J.
[0107] The video display unit 1102 is a display unit that generates an image by modulating transmitted light based on a video signal input by the video control unit 1160, which will be described later. As the video display unit 1102 (the liquid crystal display panel 11 described above), for example, a transmissive liquid crystal panel is used, but it is not limited to this. Alternatively, as the video display unit 1102, for example, a reflective liquid crystal panel that modulates reflected light or a DMD (Digital Micromirror Device: registered trademark) panel may be used.
[0108] The light source 1105 generates light for the image display unit 1102 and is a solid-state light source such as an LED (Light Emitting Diode) or a laser light source. The power supply 1106 converts the AC current input from an external source via the external power input interface 1111 into DC current and supplies power to the light source 1105. The power supply 1106 also supplies the necessary DC current to each part of the levitating image display device 1000. The secondary battery 1112 stores the power supplied from the power supply 1106. The secondary battery 1112 also supplies power to the light source 1105 and other components that require power via the external power input interface 1111 when power is not supplied from an external source. In other words, if the levitating image display device 1000 is equipped with a secondary battery 1112, the user can use the levitating image display device 1000 even when power is not supplied from an external source.
[0109] The light guide 1104 guides the light generated by the light source 1105 and illuminates the image display unit 1102. The combination of the light guide 1104 and the light source 1105 can also be called the backlight of the image display unit 1102. The light guide 1104 may be made mainly of glass. The light guide 1104 may be made mainly of plastic. The light guide 1104 may be made using mirrors.
[0110] The aerial operation detection sensor 1351 is a sensor that detects operation of the floating aerial image 3 by an object such as a user's finger. The aerial operation detection sensor 1351 senses, for example, the area that overlaps with the entire display range of the floating aerial image 3. Alternatively, the aerial operation detection sensor 1351 may sense only the area that overlaps with at least a portion of the display range of the floating aerial image 3.
[0111] Specific examples of the aerial operation detection sensor 1351 include distance sensors using invisible light such as infrared, invisible light lasers, and ultrasonic waves. The aerial operation detection sensor 1351 may also be configured by combining multiple sensors to detect coordinates in a two-dimensional plane. Furthermore, the aerial operation detection sensor 1351 may be composed of a Time of Flight (ToF) LiDAR (Light Detection and Ranging) or an image sensor.
[0112] The aerial operation detection sensor 1351 only needs to be capable of sensing touch operations by a user's finger on an object displayed as an aerial floating image 3. Such sensing can be performed using existing technologies.
[0113] The aerial operation detection unit 1350 acquires a sensing signal from the aerial operation detection sensor 1351 and, based on the sensing signal, calculates whether or not the user's finger has made contact with an object in the aerial floating image 3, and the position where the user's finger and the object made contact (contact position). The aerial operation detection unit 1350 is composed of a circuit such as an FPGA (Field Programmable Gate Array). In addition, some functions of the aerial operation detection unit 1350 may be implemented in software, for example, by an aerial operation detection program executed in the control unit 1110 or the video control unit 1160. The aerial operation detection sensor 1351 and the aerial operation detection unit 1350 may be configured as an integrated unit. The aerial operation detection unit 1350 and the control unit 1110 or the video control unit 1160 may be configured as an integrated unit.
[0114] The aerial operation detection sensor 1351 and the aerial operation detection unit 1350 may be built into the aerial floating image display device 1000, or they may be provided separately from the aerial floating image display device 1000. When provided separately from the aerial floating image display device 1000, the aerial operation detection sensor 1351 and the aerial operation detection unit 1350 are configured to transmit information and signals to the aerial floating image display device 1000 via wired or wireless communication lines or video signal transmission lines. This makes it possible to construct a system in which the aerial floating image display device 1000 without an aerial operation detection function is used as the main unit, and only the aerial operation detection function can be added as an option.
[0115] Alternatively, the aerial operation detection sensor 1351 may be a separate unit, and the aerial operation detection unit 1350 may be built into the aerial floating image display device 1000. The configuration in which only the aerial operation detection sensor 1351 is a separate unit has advantages, for example, when it is desired to have more freedom in positioning the aerial operation detection sensor 1351 relative to the installation location of the aerial floating image display device 1000.
[0116] The imaging unit 1180 is, for example, a camera with an image sensor, and captures images of the space near the floating image 3 and / or the user's face, arms, fingers, etc. Multiple imaging units 1180 may be provided. For example, the imaging units 1180 may be provided as a stereo camera. By using multiple imaging units 1180, or by using an imaging unit with a depth sensor, the aerial operation detection unit 1350 can be assisted when detecting touch operations on the floating image 3 by the user 230. The imaging unit 1180 may be provided separately from the floating image display device 1000. If the imaging unit 1180 is provided separately from the floating image display device 1000, it should be configured to transmit imaging signals to the floating image display device 1000 via a wired or wireless communication connection path or the like.
[0117] For example, if the aerial operation detection sensor 1351 is configured as an object intrusion sensor that detects whether or not an object has entered a plane (intrusion detection plane) that includes the display surface (display range) of the aerial floating image 3, the aerial operation detection sensor 1351 may not be able to detect information such as how far away an object that has not entered the intrusion detection plane (for example, a user's finger) is from the intrusion detection plane, or how close an object is to the intrusion detection plane.
[0118] In such cases, the distance between the object and the intrusion detection plane (floating image 3) can be calculated by using information such as depth calculation information of the object based on images captured by multiple imaging units 1180 and depth information of the object from a depth sensor. This various information, including depth calculation information, depth information, and the distance between the object and the intrusion detection plane, is then used for various display controls of the floating image 3.
[0119] Alternatively, instead of using the aerial operation detection sensor 1351, the aerial operation detection unit 1350 may detect touch operations on the aerial floating video 3 by the user 230 based on the image captured by the imaging unit 1180. In this case, the imaging unit 1180 may be referred to as the aerial operation detection sensor.
[0120] Furthermore, the imaging unit 1180 may capture an image of the user operating the floating video 3, and the control unit 1110 or the like may perform user identification processing. In addition, to determine if there are other people standing around or behind the user operating the floating video 3 and whether they are peeking at the user's operation of the floating video 3, the imaging unit 1180 may capture an area that includes the user operating the floating video 3 and the area surrounding the user.
[0121] In addition to the configuration and operation described in Figure 1B, the control unit 1110 also controls the various parts that realize the aerial floating image display function and the aerial operation detection function, which are newly described using Figure 1K.
[0122] As described above, the various processing units 9999 have the configuration already explained in Figure 1B, so a repeated explanation will be omitted.
[0123] As explained above, the aerial floating image display device 1000 is equipped with various functions. However, the aerial floating image display device 1000 does not need to have all of these functions; any configuration is acceptable as long as it has the function of forming the aerial floating image 3.
[0124] As described above with reference to Figures 1F to 1K, the aerial floating image display device 1000 of this embodiment can be equipped with an aerial floating image display function. In other words, an aerial floating image display device equipped with an artificial intelligence response output function can be realized.
[0125] <Example 2> Next, as Example 2 of the present invention, we will describe an example in which the artificial intelligence response output device 10010 described in Example 1 is connected to the internet and operates by connecting to a server equipped with a large-scale language model artificial intelligence via the internet. In this example, we will explain the differences from Example 1, and will omit repeated explanations of configurations similar to those in these examples.
[0126] Using Figure 2A, an example of the connection state between the artificial intelligence response output device 10010 and the large-scale language model server 20001 of Embodiment 2 of the present invention will be described. The artificial intelligence response output device 10010 according to Embodiment 2 has the function of displaying a character or avatar, and may be called a character conversation device or an avatar display device. Furthermore, the system including the artificial intelligence response output device 10010 and the large-scale language model server 20001 according to Embodiment 2 may be called a character conversation system or an avatar conversation system. Here, the display unit 10011 displayed by the artificial intelligence response output device 10010 displays an image of character 19051. The image of character 19051 is an image generated by rendering a 3D model of the character in a virtual space. Note that repeatedly writing "character or avatar" is redundant, so in the following description of this embodiment, it will simply be written as "character".
[0127] Furthermore, the character in this embodiment can provide the user with the services of a large-scale language model, which is an artificial intelligence, and can assist the user. Therefore, the character can become an artificial intelligence (AI) assistant for the user. In this case, the character conversation device or character conversation system in this embodiment may also be called an AI assistant conversation device, an AI assistant display device, an AI assistant response output device, an AI assistant conversation system, an AI assistant display system, or an AI assistant response output system.
[0128] In the example shown in Figure 2A, the voice output unit 1140 of the artificial intelligence response output device 10010 is composed of a speaker. The artificial intelligence response output device 10010 is also equipped with a microphone 1139 that can pick up the user's voice. The artificial intelligence response output device 10010 can communicate with a communication device 19011 connected to the internet 19000 via a communication unit 1132. In the example shown in Figure 2A, the communication between the communication unit 1132 and the communication device 19011 is shown as wireless, but wired communication is also acceptable. The communication path from the communication unit 1132 to the internet 19000 may have both wired and wireless sections. The artificial intelligence response output device 10010 can communicate with the large-scale language model server 20001 via the communication device 19011 and the internet 19000. Furthermore, the artificial intelligence response output device 10010 can communicate with a second server 19002, which is different from the large-scale language model server 20001, via the communication device 19011 and the internet 19000. The configuration including the artificial intelligence response output device 10010 and the large-scale language model server 20001 may be considered as a single system.
[0129] Here, the large-scale language model server 20001 is a server equipped with a large-scale language model artificial intelligence. For example, it is a server with a multimodal large-scale language model that can process not only natural language text information but also other types of information.
[0130] The artificial intelligence response output device 10010 can communicate with the large-scale language model of the large-scale language model server 20001 via the internet 19000 using an API.
[0131] The artificial intelligence response output system of Example 2 includes a mobile information processing terminal 20010 used by user 230. The mobile information processing terminal 20010 is a so-called smartphone or tablet information processing terminal.
[0132] Here, an example of a mobile information processing terminal 20010 will be described using Figure 2B. The mobile information processing terminal 20010 includes a display panel 20011 which is a touch operation input panel, a control unit 20012, a memory 20026, a non-volatile memory 20027, an external power input interface 20013, a power supply 20014, a secondary battery 20015, a storage unit 20016, a video control unit 20017, an attitude sensor 20018, a position sensor 20019, a local LLM processing unit 20028, a communication unit 20020, an audio output unit 20021, a microphone 20022, a video signal input unit 20023, an audio signal input unit 20024, an imaging unit 20025, and the like.
[0133] The display panel 20011 is equipped with a touch input sensor and can accept touch input from the user 230's finger. The display panel 20011 displays images using a liquid crystal panel or an organic EL panel and can display images. The display panel 20011 may also be called a display unit.
[0134] The communication unit 20020 can be configured with a Wi-Fi communication interface, a Bluetooth communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 20020 of the mobile information processing terminal 20010 can communicate with the communication unit 1132 of the artificial intelligence response output device 10010. The control unit 20012 controls the communication unit 20020. Furthermore, using any of the communication methods of the communication unit 20020, the communication unit 20020 can communicate with the communication device 19011 connected to the Internet 19000. As a result, the mobile information processing terminal 20010 can communicate with various servers connected to the Internet 19000.
[0135] The power supply 20014 converts the AC current input from an external source via the external power input interface 20013 into DC current and supplies the necessary DC current to each part of the mobile information processing terminal 20010. The secondary battery 20015 stores the power supplied from the power supply 20014. In addition, the secondary battery 20015 supplies power to each part that requires power via the external power input interface 20013 when power is not supplied from an external source.
[0136] The video signal input unit 20023 receives video data by connecting an external video output device. Various digital video input interfaces are possible for the video signal input unit 20023. For example, it may be configured with a video input interface conforming to the HDMI (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided. The video signal input unit 20023 may also be various USB interfaces.
[0137] The audio signal input unit 20024 receives audio data by connecting an external audio output device. The audio signal input unit 20024 may be configured as an HDMI standard audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface. The audio signal input unit 20024 may also be various USB interfaces. In the case of an HDMI standard interface, the video signal input unit 20023 and the audio signal input unit 20024 may be configured as an interface with integrated terminals and cables.
[0138] The audio output unit 20021 is capable of outputting audio based on audio data input to the audio signal input unit 20024. The audio output unit 20021 is also capable of outputting audio based on audio data stored in the storage unit 20016. The audio output unit 20021 may be configured as a speaker. The audio output unit 20021 may also output built-in operation sounds or error warning sounds. Alternatively, the audio output unit 20021 may be configured to output as a digital signal to an external device, such as the Audio Return Channel function specified in the HDMI standard.
[0139] Microphone 20022 is a microphone that picks up sounds from the surrounding area of the mobile information processing terminal 20010, converts them into signals, and generates an audio signal. The microphone may record human voices, such as the user's voice, and the control unit 20012, described later, may perform speech recognition processing on the generated audio signal to obtain text information from the audio signal.
[0140] The imaging unit 20025 is a camera having an image sensor. The camera may be provided on the front of the display panel 20011 side of the mobile information processing terminal 20010, or on the back of the display panel 20011 side. Both a front camera and a rear camera may be provided. In this embodiment, the imaging unit 20025 will be described as having both a front camera and a rear camera.
[0141] The storage unit 20016 is a storage device that records various types of information, such as video data, image data, and audio data. The storage unit 20016 may be composed of a magnetic recording medium such as a hard disk drive (HDD) or a semiconductor memory such as a solid-state drive (SSD). For example, the storage unit 20016 may have various types of information, such as video data, image data, and audio data, pre-recorded in it at the time of product shipment. The storage unit 20016 may also record various types of information, such as video data, image data, and audio data, acquired from external devices or external servers via the communication unit 20020. The video data, image data, etc., recorded in the storage unit 20016 are output to the display panel 20011. The video data, image data, etc., recorded in the storage unit 20016 may also be output to external devices or external servers via the communication unit 20020.
[0142] The video control unit 20017 performs various controls related to the video signals input to the display panel 20011. The video control unit 20017 may also be called a video processing circuit and may be composed of hardware such as an ASIC, FPGA, or video processor. The video control unit 20017 may also be called a video processing unit or image processing unit. For example, the video control unit 20017 performs control such as switching between video signals, such as which video signals to input to the display panel 20011 from among the video signals to be stored in the memory 20026 and the video signals (video data) input to the video signal input unit 20023. The video control unit 20017 may also perform control to perform image processing on the video signals input from the video signal input unit 20023 and the video signals to be stored in the memory 20026. Image processing techniques include scaling (enlarging, reducing, and transforming images), brightness adjustment (changing brightness), contrast adjustment (modifying the image's contrast curve), and retinex processing (decomposing an image into its light components and changing the weighting of each component).
[0143] The attitude sensor 20018 is a sensor composed of a gravity sensor, an acceleration sensor, or a combination thereof, and can detect the attitude of the mobile information processing terminal 20010. Based on the attitude detection result of the attitude sensor 20018, the control unit 20012 may control the operation of each connected part.
[0144] The position sensor 20019 is a sensor that measures the position of the mobile information processing terminal 20010 using GPS (Global Positioning System) or the like. If the position of the mobile information processing terminal 20010 can be estimated based on information obtained through communication by the communication unit 20020, the position estimated by that method may be used instead of the measurement result of the position sensor 20019.
[0145] The non-volatile memory 20027 stores various data used by the mobile information processing terminal 20010. The data stored in the non-volatile memory 20027 includes, for example, data for various operations displayed on the display panel 20011 of the mobile information processing terminal 20010, display icons, data and layout information for user-operated objects, etc. Memory 20026 stores video data and device control data displayed on the display panel 20011. The control unit 20012 may read various software from the storage unit 20016, expand it into memory 20026, and store it there.
[0146] The local LLM processing unit 20028 has memory capable of holding a large-scale language model (LLM) and can perform inference of the large-scale language model based on the control of the control unit 20012. The hardware can be a so-called GPU (Graphics Processing Unit). The local LLM processing unit 20028 may perform training as well as inference. Note that the local LLM processing unit 20028 is not necessarily required if the execution of large-scale language model inference in the local environment of the mobile information processing terminal 20010 is not necessary.
[0147] The control unit 20012 controls the operation of each connected component. The control unit 20012 may also work in cooperation with a program stored in the memory 20026 to perform calculations based on information acquired from various components within the mobile information processing terminal 20010. The specific configuration of the control unit 20012 may include a CPU, and it may also be referred to as a processor or control circuit.
[0148] Next, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 2 of the present invention will be described using Figure 2C. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Note that in Figure 2C, the communication path such as the internet 19000 shown in Figure 2A has been omitted. In Figure 2C, the user 230 of the artificial intelligence response output device 10010 is also shown. Note that the operation or processing of each part of the artificial intelligence response output device 10010 described in Embodiment 2 of the present invention may be controlled by the client application 8010 as described in Figure 1E. Note that the operation or processing of the large-scale language model server 20001 described in Embodiment 2 of the present invention may be controlled by the server LLM application 8020 as described in Figure 1E.
[0149] Here, we will describe an example of the sequence of operations of the artificial intelligence response output device 10010 (the first example of operation). The artificial intelligence response output device 10010 loads the character operation program stored in the storage unit 1170 or the like into the memory 1109, and the control unit 1110 executes the character operation program, thereby enabling the various processes described below. The character operation program may be the same program as the client application 8010 described above, or it may be a different program. Furthermore, the character operation program may be a part of the program that constitutes the client application 8010 described above.
[0150] <First Operation Example> First, the artificial intelligence response output device 10010 is equipped with a microphone 1139. When user 230 speaks to character 19051, the microphone 1139 picks up the user's voice (words from the user) and converts it into an audio signal. Here, the character operation program executed by the control unit 1110 extracts the text of the words spoken by user 230 from the audio signal. This text is in natural language. Note that the extraction of the text of the words spoken by user 230 may be performed continuously for all words, or it may be started when a word is uttered by the user within a predetermined period after a trigger keyword. For example, a trigger keyword could be when the user says "Hello" followed by the character's name. For example, if the name of character 19051 is "Koto", then "Hello, Koto!" can be used as the trigger keyword.
[0151] The character operation program of the artificial intelligence response output device 10010 creates an instruction sentence (prompt) based on the text of the words spoken by the user 230, and sends the instruction sentence to the large-scale language model server 20001 using an API. Here, the instruction sentence may be metadata containing information written using notation with tags such as the markup format of a markup language, notation using predetermined symbols such as the Markdown format, or object notation of a predetermined script such as JSON. The instruction sentence contains natural language text information as the main message. The types of instruction sentences sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001 include setting instruction sentences that store instructions such as initial settings, and user instruction sentences that reflect instructions from the user. Type identification information that identifies whether the instruction sentence is a setting instruction sentence or a user instruction sentence may be stored in a part of the instruction sentence other than the main message. When the character operation program of the artificial intelligence response output device 10010 creates an instruction sentence (prompt) based on the text of the words spoken by the user 230, it creates a user instruction sentence and sends it to the large-scale language model server 20001.
[0152] Next, the large-scale language model of the artificial intelligence in the large-scale language model server 20001 performs inference based on the instruction sent from the artificial intelligence response output device 10010, and generates a response containing natural language text information based on the result. The large-scale language model server 20001 sends the response to the artificial intelligence response output device 10010 using an API. The response contains natural language text information as the main message. Here, the response may also be metadata containing information written in the same format as the instruction mentioned above (notation using tags such as the markup format of a markup language, notation using predetermined symbols such as the Markdown format, or object notation of a predetermined script such as JSON). If the same format as the instruction mentioned above is used in the response, type identification information may be stored in a part other than the main message to indicate that it is a different type of information from the setting instruction and the user instruction mentioned above. For example, information indicating that it is a response from the large-scale language model may be stored.
[0153] Next, the artificial intelligence response output device 10010 receives a response from the large-scale language model server 20001 and extracts the natural language text information stored as the main message in that response. Based on the natural language text information extracted from the aforementioned response, the character operation program of the artificial intelligence response output device 10010 uses speech synthesis technology to generate natural language speech that serves as a response to the user, and outputs it from the speaker-like speech output unit 1140 so that it sounds as if it were the voice of character 20051. This process may also be described as the character "uttering".
[0154] According to the first example of operation described above, using the artificial intelligence response output device 10010 shown in Figure 2C, or the artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001, the advanced natural language processing capabilities of the large-scale language model can be utilized via the API, enabling the character to provide a more appropriate response and engage in a more suitable conversation when the user speaks to it.
[0155] Next, we will explain another example (a second example of operation) of the sequence of operations of the artificial intelligence response output device 10010.
[0156] <Second Operation Example> In the first operation example, the actions performed by user 230 to the artificial intelligence response output device 10010 were mainly voice calls from user 230. In the first operation example, a series of operations were performed starting with the process of capturing user 230's voice with a microphone. In contrast, in the second operation example, in addition to the series of operations performed by the artificial intelligence response output device 10010 starting with the process of capturing user 230's voice with a microphone, user 230 can also perform actions to the artificial intelligence response output device 10010 through user operation via the operation input unit 1107 in Figure 1B. Here, an example of the operation input unit 1107 in Figure 1B is a mouse, keyboard, touch panel, etc.
[0157] In the second example of operation, the user 230 can perform an action on the artificial intelligence response output device 10010 by touching the user, which can be detected by the touch operation input sensor of the display unit 10011 in Figure 1B.
[0158] Furthermore, in the second example of operation, user 230 can also input user 230's operation input to the artificial intelligence response output device 10010 by operating the mobile information processing terminal 20010 and communicating from the mobile information processing terminal 20010 to the artificial intelligence response output device 10010.
[0159] Alternatively, the display panel 20011 of the mobile information processing terminal 20010 may display an information-storage image, such as a two-dimensional code containing information that the user wants to transmit to the artificial intelligence response output device 10010, and the imaging unit 1180 of the artificial intelligence response output device 10010 (Figure 1B) may capture this display. The control unit 1110 of the artificial intelligence response output device 10010 may extract information from the information-storage image, such as a two-dimensional code, captured by the imaging unit 1180, and obtain the information. Alternatively, the display panel 20011 of the mobile information processing terminal 20010 may display an image that the user wants to transmit to the artificial intelligence response output device 10010, and the imaging unit 1180 of the artificial intelligence response output device 10010 (Figure 1B) may capture this display. The control unit 1110 of the artificial intelligence response output device 10010 may perform image recognition processing on the image captured by the imaging unit 1180 and obtain the result of the image recognition processing.
[0160] Thus, in the second operational example, the types of actions that the user 230 can perform on the artificial intelligence response output device 10010 are greater than those described in the first operational example. As a result, the second operational example can acquire the results of actions performed by the user 230 other than the user's voice and generate an instruction sentence (prompt) to send to the large-scale language model server 20001 based on these results. This makes it possible to more favorably include types of information other than natural language text information extracted from the user's voice in the instruction sentence sent to the large-scale language model server 20001. Examples of types of information other than natural language text information extracted from the user's voice include images, videos, and audio.
[0161] Next, in the second operational example, the artificial intelligence response output device 10010 sends an instruction to the large-scale language model server 20001 using an API. In this operational example as well, the instruction may be metadata containing information written using notation such as markup format of a markup language using tags, notation such as Markdown format using predetermined symbols, or object notation of a predetermined script such as JSON. In this operational example as well, there are two types of instruction: setting instruction statements that store instructions such as initial settings, and user instruction statements that reflect instructions from the user. Type identification information that identifies whether an instruction is a setting instruction statement or a user instruction statement may be stored in a part of the instruction statement other than the main message. In this case, the instruction statement includes natural language text information as the main message. Furthermore, in this embodiment, in addition to natural language text information, the main message of the instruction statement may include non-natural language information sources such as images, videos, or audio as a type of information other than natural language text information. A specific method for including non-natural language information sources in an instruction statement will be described later.
[0162] The large-scale language model server 20001 has a multimodal large-scale language model that can process non-natural language information sources in conjunction with natural language text information. The large-scale language model server 20001 receives an instruction from the artificial intelligence response output device 10010. Based on the instruction, the multimodal large-scale language model performs inference and generates a response that includes natural language text information as a result of the inference. Here, since the artificial intelligence of the large-scale language model server 20001 is a multimodal large-scale language model, the response can include non-natural language information sources such as images, videos, or audio in addition to natural language text information.
[0163] The artificial intelligence response output device 10010 receives a response from the large-scale language model server 20001 and extracts natural language text information and non-natural language information sources such as images, videos, or audio stored as the main message in the response. The character operation program of the artificial intelligence response output device 10010 may use speech synthesis technology to generate natural language audio that serves as a response to the user based on the natural language text information extracted from the aforementioned response, and output it from the audio output unit 1140, which is a speaker, so that it sounds as if it were the voice of the character 19051 displayed on the display screen.
[0164] Furthermore, the character operation program of the artificial intelligence response output device 10010 may display natural language characters that serve as a response to the user on the display screen of the artificial intelligence response output device 10010, based on the natural language text information extracted from the aforementioned response. In this case, the characters may be displayed together with the character 19051, superimposed on the image of the character 19051, or displayed in place of the image of the character 19051. The video control unit 1160 may perform these specific processes.
[0165] Furthermore, the character operation program of the artificial intelligence response output device 10010 may display an image on the display screen of the artificial intelligence response output device 10010 in order to present it to the user, based on the image information of the non-natural language information source extracted from the aforementioned response. In this case, the image may be displayed together with the character 19051, superimposed on the image of the character 19051, or displayed in place of the image of the character 19051. These specific processes can be executed by the image control unit 1160.
[0166] Furthermore, the character operation program of the artificial intelligence response output device 10010 may display the video information of the non-natural language information source extracted from the aforementioned response on the display screen of the artificial intelligence response output device 10010 in order to present it to the user. In this case, the video may be displayed together with the character 19051, superimposed on the video of the character 19051, or displayed in place of the video of the character 19051. The video control unit 1160 may perform these specific processes.
[0167] Furthermore, the character operation program of the artificial intelligence response output device 10010 may output speech generated based on the speech information of the non-natural language information source extracted from the aforementioned response from the speech output unit 1140, which is a speaker.
[0168] According to the second operational example of the artificial intelligence response output device 10010 shown in Figure 2C, or the artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001, as described above, the advanced natural language processing and non-natural language information processing capabilities of the multimodal large-scale language model can be utilized via the API. In addition to responses based on natural language text, responses based on non-natural language information sources can be provided in response to user actions directed at the character, enabling more appropriate conversations.
[0169] <Example of Operation> Next, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 2 of the present invention will be described using Figure 2D. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 2D shows an example of natural language text and non-natural language information sources such as images for the main message of an instruction sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, and an example of natural language text and non-natural language information sources such as images for the main message of the server response that is the response to it. In this embodiment, images, videos, audio, etc. can be used as non-natural language information sources, but in Figure 2D, an example of an image is shown as a non-natural language information source. The series of operations or processes described using this figure may be controlled by the cooperation of the client application 8010 and the server LLM application 8020, as described in Figure 1E.
[0170] Furthermore, Figure 2D shows the exchange of instructions and responses in chronological order, from the first round of setting instructions and user instructions and their responses to the second round of user instructions and their responses. Here, the instructions and responses shown in the example of Figure 2D include non-natural language information sources 20061 and 20062. In the example of Figure 2D, both non-natural language information sources 20061 and 20062 are images.
[0171] In Figure 2D, for the sake of simplicity, an image of the non-natural language information source 20061 is shown embedded within the instruction text. However, there are multiple methods for transmitting or specifying the data of the non-natural language information source 20061 in the instruction text sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001. The artificial intelligence response output device 10010 may use any one of these methods, or switch between them. An example of each method will be described below.
[0172] The first method for transmitting or specifying non-natural language information source data in an instruction is used, for example, when the non-natural language information source to be specified is located on a server or other location connected to a network such as the Internet. A specific example of the first method is to specify a non-natural language information source file located on a network such as the Internet using information such as tags and symbols within the instruction, along with the network location information (so-called URL, etc.) and file name.
[0173] For example, a tag used to specify an image in a markup language.<img src=“****”> By using this tag and writing the location and filename information of the image file in the **** part, you can specify an image that exists on a network such as the internet. Alternatively, you can use a tag that specifies a video in a markup language.<video src=“****”> You can also specify a video that exists on a network such as the internet by using the **** part and writing the location information and file name information of the video file. Alternatively, you can use a tag that specifies audio in a markup language.<audio src=“****”> You can specify audio files located on a network such as the internet by using the format and writing the location and filename information of the audio file in the **** section. Alternatively, if using the JSON format notation, you can specify images located on a network such as the internet by preparing a key such as img_src and writing the location and filename information of the image file as its value. For video and audio files, you just need to prepare the respective keys and values. The example given is just one example, and you may use other proprietary formats. In any case, the information specifying the location and filename information of the non-natural language information source file should be stored in the instruction statement.
[0174] As in the first method, when information specifying the location and filename of the non-natural language information source file is stored in the instruction statement, it is not necessary to store the data of the non-natural language information source file in the instruction statement. Therefore, the amount of data in the instruction statement can be reduced. In the first method, the large-scale language model server 20001 that receives an instruction statement specifying non-natural language information source data can use the location and filename information of the non-natural language information source file stored in the instruction statement to obtain the non-natural language information source file located at a server or other location connected to a network such as the Internet.
[0175] Here, we will explain how location information and file name information are input when the artificial intelligence response output device 10010 specifies non-natural language information source data in an instruction sentence using the first method. As explained in Figure 2C, there are types of actions that the user 230 can take with the artificial intelligence response output device 10010 other than the user 230's voice. Therefore, for example, the user 230 may input location information such as a URL for specifying non-natural language information source data, file name information, etc., through user operation (e.g., mouse, keyboard, touch panel) via the operation input unit 1107 in Figure 1B.
[0176] Furthermore, in the artificial intelligence response output device 10010, the control unit 1110 may cooperate with the memory 1109 to execute a web browser program and display the GUI of the web browser program on the display screen of the artificial intelligence response output device 10010. User operations on the GUI of the web browser program may be received via the operation input unit 1107 (for example, mouse, keyboard, touch panel) or by user touch operations detectable by the touch operation input sensor of the display unit 10011, and non-natural language information source data such as images, videos, and audio selected on the browser screen of the web browser program may be used as the data to be specified in the instruction statement. In this case, the web browser program should acquire the location information and file name information of the non-natural language information source data and pass them to the character operation program.
[0177] Alternatively, user 230 may operate the mobile information processing terminal 20010 to communicate with the artificial intelligence response output device 10010, thereby inputting location information such as a URL for specifying non-natural language information source data to the artificial intelligence response output device 10010. Alternatively, as explained in Figure 2C, location information such as a URL for specifying non-natural language information source data, file name information, etc., may be input by displaying an information-storing image such as a two-dimensional code on the display panel 20011 of the mobile information processing terminal 20010, performing image recognition processing on the image captured by the imaging unit 1180 of the artificial intelligence response output device 10010, and obtaining the result of the image recognition processing.
[0178] Furthermore, the use of the first method for transmitting or specifying non-natural language information source data in an instruction statement is not limited to cases where the non-natural language information source file already exists in a location such as a server connected to a network such as the Internet. For example, if it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 1170 of the artificial intelligence response output device 10010 in an instruction statement, the artificial intelligence response output device 10010 may upload the non-natural language information source data to a second server 19002 via the Internet 19000 and include the Internet location information (so-called URL, etc.) and file name of the uploaded non-natural language information source data on the second server 19002 in the instruction statement. In this case, the second server 19002 functions as a so-called intermediate server.
[0179] Similarly, if it is desired to include non-natural language information source data such as images, videos, and audio stored in the storage unit 20016 of the mobile information processing terminal 20010 in the instruction text, the mobile information processing terminal 20010 may upload the non-natural language information source data to the second server 19002 via the internet 19000. The mobile information processing terminal 20010 or the second server 19002 may transmit the internet location information (so-called URL, etc.) and file name of the non-natural language information source data on the second server 19002 to the artificial intelligence response output device 10010, and the character operation program of the artificial intelligence response output device 10010 may include the internet location information (so-called URL, etc.) and file name of the non-natural language information source data uploaded to the second server 19002, which it has acquired, in the instruction text.
[0180] Furthermore, the character operation program of the artificial intelligence response output device 10010 may work in cooperation with the memory 1109 and the storage unit 1170 to construct a media server within the artificial intelligence response output device 10010 that can be accessed from other servers via the internet 19000. In this case, when the artificial intelligence response output device 10010 specifies non-natural language information source data in an instruction statement using the first method, it only needs to store in the instruction statement location information on the internet (such as a URL) indicating the media server constructed within the artificial intelligence response output device 10010 itself, and the file name of the corresponding non-natural language information source data.
[0181] Next, a second method for specifying the transmission or designation of non-natural language information source data in an instruction is, for example, simply to store (attach) the non-natural language information source data itself in the instruction (prompt) and send it. Generally, non-natural language information source data such as images, videos, and audio are larger in data size than natural language text information. Therefore, in this case, the data size of the instruction (prompt) will be larger than that of the first method. The character operation program of the artificial intelligence response output device 10010 can store the non-natural language information source data that it wants to store (attach) in the instruction (prompt) in memory 1109, and when sending the instruction (prompt), it can store (attach) the data in the instruction (prompt) from memory 1109 via the communication unit 1132 and output it to the large-scale language model server 20001. The non-natural language information source data that the character operation program of the artificial intelligence response output device 10010 stores in memory 1109 may be acquired by the communication unit 1132 via the internet 19000, acquired by the communication unit 1132 from the mobile information processing terminal 20010, or read from the storage unit 1170 and stored in memory 1109.
[0182] As described above, the artificial intelligence response output device 10010 can transmit or specify non-natural language information source data using instruction sentences.
[0183] The large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information. As shown in the example in Figure 2D, through the first round of user instructions, it can acquire images of a swimming pool and poolside, which are non-natural language information sources 20061, and natural language text information. As a result of this inference, it can output natural language text information as shown in the figure, in response to the first round of user instructions.
[0184] Furthermore, since the large-scale language model server 20001 is a multimodal large-scale language model that can process non-natural language information sources together with natural language text information, as shown in the example in Figure 2D, in the response to the second round of user instructions, the large-scale language model server 20001 can include the non-natural language information source 20062 generated by the inference of the multimodal large-scale language model in its response and transmit it to the artificial intelligence response output device 10010. In Figure 2D, the non-natural language information source 20062 is an example of an image in which a circle image 20063 is attached to an image of a swimming pool and poolside, which is the non-natural language information source 20061. Note that the non-natural language information source 20062 stored in the response is not limited to the image shown in Figure 2D, but may also be a video or audio.
[0185] In cases where the response from the large-scale language model server 20001 includes non-natural language information sources other than natural language text information, the first method or a method similar to the second method described above, in which the artificial intelligence response output device 10010 transmits or specifies non-natural language information source data in the instruction statement, may also be used.
[0186] Specifically, in a method similar to the first method described above, the large-scale language model server 20001 may store information specifying the location and file name of the non-natural language information source file in the instruction statement in its response. The non-natural language information source 20062, such as images, videos, and audio, may be kept by the large-scale language model server 20001, or it may be transferred to and kept by the second server 19002, which functions as an intermediate server. In either case, the large-scale language model server 20001 may store information specifying the location and file name of the non-natural language information source file in the instruction statement in its response. The artificial intelligence response output device 10010, having received the response, may use the location and file name information of the non-natural language information source file described in the instruction statement to access the large-scale language model server 20001 or the second server 19002 to obtain the non-natural language information source 20062.
[0187] Furthermore, specifically, as a method similar to the second method described above, the large-scale language model server 20001 may store (attach) the file data of the non-natural language information source 20062 itself in the response and send it to the artificial intelligence response output device 10010. The artificial intelligence response output device 10010 can acquire the data of the non-natural language information source 20062 stored (attached) in the instruction and use it for various outputs to the user 230.
[0188] As described above with reference to Figure 2D, the operation of the artificial intelligence response output device 10010 and artificial intelligence response output system of Embodiment 2 involves the transmission and reception of instruction sentences and responses between the character displayed on the artificial intelligence response output device 10010 and the user 230, enabling conversation using natural language information and / or non-natural language information such as images, videos, and audio. This makes it possible to achieve more sophisticated and natural conversations, as shown in each message in Figure 2D.
[0189] <Example of Display> Next, an example of the display of the artificial intelligence response output device 10010 of Embodiment 2 of the present invention will be described using Figure 2E. In the example in Figure 2E, an example is shown of displaying the response from the large-scale language model to the user instruction sentences described in Figures 2A to 2D on the display unit 10011 of the artificial intelligence response output device 10010. Specifically, this is an example of displaying the text 10063 of the natural language information source data, the image 10064 of the non-natural language information source data, and / or the video 10065 of the non-natural language information source data, which are the response from the large-scale language model, together with the video of the character 19051 on the display unit 10011. The text 10063, the image 10064, and / or the video 10065, which are the response from the large-scale language model, may be displayed superimposed in front of the video of the character 19051, as shown in Figure 2E.
[0190] Furthermore, the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, may be displayed together with the video of character 19051 without being superimposed on it. The display in Figure 2E is just one example, but for example, if user 230 adjusts the volume of the audio output of the audio output unit 1140 of the artificial intelligence response output device 1001 to the minimum or sets the audio output to OFF by operating via the operation input unit 1107 or the touch operation input sensor of the display unit 10011, user 230 will not be able to confirm the response from the large-scale language model by voice. In this case, the control unit 1110 may control the system to start a display mode in which the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, are displayed together with the video of character 19051, as shown in Figure 2E.
[0191] In this way, even when the user 230 wishes to reduce voice output, the artificial intelligence response output device 10010 can be used more favorably. Furthermore, the user 230 may manually switch ON / OFF the display mode, which displays the text 10063, image 10064, and / or video 10065, which are responses from the large-scale language model, together with the image of the character 19051, via the operation input unit 1107 or the touch operation input sensor of the display unit 10011. As shown in the display example in Figure 2E, the multimodal artificial intelligence response output device 10010 can more favorably output responses from the large-scale language model.
[0192] As described above, the artificial intelligence response output device 10010 and artificial intelligence response output system according to Embodiment 2 can provide users with a more advanced conversational experience that includes not only natural language information but also non-natural language information by using a multimodal large-scale language model.
[0193] In the above description of Example 2, an example was described in which the large-scale language model possessed by the large-scale language model server 20001 is used as the large-scale language model. In contrast, the artificial intelligence response output device 10010 may be equipped with a local LLM processing unit 10028 as shown in Figure 1B, and the multimodal large-scale language model possessed by the local LLM processing unit 10028 may be used. In this case, the multimodal large-scale language model possessed by the local LLM processing unit 10028 may be used instead of the multimodal large-scale language model possessed by the large-scale language model server 20001. In this case, the operation or processing of the local LLM processing unit 10028 may be controlled by the local LLM application 8015 as described in Figure 1E.
[0194] In this case, in the above description of Embodiment 2, the multimodal large-scale language model possessed by the large-scale language model server 20001 can be replaced with the multimodal large-scale language model possessed by the local LLM processing unit 10028 of the artificial intelligence response output device 10010. In this case as well, by using the multimodal large-scale language model, it is possible to provide the user with a more advanced conversational experience that includes not only natural language information but also non-natural language information.
[0195] <Example 3> Next, Example 3 of the present invention will be described. The artificial intelligence response output device 10010 of Example 3 of the present invention is an artificial intelligence response output device 10010 that has the function of displaying a character or avatar as described in Example 2, and further has the function of switching the character or avatar to be displayed. In this example, the differences from Example 2 will be explained, and the same configuration as in these examples will not be repeated. Note that since it is redundant to repeatedly write "character or avatar", in the following description of this example, it will simply be written as "character".
[0196] Here, using Figure 3A, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 3 of the present invention will be described. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3A shows an example of operation in which the character to be displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched from among a plurality of character candidates. The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 can switch the displayed character based, for example, on operation input input to the operation input unit 1107 or operation detected by the touch operation input sensor of the display unit 10011.
[0197] In the example in Figure 3A, in addition to character 19051 (named "Koto") used in the explanations of Figures 2A to 2E, characters 19052 (named "Tom") and 19053 (named "Necco") are shown. Characters 19051 (named "Koto") and 19052 (named "Tom") are human characters, while character 19053 (named "Necco") is a cat character. Switching the display of characters on the display unit 10011 can be done by rendering different virtual 3D characters for each character and then switching the displayed image on the display unit 10011.
[0198] Furthermore, when the character operation program executed by the control unit 1110 switches the display of the characters shown on the display unit 10011, it is preferable that the synthesized voice used for each character's "utterance" is also changed. This can be done by pre-storing synthesized voice data with corresponding voice to each character in the storage unit 1170, and then performing the synthesized voice change process when switching the display of the characters.
[0199] In the example shown in Figure 3A, the system is configured so that user 230 can converse with any of the characters. The artificial intelligence response output device 10010 in Figure 3A assigns different roles, names, conversational characteristics, or personalities to each of these characters. Furthermore, the memories of each character based on their conversation history are managed separately for each character.
[0200] Therefore, the artificial intelligence response output device 10010 constructs a database shown in Figure 3B in the storage unit 1170, and uses this database to manage character settings and the character's conversation history.
[0201] Next, an example of the operation of the artificial intelligence response output device 10010 of Embodiment 3 of the present invention will be described using Figure 3B. This can also be described as an example of the operation of an artificial intelligence response output system including the artificial intelligence response output device 10010 and the large-scale language model server 20001. Specifically, Figure 3B is an explanatory diagram of the database 19200 for managing the character settings and conversation history of multiple characters displayed on the display unit 10011 of the artificial intelligence response output device 10010.
[0202] The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 constructs the database 19200 in the storage unit 1170, for example. The character ID is an identification number that identifies each of the multiple characters that can be displayed by the artificial intelligence response output device 10010, and may be a natural number or use the alphabet, etc. The name is data of the name of each of the multiple characters that can be displayed by the artificial intelligence response output device 10010.
[0203] The configuration instruction is natural language text information that describes the settings of each of the multiple characters that can be displayed by the artificial intelligence response output device 10010, such as their role, name, conversational characteristics, or personality. Since this configuration instruction is the main data of the configuration instruction transmitted from the artificial intelligence response output device 10010 to the large-scale language model server 20001, it is desirable that the content be such that the large-scale language model of the artificial intelligence in the large-scale language model server 20001 can read it directly.
[0204] The conversation history, which continues as Conversation History 1, 2, ..., is a record of the conversation between each character and the user, and is recorded separately for each character. Since this conversation history will be included in the natural language text information, which is the main data of the setting instruction message sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, it is desirable that the content be such that the large-scale language model of the artificial intelligence in the large-scale language model server 20001 can read it directly.
[0205] The character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010, when the character displayed on the display unit 10011 of the artificial intelligence response output device 10010 is switched, uses the database 19200 in Figure 3B to select and switch the setting instructions and conversation history used for the natural language text information, which is the main data of the setting instructions sent from the artificial intelligence response output device 10010 to the large-scale language model server 20001, so that they correspond to the character displayed on the display unit 10011 of the artificial intelligence response output device 10010. In addition, each time a conversation takes place between the user 230 and the character, the character operation program records the history of that conversation in the conversation history area of the database 19200 in Figure 3B that corresponds to the character displayed on the display unit 10011.
[0206] By using the database 19200, the character operation program executed by the control unit 1110 of the artificial intelligence response output device 10010 uses the same large-scale language model of the same artificial intelligence on the same large-scale language model server 20001 to establish a conversation between the user 230 and the character. However, from the user's perspective, the uniqueness of each character's personality and other settings is preserved, and the memory of different conversations for each character continues. From the user's perspective, this is preferable because it is perceived that the identity of the character's role, name, conversational characteristics, or personality settings and memories from previous conversations is more readily maintained for each character. This can also be expressed as ensuring a pseudo-identity of each character from the user's perspective.
[0207] Therefore, even when the artificial intelligence response output device 10010 is configured to switch between displaying characters from among multiple character candidates on the display unit 10011, the operation using the database 19200 described above will result in a less jarring experience for the user in conversations with each character, and will allow them to share memories with each of the multiple characters, providing a more enjoyable character conversation experience.
[0208] Furthermore, if the user is prevented from editing the setting instructions for multiple characters, the settings for each character, such as their role, name, conversational characteristics, or personality, can be maintained in a state close to the intentions of the provider of the artificial intelligence response output device 10010 or the creator of the character's content. Alternatively, the user may be allowed to edit the character setting instructions in response to input from the operation input unit 1107 or the like. In this case, the user can customize the character's role, name, conversational characteristics, or personality, and converse with a character they have set up themselves. In this case, the character's 3D model, its rendered image, and the type of synthesized voice of the character may also be replaced accordingly.
[0209] Next, an example of a response template database (DB) in the artificial intelligence response output device 10010 capable of displaying multiple characters, as described in Figures 3A and 3B, will be explained using Figure 3C. In the example in Figure 3C, the response template DB 19300 stores the template responses that the artificial intelligence response output device 10010 outputs for each condition assigned a condition number. For these conditions, in the example in Figure 3C, individual response templates are set for each of the multiple characters. For example, for each of the three characters described in Figures 3A and 3B—character 1: Koto, character 2: Tom, and character 3: Necco—a response template for each condition is stored.
[0210] Here, a specific example of a standard response for character 1: Koto, stored in the response template DB 19300 in Figure 3C, and the operation of the artificial intelligence response output device 10010 using it, is as follows.
[0211] First, for example, as in condition number 1, if the user inputs "Good morning" via the touch panel, microphone 1139, or operation input unit 1107, the response should be output using "Good morning" or "Today is [Month] [Day]." as a standard response. The part marked with "[Month] [Day]" can be generated using information stored in the memory 1109 or the like of the artificial intelligence response output device 10010.
[0212] Furthermore, in the example of a standard response phrase in the database shown in Figure 3C, if multiple standard response phrases separated by a slash ( / ) are stored, the control unit 1110 can be controlled to randomly select one of the standard response phrases using a random number or the like and output the response. This eliminates and improves the situation where the response becomes monotonous under the same conditions. The explanation for the examples of condition numbers 2, 3, and 4 is the same as for the example of condition number 1. The control unit 1110 should be controlled to output using the standard response phrases shown in Figure 3C for each example of the condition content shown in Figure 3C.
[0213] Next, we will explain an example of condition number 5 shown in Figure 3C. Condition number 5 is an example of control in which, when the control unit 1110 cannot understand the meaning of user input obtained via the touch panel, microphone 1139, or operation input unit 1107 as natural language, or when there is an obvious grammatical error in the user input, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that" or "I may not know about that." By responding in this way, the user can be prompted to input again, and the system can wait for corrected user input.
[0214] Next, an example of condition number 6 shown in Figure 3C will be explained. Condition number 6 is an example where the control unit 1110 detects an error (abnormal state) in any of the parts constituting the artificial intelligence response output device 10010 shown in Figure 1B, and user input is received via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 controls the system to output a response using the standard response phrase "It seems to be malfunctioning." By responding in this way, the system can explain to the user that the artificial intelligence response output device 10010 is malfunctioning and prompt the user to take action to address the error.
[0215] The above describes a specific example of a standard response for character 1: Koto and the operation of the artificial intelligence response output device 10010 using it. However, in the example of the standard response DB 19300 in Figure 3C, standard responses for other characters such as character 2: Tom or character 3: Necco are also stored.
[0216] In the example shown in Figure 3C, the control unit 1110 selects a corresponding predefined response from the predefined response database 19300 based on the character displayed in the artificial intelligence response output device 10010 and the current conditions, and uses it to control the output as a response from the character. For example, in the example of the predefined response database 19300 in Figure 3C, even under the same conditions, the predefined response is changed to an expression or content that corresponds to the personality of each character. As a result, the artificial intelligence response output device 10010 can provide the user with conversations that correspond to the personality of the displayed character. The user can feel that each character is a being with a more consistent personality. This makes it possible to realize an artificial intelligence response output device 10010 that gives multiple characters a greater sense of reality.
[0217] The artificial intelligence response output device 10010 may output a response using the response template database described with reference to Figure 3C, instead of a response from a large-scale language model such as the local large-scale language model provided by the artificial intelligence response output device 10010, the large-scale language model provided by the large-scale language model server 19001, or the multimodal large-scale language model provided by the large-scale language model server 20001. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the response template database.
[0218] The response template DB 19300 shown in Figure 3C, as described above, is stored in the storage unit 1170, and the control unit 1110 of the artificial intelligence response output device 10010 can use it. However, the response template DB 19300 shown in Figure 3C may also be provided on the large-scale language model server 20001 side. In this case, for example, in the configuration of Figure 1D, the response template DB 19300 shown in Figure 3C is stored in the various DBs 22021 of the storage unit 22020, and the control unit 22030 of the large-scale language model server 20001 can generate responses using the response template DB 19300. The control unit 22030 of the large-scale language model server 20001 can then send the response generated using the response template DB 19300 stored in the various DBs 22021 of the storage unit 22020 to the artificial intelligence response output device 10010, instead of the response generated by the large-scale language model. In this way, even if the artificial intelligence response output device 10010 is not equipped with a response template database, it becomes possible to generate responses using a response template database.
[0219] Here, using Figures 3D, 3E, and 3F, we will describe an example of control that has been further improved to allow the user to perceive each character as having a more consistent personality in the artificial intelligence response output device 10010 and artificial intelligence response output system according to Embodiment 3 of the present invention. Figures 3D, 3E, and 3F describe an example of control regarding the characteristics of character conversations in the artificial intelligence response output system when a client application and an LLM application cooperate to perform a series of artificial intelligence response outputs. Specifically, this is an example in which the response generation process of the artificial intelligence response output device 10010 more preferably combines response generation processing using a large-scale language model on the network or response generation processing using a local large-scale language model (local LLM processing unit 10028, etc.) that the artificial intelligence response output device 10010 has internally, and response generation processing using a response template database to generate response output.
[0220] First, an example of the operation of the artificial intelligence response output device and artificial intelligence response output system will be explained using Figure 3D. The configuration of the artificial intelligence response output system in Figure 3D is an improvement on the artificial intelligence response output system shown in Figure 2C. Components with the same reference numerals as in Figure 2C are the same as those in Figure 2C, so repeated explanations will be omitted.
[0221] In the example shown in Figure 3D, as described in explanation 7010, the artificial intelligence response output device 10010 includes a client application as software that is deployed in the memory 1109 shown in Figure 1B and executed by the control unit 1110. This corresponds to the client application 8010 in Figure 1E.
[0222] The client application can control the input of user instructions as described in the figures of the embodiments described above. The client application also uses user instructions to send and receive information with the large-scale language model. The client application obtains a response from the large-scale language model. Based on this, the client application can control the parts shown in Figure 1B, including the display unit 10011 and the audio output unit 1140, to output a response to the user. The user instructions can be generated by the client application in response to user input. As an example of a user input interface that accepts user input, as described in Figure 2C, an example of the operation input unit 1107 in Figure 1B is a mouse, keyboard, touch panel, etc. The microphone 1139 that picks up the user's voice can also be considered a user input interface. The communication unit 1132 that communicates with the mobile information processing terminal 20010 used by the user, which is a smartphone or tablet information processing terminal, can also be considered a user input interface. Furthermore, the output interface for the client application to output a response to the user includes a display unit 10011 that outputs the response from the large-scale language model in media such as text, images, or video, and an audio output unit 1140 that outputs the response from the large-scale language model in audio.
[0223] Here, the client application can control the insertion of predefined phrases into the responses output by the artificial intelligence response output device 10010. This includes the output control of predefined phrases as explained in Figure 3C. Examples of such predefined phrase insertions include the insertion of predefined greeting phrases as explained in Figure 3C, as well as the insertion of predefined acknowledgment phrases. These predefined phrases can be stored in the storage unit 1170 in Figure 1B of the artificial intelligence response output device 10010, and then loaded into the memory 1109 for use by the client application.
[0224] Furthermore, in the example of Figure 3D, as explained in 7020, there is an LLM application on the large-scale language model side that controls the input and output of information to the large-scale language model. As in the example of Figure 3D, when the LLM application is on the large-scale language model server 20001 side, the LLM application is loaded into the memory 22031 of the large-scale language model server 20001 and executed by the control unit 22030 of the large-scale language model server 20001. The LLM application can accept preset instructions to adjust the output of the large-scale language model to a specific specification. While user instructions are instructions whose content changes each time a user makes an input, the preset instructions for adjusting the output of the large-scale language model to a specific specification do not change with each user input and are given on a regular basis, so these preset instructions may also be called regular instructions. Alternatively, these may also be called instructions for customizing the output of the large-scale language model. For example, a preset instruction is possible to set the endings of the natural language response sentences generated by the large-scale language model to a predetermined catchphrase. Furthermore, preset instructions can be used to insert predetermined interjections into the natural language response sentences generated by the large-scale language model. The preset instructions for the LLM application may be set by sending instructions from the client application of the artificial intelligence response output device 10010 to the LLM application. Alternatively, the preset instructions for the LLM application may be set by communicating with the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user.
[0225] In the example shown in Figure 3D, an example is shown where the large-scale language model has an LLM application that controls the input and output of information to the large-scale language model. This corresponds to the server LLM application 8020 in Figure 1E. In contrast, as another modification, the artificial intelligence response output device 10010 may be configured to have an LLM application that controls the input and output of information to the local LLM processing unit 10028 of the artificial intelligence response output device 10010. This corresponds to the local LLM application 8015 in Figure 1E. In this case, the LLM application may be software that is loaded into memory 1109 and executed by the control unit 1110. In this case, the LLM application that controls the input and output of information to the local LLM processing unit 10028 and the client application described above may be different applications, but they may also be the same application.
[0226] In the example shown in Figure 3D, the artificial intelligence response output device 10010 can send not only instruction text but also control information to the large-scale language model server 20001. This control information may be sent together with the instruction text. The transmission of the control information may occur before the transmission of the instruction text. The client application can control the transmission of this control information and instruction text.
[0227] The following describes an example of the details of the control information. First, the control information may contain authentication information for logging into the LLM application. For example, a user may create an account in the LLM application in advance and generate authentication information including identification information and a password. The creation of the account may be performed by communicating with the LLM application via the communication unit 1132 of the artificial intelligence response output device 10010. Alternatively, the creation of the account may be performed by communicating with the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user. The communication unit 1132 of the artificial intelligence response output device 10010 sends the authentication information to the LLM application and performs login through the authentication process. If authentication is successful, the artificial intelligence response output device 10010 and the LLM application establish communication as the user corresponding to the authentication information. As a result, the user can use the information available to their account from the information stored in the memory area of the LLM application. In the case of the LLM application on the large-scale language model server 20001 (server LLM application 8020), the memory area for the LLM application can be provided in the memory 22031 or storage unit 22020 of the large-scale language model server 20001. Furthermore, if the LLM application is a local LLM application (local LLM application 8015) that controls the input and output of information to and from the local LLM processing unit 10028 of the artificial intelligence response output device 10010, the memory area for the LLM application can be provided in the storage unit 1170 or memory 1109 shown in Figure 1B.
[0228] The information available to the user's account includes the preset instructions mentioned above. Furthermore, as shown in Figures 3A, 3B, or 3C, if the artificial intelligence response output device 10010 can switch between and display multiple characters, the LLM application may store information corresponding to each of the multiple characters in its memory area. In this case, if the control information transmitted from the artificial intelligence response output device 10010 to the LLM application includes a character ID to identify the character, the LLM application can determine which character's preset instruction the user wants to apply. In other words, this is equivalent to including character-specific preset instruction switching information in the control information and enabling the LLM application to switch preset instructions according to that switching information. Also, if the user wants to set character-specific preset instructions for the LLM application, they can store the preset instruction setting information in the control information transmitted from the artificial intelligence response output device 10010 to the LLM application and send it. Alternatively, the preset instruction setting information may be transmitted to the LLM application via the communication unit of an information terminal such as a smartphone or personal computer used separately by the user. The LLM application, upon receiving the preset instruction setting information, stores the preset instruction information based on that setting information in its memory. As described above, the preset instruction information may be stored for each character. Alternatively, it may be stored for each user account. Once the preset instruction settings are complete, the user transmits a character ID identifying the desired character from the artificial intelligence response output device 10010 to the LLM application, allowing the LLM application to select the preset instruction to apply to that character and apply it to the large-scale language model. In this case, it is not necessary to transmit information corresponding to the preset instruction to the LLM application each time an instruction is sent, thus reducing the amount of communication data.
[0229] Another example is that, each time an instruction is sent, the control information transmitted from the artificial intelligence response output device 10010 to the LLM application may contain information equivalent to a preset instruction and transmit it. In this case, although the transmission frequency of information equivalent to a preset instruction will increase, it will be possible to omit preparations such as prior user account registration and prior setting of preset instructions. Alternatively, each time an instruction is sent, the information equivalent to the preset instruction may be stored in the setting instruction area of the instruction rather than the user instruction area and transmitted. In this case as well, it will be possible to omit preparations such as prior user account registration and prior setting of preset instructions.
[0230] Next, using Figure 3E, we will explain a specific example of how the client application and the LLM application work together to generate response output by more favorably combining the response generation process using a large-scale language model and the response generation process using a response template database.
[0231] Figure 3E is an explanatory diagram of a control example when the artificial intelligence response output system is a character conversation system or an AI assistant system. Furthermore, the example in Figure 3E is an example where the artificial intelligence response output system can switch between and display multiple characters or AI assistants, as shown in Figures 3A, 3B, or 3C. Here, the example in Figure 3E uses two character examples for explanation. Using only two character examples simplifies the explanation; it is possible to switch between and display three or more characters, including the other characters described in previous embodiments. These two characters include a character with character ID 3 and name Necco. This character is the same character described in the embodiments described previously. Additionally, the above two characters include a new character with character ID 4 and name Airia.
[0232] Here, in the table of FIG. 3E, the character IDs, names, and display examples of these two characters are shown. Further, in the table of FIG. 3E, for each character, examples of the fixed phrases for greetings and responses that the client application inserts are shown. The process of the client application inserting the fixed phrases for greetings and responses has been described in FIG. 3C, so repeated description is omitted.
[0233] In the table of FIG. 3E, in order to give these characters personality, characteristic expressions (keywords) are used in both these fixed phrases for greetings and the fixed phrases for responses. For example, since Necco is a character modeled after a cat, "nya" is attached to the end of the fixed phrase. Also, Necco's first-person pronoun is "boku". For example, Airia has a tone where the end of the fixed phrase is "desuwa" or "masuwa". That is, it is set so that the end becomes "wa". Also, Airia's first-person pronoun is "watakushi". These keywords may be referred to as keywords indicating the personality of the character. These keywords can express the personality of the character by being commonly used in the conversations of each character. When performing the character switching process as described in FIG. 3A, the client application may switch the keywords according to the setting information such as the table of FIG. 3E.
[0234] Here, in the example of FIG. 3E, in addition to the setting of the fixed phrases inserted by the client application, in the preset instructions of the LLM application, there are settings for speech habits and responses. The preset instructions of the LLM application can be set by sending control information from the client application to the LLM application. Since this process is as described in FIG. 3D, repeated explanations are omitted. Here, in the example of FIG. 3E, in the settings of the speech habits and responses of the preset instructions of the LLM application, instructions are included to also output the characteristic expressions (keywords) common to the greeting fixed phrases and response fixed phrases inserted by the client application in the responses of the large language model. Thereby, the personality of the character can also be reflected in the response output of the large language model. Specifically, since Necco is a character modeled after a cat, in the settings of the speech habits and responses of the preset instructions of the LLM application, similar to the settings of the greeting fixed phrases and response fixed phrases inserted by the client application, the instruction is set so that "nya" is added to the end of the word. Also, similar to the settings of the greeting fixed phrases and response fixed phrases inserted by the client application, in the setting of the speech habits of the preset instructions of the LLM application, the instruction is set so that Necco's first person is "boku". Similarly, for the character Airia, in the settings of the speech habits and responses of the preset instructions of the LLM application, similar to the settings of the greeting fixed phrases and response fixed phrases inserted by the client application, the instruction is set so that "wa" is added to the end of the word. Also, similar to the greeting fixed phrases and response fixed phrases inserted by the client application, in the setting of the speech habits of the preset instructions of the LLM application, the instruction is set so that Airia's first person is "watakushi". That is, in the example of the table information shown in FIG. 3E, the characteristic expressions (keywords) that overlap with the characteristic expressions (keywords) indicating the personality of each character included in the settings of the greeting fixed phrases and response fixed phrases inserted by the client application are also included in the settings of the speech habits and responses of the preset instructions of the LLM application.Furthermore, if the character switching process described in Figure 3A is performed, the client application should switch the settings related to the character's conversational characteristics, as shown in Figure 3E, including the settings for standard greetings, standard interjections, and the preset instructions for catchphrases and interjections in the LLM application, according to the character switching process.
[0235] The information shown in the table in Figure 3E can be stored as table information in the storage unit 1170 or memory 1109 of the artificial intelligence response output device 10010 shown in Figure 1B. Client applications can then utilize this information. The information shown in the table in Figure 3E may also be referred to as setting information related to the characteristics of the character's conversation.
[0236] Conversation examples that apply the control of standard phrases and preset instructions for the table information shown as a table in Figure 3E, which incorporates these improved settings, will be explained using Figures 3F and 3G. For clarity, the conversation examples in Figures 3F and 3G are shown as examples of conversations that apply the control of standard phrases and preset instructions for the table information shown as a table in Figure 3E.
[0237] First, in Figures 3F and 3G, the right side shows examples of user instructions that the user can input. The details of how to input user instructions have already been explained, so a repeated explanation will be omitted. On the left side of Figures 3F and 3G, examples of output from the artificial intelligence response output device 10010 of the artificial intelligence response output system are shown. Both the user instruction input on the right and the output from the artificial intelligence response output device 10010 on the left are shown in chronological order from top to bottom.
[0238] Here, the example output from the artificial intelligence response output device 10010 on the left first shows a standard phrase response 1 (greeting). This is an example of outputting a standard phrase greeting from the artificial intelligence response output device 10010 using the standard phrase output control explained in Figure 3C, etc. This standard phrase output control can be performed by the client application 8010. Next, when a user instruction sentence 1 is input to the artificial intelligence response output device 10010, the client application 8010 acquires it and sends the user instruction sentence 1 and control information to the large-scale language model. The large-scale language model that receives the user instruction sentence 1 performs inference at the timing of the star mark 9001 and generates an LLM response 1. This large-scale language model may also be the large-scale language model of the large-scale language model server 20001. In this case, the input and output of information to and from this large-scale language model is controlled by the server LLM application 8020. Alternatively, this large-scale language model may also be the large-scale language model of the local LLM processing unit 10028. In this case, the input and output of information to the large-scale language model is controlled by the local LLM application 8015. The LLM response 1 generated by the large-scale language model is sent to the client application 8010 by these LLM applications and output as an artificial intelligence response from the artificial intelligence response output device 10010. If the artificial intelligence response output device 10010 is a character conversation device, the artificial intelligence response output is recognized by the user as a response from the displayed character.
[0239] Next, an example is shown in which a user instruction 2 is input to the artificial intelligence response output device 10010 in response to the LLM response 1. Here, an example is shown in which the client application 8010, which has acquired the user instruction 2, outputs a standard response 2 (acknowledgment). This is an example in which the artificial intelligence response output device 10010 outputs a standard acknowledgment using the standard output control described in Figure 3C and other figures. Multiple types of standard acknowledgments can be prepared in advance, and the system can be controlled to output them randomly with an interval of a predetermined period or longer to avoid an unnaturally high frequency. Meanwhile, while the client application 8010 is outputting the standard response, it transmits the user instruction 2 and control information to the LLM application (server LLM application 8020 or local LLM application 8015) as described above. The LLM application that has received user instruction 2 controls the large-scale language model (the large-scale language model of the large-scale language model server 20001 or the large-scale language model of the local LLM processing unit 10028) to perform inference at the timing of the star mark 9002. The large-scale language model's inference generates an LLM response 2. The LLM response 2 generated by the large-scale language model is sent to the client application 8010 by these LLM applications and output as an artificial intelligence response from the artificial intelligence response output device 10010. If the artificial intelligence response output device 10010 is a character conversation device, the artificial intelligence response output is recognized by the user as a response from the displayed character.
[0240] Through the above series of processes, a conversation takes place between the artificial intelligence response output system and the user. Figures 3F and 3G show an example of a series of conversations between the user and the AI regarding Haneda Airport in Japan, with greetings and acknowledgments mixed in. The series of AI response outputs from the AI response output device 10010 are perceived by the user as a series of responses with a certain degree of consistency. However, these series of responses are composed of response sentences generated by a large-scale language model whose information input and output are controlled by the LLM application and whose output is controlled by the client application's control of fixed phrases. In other words, the collaboration between the client application and the LLM application results in AI response outputs that are more suitable for the user.
[0241] Here, we will explain the details of the control using the table information shown in Figure 3E for each of the conversation examples in Figure 3F and Figure 3G.
[0242] First, Figure 3F is an example of a conversation in which the control of the standard phrases and preset instruction phrases from the table information shown in Figure 3E is applied. Figure 3F shows an example of a conversation in which the client application controls the output of an artificial intelligence response from the artificial intelligence response output device 10010 as a response from the character Necco. The client application reads the information of the character Necco's standard greeting phrases and standard interjection phrases from the table information shown in Figure 3E and uses it for outputting standard phrases for the artificial intelligence response. Then, as shown in the conversation example in Figure 3F, standard phrase response 1 uses the character Necco's standard greeting phrase, with the first-person pronoun being "boku" and ending with "nyaa". Also, standard phrase response 2 uses the character Necco's standard interjection phrase, ending with "nyaa". Here, in the table in Figure 3E, the settings for catchphrases and interjections based on the LLM application preset instructions include characteristic expressions (keywords) that overlap with characteristic expressions (keywords) that represent the personality of character Necco, namely "meow" and the first-person pronoun "boku" which are included in the settings for sentence endings and interjections. As a result, in the conversation example in Figure 3F, LLM response 1 ends with "meow". In LLM response 2, the sentence ending is "meow" and the first-person pronoun is "boku". In addition, LLM response 2 includes the interjection "Hmm meow" which was generated by the large-scale language model according to the LLM application preset instructions. The ending of this interjection is also "meow". Thus, in a series of conversations, whether the output from the artificial intelligence response output device 10010 as the response of character Necco is a fixed phrase output by the client application or a response output by the large-scale language model controlled by the LLM application, a common characteristic expression (keyword) is consistently used. As a result, users can get a more consistent impression of the character's personality from the content output from the artificial intelligence response output device 10010 as the response of the character Necco.
[0243] Similarly, Figure 3G is another example of a conversation in which the control of the standard phrases and preset instruction phrases from the table information shown in Figure 3E is applied. Figure 3G shows an example of a conversation in which the client application controls the output of an artificial intelligence response from the artificial intelligence response output device 10010 as a response from the character Airia. The client application reads the information of the character Airia's standard greeting phrases and standard interjection phrases from the table information shown in Figure 3E and uses it for outputting standard phrases for the artificial intelligence response. Then, as shown in the conversation example in Figure 3G, standard phrase response 1 uses the character Airia's standard greeting phrase, with the first-person pronoun being "watakushi" and the ending being "wa". Also, standard phrase response 2 uses the character Airia's standard interjection phrase, with the ending being "wa". Here, in the table in Figure 3E, the settings for catchphrases and interjections based on the LLM application preset instructions include characteristic expressions (keywords) that overlap with characteristic expressions (keywords) that represent the personality of character Airia, namely "wa" and the first-person pronoun "watakushi" which are included in the settings for sentence endings and interjections. As a result, in the conversation example in Figure 3G, the sentence ending in LLM response 1 is "wa". Also, in LLM response 2, the sentence ending is "wa" and the first-person pronoun is "watakushi". Furthermore, LLM response 2 includes the interjection "Uuun desu wa" which was generated by the large-scale language model according to the LLM application preset instructions. The sentence ending of this interjection is also "wa". Thus, in a series of conversations, whether the output from the artificial intelligence response output device 10010 as the response of character Airia is a fixed sentence output by the client application or a response output by the large-scale language model controlled by the LLM application, a common characteristic expression (keyword) is consistently used. As a result, users can get a more consistent impression of the character's personality from the content output from the artificial intelligence response output device 10010 as the response of character Airia.
[0244] As described above, in the artificial intelligence response output system explained using Figures 3D to 3G, the client application and the LLM application cooperate to perform a series of artificial intelligence response outputs. Furthermore, in the artificial intelligence response output system explained using Figures 3D to 3G, the system controls the output of standardized texts by the client application and the response output of the large-scale language model controlled by the LLM application to consistently use common characteristic expressions (keywords) for each character. This allows the user to get a more consistent impression of the character's personality in the content output from the artificial intelligence response output system. In addition, in this system, the artificial intelligence response output device 10010 can control the output of standardized texts by the client application and the response output of the large-scale language model controlled by the LLM application to consistently use common characteristic expressions (keywords) for each character by having the client application it executes send and receive control information with the LLM application. This control by the client application executed by the artificial intelligence response output device 10010 allows the user to get a more consistent impression of the character's personality in the content output from the artificial intelligence response output system.
[0245] In the conversation examples shown in Figures 3F and 3G, the output of the standard phrase response 2, which is a set phrase for acknowledgment, begins after receiving the user instruction sentence 2, but before the LLM response output 2, which is the response output of the large-scale language model controlled by the LLM application. Since the inference of the large-scale language model takes a certain amount of time, the client application is controlled to insert the standard phrase for acknowledgment at this timing in order to prevent an unnatural gap in the response output from the artificial intelligence response output device 10010 to the user. In other words, the time immediately before the start of the inference of the large-scale language model is one of the suitable timings for outputting a standard phrase response.
[0246] Furthermore, the term "character" in each embodiment of the present invention described above includes the concepts of an AI assistant and an avatar.
[0247] The table information described in Figure 3E above contains various settings related to the characteristics of character conversations, but the output medium for character conversations is mainly natural language. Therefore, if the language output by the character conversations is different, the various settings information related to the characteristics of these character conversations also needs to be different. Accordingly, if the artificial intelligence response output device 10010 can switch between multiple languages as the language output by the character conversations, it is sufficient to store various settings information related to the characteristics of character conversations, such as the table information described in Figure 3E, corresponding to each of the multiple languages, in the storage unit 1170 or memory 1109. When the artificial intelligence response output device 10010 switches the language output by the character conversations, it is desirable to also switch to the table information corresponding to that language.
[0248] <Example 4> Example 4 will now be described. In this example, the differences from Examples 1 to 3 will be explained, and the same configurations as in those examples will not be repeated. Example 4 realizes a function in an artificial intelligence response output device such as an aerial floating image display device that uses artificial intelligence, i.e., AI, for example, a multimodal LLM, to set, generate, and output a character / avatar that reflects the desired features specified by the user. The artificial intelligence response output device will be referred to as the AI response output device below.
[0249] [Overview] Figure 4 shows the overall system configuration and overview of Embodiment 4, including the response output device 10010, illustrating the flow from user operation to character display. The response output device 10010 in Embodiment 4 is a display device, for example, an aerial levitation image display device 185 (Figure 1C). In addition to the response output device 10010 in Embodiment 4, various other display devices such as those shown in Figure 1C can also be applied. The user 230 operates and uses the response output device 10010. The response output device 10010 is connected to an LLM application 8020 (Figure 1E), such as a server. Alternatively, it may be connected to a local LLM application 8015. Furthermore, the user 230's mobile information processing terminal 20010 may be connected to the response output device 10010.
[0250] In Embodiment 4, the response output device 10010, such as the aerial levitation image display device 185, displays and outputs a character, or in other words, a character image 19051, through the exchange of instructions and responses with the AI, LLM (LLM application 8020). In Embodiment 4, such a character 19051 is set, generated, and displayed based on the user 230's operation input to the interface of the response output device 10010, or in other words, the user interface, AI interface, etc. The user 230 specifies and inputs the characteristics of the character to be set, generated, and displayed (referred to as characteristic information / character input information 43) to the interface of the response output device 10010. The response output device 10010 uses the characteristic information / character input information 43 to create and transmit an instruction message 41 to the LLM.
[0251] The LLM (LLM application 8020) generates character data that reflects the specified features through inference according to the instruction statement 41. Specifically, the LLM generates a dataset including a 3D model and rendered video data related to the 3D character / avatar. The LLM returns the generated character data as a response. The response output device 10010 performs processing in response to output the character to the screen and interface of the display unit 10011. The character output includes at least image display. As a result, the appearance, movements, and speech of the character that reflect the features specified by the user 230 are output.
[0252] In Figure 4, the specific processing details and flow are as follows. Explanation 4001 (1) describes user operation input and feature input. First, user 230 inputs the features of the character that they want to display and output to the response output device 10010, i.e., feature information / feature input information 43, to the interface of the response output device 10010. In other words, user 230 performs operation input to specify the features of the desired character.
[0253] The characteristics of this character can be input via touch operation on the display unit 10011 of the response output device 10010, for example, via a touch panel, on an aerial levitation image, or via aerial operation. This input can also be done in various ways, such as text input, image input, or voice input. Alternatively, this input may be done by entering information / values into each characteristic item of a pre-prepared sheet with a predetermined format, such as the character sheet described later.
[0254] Alternatively, user 230 may input characteristic information on the mobile terminal 20010 and transmit the characteristic information from the mobile terminal 20010 to the response output device 10010.
[0255] Alternatively, the above feature input may be performed by reading image data containing a picture of a character desired by the user 230, or characters representing the character's characteristics. For example, the user 230 may take a picture of a desired character using the camera (imaging unit 1180) provided in the response output device 10010 and input it as an image. The user 230 may also take a picture of a desired image with the mobile terminal 20010 and transmit the captured image to the response output device 10010 via communication from the mobile terminal 20010 for input.
[0256] The image representing the character's features may be created by user 230 by drawing it themselves, or licensed image data (digital illustrations, photographs, etc.) may be reused.
[0257] When performing feature input using text input or voice input, specific examples include the following:
[0258] Example 1: User 230 inputs a request to the interface in text or voice (natural language), such as: "Create a character with the following characteristics!!" The "..." part is the text / voice representing the characteristics desired by User 230. For example, "A girl named Koto-chan with long, red, tufted hair and wearing a dress with a ribbon."
[0259] Section (2) of Explanation 4002 describes the instruction statement 41. The response output device 10010 (particularly the client application 8010 in Figure 1E) uses the feature input information 43 to create an instruction statement (prompt) 41 for the AI (LLM). In this case, the response output device 10010 ensures that the instruction statement 41 includes feature data such as natural language (text), image data, and audio data, which correspond to the feature input information 43 provided by the user 230 using the method and manner of (1). In the illustrated example, the instruction statement 41 includes natural language (text / audio) and feature data 4012 such as character sheet / image data. The response output device 10010 may create an instruction statement 41 that includes the feature input information 43 by using the feature input information 43 as is, or it may create feature data 4012 for the instruction statement 41 by processing the feature input information 43 and then create an instruction statement 41 that includes the feature data 4012.
[0260] An example of natural language for instruction 41 is shown in Example 1. Another example (Example 2) is, "This image summarizes the characteristics of a character. Please return a character that reflects the characteristics depicted in the image." As an example of including information representing the characteristics of a character in instruction 41, the characteristic data 4012 may be, for example, data from a character sheet in which the user 230 has entered characteristics on the screen (GUI described later), or it may be image or photograph data in which pictures representing the characteristics are drawn.
[0261] The LLM performs inference processing according to instruction 41 and generates character data as an inference result. Explanation 4003 (3) is an explanation of the response 42 from the AI (LLM). The response 42 has a dataset that includes a 3D model of the character that reflects the features input / instructed in (1) and (2), or a rendered image (character image) 4013 generated by rendering processing from the 3D model, and audio data of conversations and speeches made by the character. The rendered image (character image) 4013 may be a still image or a video. The character image 4013 may be accompanied by audio data, or it may be just video data.
[0262] Section (4) of Description 4004 describes the processing of response and output. The response output device 10010 performs character output processing as processing in response to response 42. The processing in response to response 42 may also be called response response processing. Based on the data and information in (3) generated by the LLM, the response output device 10010 displays the character 19051, that is, the character image 4014, on the screen of the display unit 10011.
[0263] When LLM generates character images, the response output device 10010 simply displays and outputs the character image 4013 obtained from response 42 as character image 4014. When the response output device 10010 generates character images, it performs rendering processing based on a dataset such as a 3D model obtained from response 42 to generate, display, and output character images (rendered images) 4014.
[0264] The appearance of character 19051, or character video 4014, displayed and output in (4), will reflect the characteristics input and instructed in (1) and (2). The output character 19051 will perform actions and speak according to the instructed characteristics. For example, if "active" is specified as a characteristic, character video 4014 will be displayed showing character 19051 moving around. In addition, character 19051 will speak greetings and self-introductions, etc., based on the name and other settings corresponding to the specified characteristics.
[0265] In this way, when user 230 inputs desired features, a character reflecting those features is displayed and output. User 230 can then enjoy the character. Up to (4) above, a character reflecting the desired features is output only once, but the data of the character set, generated, and output in this way can be saved and used and reused by user 230.
[0266] One example of how the generated character can be used is to have a conversation between the user 230 and the character, as shown in the example in Figure 3F above.
[0267] [Examples of Use] The characters generated by the above-mentioned functions can be used in a variety of ways. For example, the following use cases can be envisioned.
[0268] (1) 3D avatar generation for users such as virtual YouTubers (VTubers), streamers, and metaverse users. The above characters correspond to 3D avatars. Users 230 can easily create 3D avatars. Furthermore, users 230 have the advantage of being able to easily check and change the degree to which their own personality and style are reflected as characteristics.
[0269] (2) Simulation for corporate recruitment departments. Based on information such as resumes, virtual avatars of candidates can be created and interview simulations can be conducted. The above characters correspond to the virtual avatars. The scenario can be finalized in advance, including what questions will be asked.
[0270] (3) Simulation for fashion designers and beauticians. It allows users to create avatars (models) wearing designed clothing, and avatars (models) with different hairstyles and makeup. It has the same advantages as (1). It can also be applied to games for children.
[0271] (4) For educators. It allows you to build a virtual environment and avatars that simulate the educational setting, enabling you to conduct lesson simulations and role-playing. You can easily generate characters that reflect the needs of learners and create a variety of situations.
[0272] [Example of an Interface] Figure 5 shows an example of a user interface (UI) when user 230 inputs character features, i.e., feature information / feature input information 43, in Figure 4 (1). The response output device 10010 provides such a UI using hardware and software such as the display unit 10011. Figure 5 shows the case where features are input in natural language, in particular the case where text representing features is input / selected for each item on the character sheet. Figure 5 shows an example of a graphical user interface (GUI) on the screen of the display unit 10011 in that case.
[0273] The GUI / screen in Figure 5 includes a mode field 501, a confirmation button 502, an instruction input field 503, a feature item list 504, a scroll bar 505, etc. If all information cannot be displayed on the screen, the screen display content can be changed by operating the scroll bar 505 or page buttons, etc. In this embodiment, the feature item list 504 corresponds to a character sheet, or in other words, a feature sheet.
[0274] The mode field 501 displays the current mode. The mode field 501 includes an indicator display 501a that indicates the current mode / status. For example, the text "List" on the indicator display 501a indicates that the feature item list 504 is displayed in list mode. The confirmation button 502 is pressed by touch or the like when the user 230 confirms the input of features and executes / completes the instruction input to the LLM. When the user 230 inputs features by operating the input field 504b of the feature item list 504, the indicator display 501a becomes, for example, "Input," indicating that the text / character input mode for features is active.
[0275] The instruction input field 503 is a field where natural language (text) can be directly entered to form the instruction 41 for the LLM. The user 230 may directly enter the instruction in natural language (text) in the instruction input field 503. Alternatively, the response output device 10010 may prepare and display a default natural language (text) in the instruction input field 503, allowing selection input (Figure 19). In this embodiment, the natural language (text) for the instruction 41 is entered in this field as, "Please return a character that reflects the features." If a default natural language is prepared, the effort required for user 230 to input can be reduced.
[0276] The instruction input field 503 can normally be used to input text for the instruction 41 to the LLM (especially user-inputted instructions), that is, text instructing the LLM to generate a character that reflects the desired characteristics of the user 230. In addition, if any special conditions (e.g., additional instructions) are to be imposed on the LLM, text representing those conditions can be entered in the instruction input field 503. Examples of such conditions include, "Please fill in the blank characteristics as appropriate," and "Please summarize the circumstances and past that led to the entered personality within 100 characters." For example, if there is a blank in the input field 504b of the characteristic item 504a, the LLM will fill in the characteristics based on the former condition. If there is a personality input value in the input field 504b of the "Personality" characteristic item 504a, the LLM will fill in the circumstances and past that led to that personality based on the latter condition. In this embodiment, the input text (natural language) in the instruction input field 503 and the input information in the feature item list 504 (character sheet) are related, and these are handled together to create the instruction 41.
[0277] In addition, if user 230 has any modifications to the characteristics of a character that has already been generated and output (as described later), they can enter instructions for those modifications in the instruction input field 503. For example, "Please make the hair color a little lighter" or "Please remove the ribbon." The character can also be regenerated by modifying the input value in the input field 504b for characteristic item 504a and giving instructions again, but it can also be regenerated by giving additional modification instructions in the instruction input field 503 mentioned above.
[0278] Furthermore, the conditions for LLM may be described in advance using system setting instructions, or in other words, control information, separate from the instruction statements 41 for each instance.
[0279] The feature item list 504 contains a list of multiple feature items 504a that the system has pre-set. In other words, feature items 504a are the feature items that make up a character sheet. Feature items 504a are a list of various features that can be attached to a character, in other words, the types of features. Examples of feature items 504a include "Name," "Gender," "Age," "Blood Type," "Birthday," "Height," "Weight," "First-person pronoun," "Hair Color," "Image Color," "Personality," "Likes," "Dislikes," and "Other." Each feature item 504a has an input field 504b. Feature item 504a corresponds to a variable, and the input field 504b corresponds to a value. By operating the input field 504b (touch operation, etc.), the user can input or select information representing the feature. For each feature item 504a, it may also be possible to select from predefined options.
[0280] For example, in the "Name" input field 504b, the user can enter a name for the character as text. The "Name" input field 504b is initially blank when the user 230 has not entered anything, and in response to the input operation, the input text, such as "Koto-chan," is displayed. In the "Gender" input field 504b, the user can select the character's gender from the options. In the "Age" input field 504b, the user can enter the character's age. Units such as "years," "age," "months," and "days" may be displayed in advance for each input field 504b. Furthermore, when the user 230 performs a selection operation (such as a touch operation) on a desired feature item 504a or input field 504b, it may be highlighted as an input target, for example, as shown in input field 504c, by changing its color. When a selection operation is performed on this input field 504c, it may transition to a predetermined input mode (described later) corresponding to the input format of the input field 504c. This may be a screen transition or a pop-up display.
[0281] The user 230 inputs the necessary feature information on this screen and operates the confirmation button 502. It is sufficient to input features into at least some of the feature items 504a, if not all of them. The response output device 10010 then creates an instruction statement 41 using the input information on this screen as feature input information 43 and the corresponding feature data 4012. Other GUI elements, such as buttons to clear / cancel input, may also be provided.
[0282] Feature item 504a may be made customizable by the user 230, such as by adding items. Figure 18 shows an example of a GUI for a function that allows the user 230 to customize feature item 504a. At the end of the feature item list 504, there is an "Add Feature Item" button 504e. If the user 230 wants to add a feature item 504a that is not in the feature item list 504, the user 230 operates the "Add Feature Item" button 504e. In response to this operation, the response output device 10010 adds and displays a new feature item 504a and an input field 504b to the feature item list 504. The user 230 inputs and sets the name of the new feature item 504a using a predetermined operation (such as a character input mode). For example, when adding a feature item 504a for "suffix", "suffix" is input and set. Then, the user 230 can input, for example, "nyan" as the value of the "suffix" feature in the corresponding input field 504b.
[0283] In other GUI examples, existing feature items 504a in the feature item list 504 may be deleted, renamed, or reordered by the user 230 through predetermined operations.
[0284] Figure 19 shows another example of the configuration of the instruction input field 503. It has the following functions to reduce the effort required to input instructions in natural language. First, even if there is no user input, the instruction input field 503 initially displays a default natural language, for example, "Please return a character that reflects the features." The user 230 may select other default natural language by touching the list box or the like in the instruction input field 503.
[0285] Furthermore, the response output device 10010 saves previously entered text as history, allowing user 230 to select and reuse it from the history if they wish to re-enter similar text. In this embodiment, when user 230 touches the list box or the like in the instruction input field 503, a history of characteristic information (natural language / text representing it) previously entered by user 230 is displayed as options. User 230 can select from these characteristic information, and the selected characteristic information is displayed in the instruction input field 503. User 230 may reuse the characteristic information as is, or may edit it as appropriate to create different characteristic information. In the illustrated example, "A girl with red, long hair. She uses 'watashi' as her first-person pronoun, and ends her sentences with 'nyan'." is selected from the options as characteristic information entered on December 1, 2024.
[0286] Note that Figure 5 shows a case where feature input in the instruction input field 503 and feature input in the feature item list 504 are related and used together, but these may also be separate functions.
[0287] Figure 6 shows an example of another UI, specifically a screen for text / character input mode. The user transitions from the screen in Figure 5 to the screen in Figure 6. When user 230 touches the instruction input field 503 or the input field 504b for feature item 504a in Figure 5, the response output device 10010 transitions to a screen like the one in Figure 6. The input mode screen in Figure 6 has a GUI that allows input of text including characters, natural language, etc. The indicator display 501a shows "Input" to indicate the input mode. The screen in Figure 6 has a "Back" button 506. When the "Back" button 506 is pressed, the user returns to the screen in Figure 5.
[0288] In the input mode screen shown in Figure 6, an input instruction 507 is displayed. The input instruction 507 is a text input instruction to the user 230, and is, for example, an input instruction for a feature corresponding to the feature item 504a. For example, in the input screen corresponding to the operation of the input field 504b for the "first person" feature item 504a, the input instruction 507 will be displayed as "Please enter the first person." In another example, in the input screen corresponding to the operation of the instruction text input field 503, the input instruction 507 will be displayed as "Please enter the instruction text."
[0289] In the input display area 508 below the input instruction 507, the user 230 can input characters, natural language, or text by touch operation, and the entered text is displayed. Below the input display area 508 is a software keyboard 509 for character input. By operating the software keyboard 509 (by touching the keys, etc.), the user 230 can input and display characters in the input display area 508. The keys of the software keyboard 509 may be highlighted, such as by changing color, when they are operated.
[0290] In the illustrated example, when user 230 inputs and sets the first-person pronoun to "watashi" using the Japanese Romanization input method, the keys are entered in the following order: w, a, t, a, s, i, and the confirm key is pressed. In the illustrated example, the input is complete up to the key [s], but before the key [i] is entered, i.e., the input is still in progress. The input display field 508 shows "watashi" (underlined to indicate conversion and unconfirmed), which is a combination of the Japanese conversion "watashi" (w) + [a] and [t] + [a] and the original [s].
[0291] The input of features is not limited to the example above. When using voice input, the user 230 can input voice representing the features through the microphone 1139. The response output device 10010 may convert the input voice through the microphone 1139 into text by speech recognition processing. The feature data attached to the instruction statement 41 may be either text or voice data. When the feature data of the instruction statement 41 is voice data, the LLM is a multimodal LLM capable of recognizing and processing voice.
[0292] [Example of Image Input] Figure 7 shows an example of a GUI screen for inputting an image (image data) that represents features, as an example of another UI. As shown on the left side of Figure 7, user 230 prepares a piece of paper 701 with the features of their desired character drawn on it. In this embodiment, a handwritten illustration (picture) 702 representing the character is drawn on the paper 701. Also, the user 230 has written a handwritten memo 703 on the paper 701 that describes the features. The illustration 702 is, for example, a front view of a girl character, focusing on the face and the part from the shoulders up. The memo 703 contains natural language that describes the features, for example, "Name Koto-chan", "Girl", "Cheerful, smiling", "Ends sentences with "nyan". The memo 703 also contains text describing the features, with leader lines pointing to parts of the illustration 702. For example, "Long hair" and "Floating hair" are added for the hair, and "Ribbon" is added for the clothing around the neck. Thus, on paper 701, the characteristics of the character desired by user 230 are described, including physical characteristics, name, gender, personality, speech patterns, etc.
[0293] As mentioned above, image data such as digital illustrations or photographs may be provided, not limited to illustrations 702 on paper 701. In that case, the image data should be input to the interface of the response output device 10010.
[0294] User 230 takes a picture of the illustration 702 on paper 701 using the camera (imaging unit 1180) of the response output device 10010 and has it read as image data. Alternatively, user 230 may input image data such as a digital illustration that they have prepared into the response output device 10010. Alternatively, image data may be input using a mobile terminal 20010 and sent to the response output device 10010 via communication.
[0295] The right side of Figure 7 shows an example of the GUI screen of the display unit 10011. This embodiment is an example of a GUI used when a user 230 inputs features by loading a piece of paper 701 on which an illustration 702 of a character they want to display on the response output device 10010 is drawn as image data.
[0296] The screen 700 in Figure 7 includes a mode field 711, a confirmation button 712, an instruction input field 713, an image data display field 714, etc. The indicator display in the mode field 711 shows, for example, "Image Loading". "Image Loading" indicates that the current mode is a mode for inputting features by loading image data. In the instruction input field 713, natural language (text) representing an instruction is entered and displayed, for example, "This image summarizes the features of a character. Please return a character that reflects the features depicted in the image."
[0297] In the image data display area 714, the image (image data) 715 read based on the paper 701 using the method described above is displayed.
[0298] In the instruction input field 713, an instruction is typically entered, particularly a user-input instruction, to generate and output a character that reflects the features described in image 715. Also, as with the text input described above, additional instructions or conditions may be entered in the instruction input field 713. For example, "Please appropriately fill in the part of the illustration below the shoulders. However, please make sure the character is wearing long pants."
[0299] Instruction 41 may be a system setting instruction, separate from the user input instruction.
[0300] Furthermore, the instruction input field 713 allows the user 230 to input instructions for modifications to the characteristics of a character that has already been generated and outputted. For example, "Please make the hair longer, down to the waist," or "Please add eyelashes." More details will be provided later.
[0301] To confirm feature input using the image 715 displayed in the image data display field 714 and the text entered in the instruction input field 713, the user 230 operates the confirmation button 712. This creates the instruction 41. The input image can also be changed using a clear / cancel button (not shown), etc.
[0302] As an alternative, this screen may be modified to allow the creation and editing of an image 715 for use as feature data. An image editing field may be provided in place of, or in addition to, the image data display field 714. In the image editing field, it may be possible to edit a part of the image 715, for example, by deleting or adding lines. The image editing field may also be a function of a general illustration creation application. The user 230 can draw illustrations or text on the image editing field using touch operations, etc., and save it as an image.
[0303] A multimodal LLM capable of processing images can recognize features (appearance, name, etc.) from image data (feature data 4012) attached to the instruction statement 41 through inference, and generate character images, etc., that reflect those features.
[0304] If the generated character includes voice data, a possible modification is to allow specifying the characteristics of the character's speech and conversation, such as vocal range and tone. The character's voice characteristics can be specified using text, images, or voice input. For example, an instruction such as "Make it a high-pitched voice" would suffice. Alternatively, licensed voice data with the desired tone can be input as characteristic data.
[0305] [Character Display Examples] Figure 8 shows examples of character displays that can be generated and output by LLM. Figure 8 shows an example of character display when features are input using one of the aforementioned input methods and instructions are given to generate and display a character that reflects those features. The table in Figure 8 summarizes several character display examples. The first row of the table shows the features (feature items), the second row shows the instructions, i.e., the text representing the features, and the third row shows the character display examples in correspondence.
[0306] For example, if you input or specify "female" for the "gender" characteristic, a character that appears female will be generated and displayed. If you input or specify "male," a character that appears male will be generated and displayed.
[0307] For example, if you input or instruct "likes singing" in the "personality" or "likes" characteristic, a character performing a singing motion will be generated and displayed. During this singing motion, not only the character's image but also text or sounds representing songs or music may be output, or dialogue suggesting that the character enjoys singing may be output. For example, the line "I really like to sing" may be output.
[0308] For example, if the characteristic "Personality" or "Likes" is entered as "Active / Likes running," a character running with a smile will be generated and displayed. The output may include video of the character moving back and forth within the screen area. Additionally, the output may include dialogue that suggests the character enjoys running. For example, "I really love running!", "My dream is to become a track and field athlete!", or "I always practice long-distance running on the field after school!"
[0309] For example, if you input or specify "off-shoulder" in the "clothing" feature, a character wearing off-shoulder clothing will be generated and displayed. If you input or specify "dress with ribbon," a character wearing a dress with a ribbon will be generated and displayed.
[0310] [Character Display Screen] Figure 9 shows an example of character display output on the screen of the display unit 10011, corresponding to step (4) in Figure 4. An example of input corresponds to the case where a feature is read from image data and an instruction is input, as shown in Figure 7. In the screen of Figure 9, the indicator display in the mode field 901 is, for example, "Display," and this "Display" indicates that it is the character display mode, a screen in which a character generated by AI is displayed. The screen of Figure 9 also has a re-instruction button 902, a re-load button 903, etc.
[0311] In the main area of the screen in Figure 9, a character (character image) 900 generated by LLM, reflecting the input and instructed features, is displayed. This character 900 may be a video that performs predetermined movements on the screen, or it may be accompanied by the output of voice or text such as conversation. In the illustrated example, character 900 is a full-body image of a girl character and is accompanied by the output of speech voice 900b. In the illustrated example, the input features instructed are the name "Koto-chan", first-person pronoun "watashi", and sentence ending "nyan". Therefore, the output image of character 900 will output voice 900b that reflects the specified features (name, first-person pronoun, sentence ending), for example, as a self-introduction voice 900b, such as "Hello nyan!!" and "My name is Koto-chan nyan!!".
[0312] For content and items for which user 230 did not provide feature input instructions, these are supplemented by inference on the LLM side. For example, in the input image 715, the part below the shoulders was missing, but in the output character video 900, the part below the shoulders (torso, limbs, clothing, etc.) is also supplemented.
[0313] Furthermore, the user 230 can change the display state of the character 900 on the screen by manipulating it with their fingers or other means. For example, the user 230 can perform a swipe operation by placing their finger on the screen and sliding it in any direction. As a result, the character 900 rotates in the direction the finger is slid. Since the character 900 is an image generated by rendering based on a 3D model, it is possible to generate and display an image from a desired viewing direction (corresponding rotation angle). This allows the user 230 to examine the details of the displayed character 900 from multiple angles, such as the front, back, and top, as desired viewing directions.
[0314] Furthermore, for example, user 230 can enlarge or reduce the display of character 900 by performing a pinch-in operation, where they place two fingers on the screen and narrow the distance between them, or by performing a pinch-out operation, where they spread their fingers apart and widen the distance between them. This allows user 230 to examine the details of the displayed character 900, or to view the character 900 as if it were a picture taken from further away.
[0315] After user 230 has viewed the character generated by LLM on a screen such as Figure 9, if they like it, they can save the character's data with a name and re-output and reuse it later. The character data can be stored on the response output device 10010 or an external device (such as a server) connected to it. The character data may also be saved on user 230's mobile terminal 20010.
[0316] The response output device 10010 may create URL information for the storage location of the character data and provide it to the user 230 or an external device. For example, the response output device 10010 may provide the URL information of the storage location as a QR code. The response output device 10010 uploads and registers the character data to the storage location specified by the URL, etc. The user 230 reads the QR code with, for example, a mobile terminal 20010. Then, the URL information is obtained by decoding the QR code. The mobile terminal 20010 can access the server, etc. at that URL and obtain the character data from the server, etc. to the mobile terminal 20010. Then, the character can be output on the mobile terminal 20010. Not limited to the mobile terminal 20010, any authorized device can similarly obtain, save, output, and use the character data.
[0317] Furthermore, the character data in these cases becomes easier to apply if it is made into a dataset that includes a 3D model. That is, a device that acquires a character dataset can generate and output video data of a character / avatar with desired movements and speech by rendering the 3D model, etc.
[0318] If the re-instruction button 902 in Figure 9 is operated, the user returns to the "List" screen in Figure 5, and can instruct the system to modify the input of character characteristics. By modifying the characteristic input (characteristic information) and issuing another instruction, a new character can be generated by LLM, or in other words, regenerated. The process in this case is the same as described above.
[0319] If the reload button 903 is operated, the display of the current character 900 is reset, and while the previous feature input instructions remain the same, LLM can generate and display a new character, or in other words, a variation of the output.
[0320] Furthermore, examples of controlling the endings of sentences and first-person pronouns in character speech and conversation can be applied as described in the previous embodiment.
[0321] [Example Configuration for Generating Character Data] Figure 20 shows an example configuration in which an external device, the LLM server, i.e., the multimodal LLM server 20001 in Figure 1A, and the response output device 10010, the aerial floating image display device 185, are connected by communication. In particular, it shows an example configuration in which character data is generated by rendering a 3D model.
[0322] The LLM server 20001 includes the LLM application 8020, as well as a video processing unit 2001 and a database (DB) 2002. The LLM server 20001 stores a dataset 2003 containing a 3D model, rendered video data, audio data, etc., in memory resources such as DB 2002. The dataset 2003 may also contain motion information, virtual light source information, virtual camera information, etc. Furthermore, the dataset 2003 may store associated feature information attached to the instruction statement 41 and feature information representing the characteristics of the generated character.
[0323] The aerial floating image display device 185 includes a client application 8010, as well as an image processing unit 2011, a memory 2012, a display unit 2013, an optical system 2014, a user operation detection mechanism 2015, and the like. The image processing unit 2011 corresponds to the control unit 1110 and the image control unit 1160 described above (Figure 1B).
[0324] Regarding the generation of character video data, there are, for example, the following two cases, and either configuration is acceptable.
[0325] (1) First case: The LLM of the LLM server 20001 generates a character that reflects the characteristics based on the instruction message 41 from the aerial floating image display device 185. At that time, the video processing unit 2001 of the LLM server 20001 generates a 3D model and generates video data (rendered video data) of the character by rendering based on the 3D model. The LLM server 20001 sends a response 42 containing the video data to the aerial floating image display device 185. The aerial floating image display device 185 displays the character on the aerial floating image 195 based on the video data of the response 42.
[0326] (2) Second case: The LLM of the LLM server 20001 generates a character that reflects the characteristics based on the instruction message 41 from the aerial floating image display device 185. At that time, the video processing unit 2001 of the LLM server 20001 generates a three-dimensional model. The LLM server 20001 sends a response 42 including the three-dimensional model to the aerial floating image display device 185. The video processing unit 2011 of the aerial floating image display device 185 generates video data (rendered video data) of the character by rendering based on the three-dimensional model in the response 42. The aerial floating image display device 185 displays the character on the aerial floating image 195 based on the video data.
[0327] The aerial levitation image display device 185 stores a dataset, such as a 3D model or rendered image data, acquired from the LLM server 20001, in its memory 2012. The image processing unit 2011 drives and controls the display unit 2013 (display unit 10011 in Figure 1B, display device 193 in Figure 1E) based on the rendered image data, thereby displaying an image on the screen of the display unit 2013, for example, on the display surface of a liquid crystal display panel. The image light emitted from the display unit 2013 passes through the optical system 2014 (such as the imaging optical plate 194 in Figure 1E) to form an aerial levitation image 195 (optical image), which is a real image, at a predetermined position.
[0328] The user operation detection mechanism 2015 is a mechanism that detects aerial operations on the aerial floating image 195 as user operations, and corresponds to, for example, the aerial operation detection sensor 1351 and the imaging unit 1180 in Figure 1K. The video processing unit 2011 performs display control processing of the aerial floating image 195 based on the user operation detection information from the user operation detection mechanism 2015 and other user setting information.
[0329] [Example of Feature Modification and Re-output] This section explains how to modify the features of a displayed character, or in other words, how to re-instruct the system to regenerate and re-output the character. An example of modifying features and re-instructing the system using the re-instruction button 902 in Figure 9 is shown below.
[0330] First, let's explain how to generate a character for the first time. On the screen shown in Figure 5 or Figure 7, the user 230 inputs the characteristics of the character. The characteristics may also be entered in the instruction input field 503 or the characteristic item list 504 (character sheet) in Figure 5. An image may also be entered on the screen shown in Figure 7. Here, we will explain the case where characteristic information is entered using an image 715 on the screen shown in Figure 7. In this case, the response output device 10010 creates an instruction statement 41 that includes characteristic data 4012 (especially image data) based on the input image 715, and obtains the character data generated by LLM as a response 42.
[0331] As a result, the generated character 900 is displayed on a screen like the one in Figure 9. Suppose user 230 looks at this character 900, checks it, modifies some of its features, and wants to generate and display the character again. In this case, user 230 presses the re-instruction button 902. This returns to the "list" screen, similar to Figure 5, where the character's features are entered.
[0332] Figure 10 shows an example of returning to the "list" screen for inputting character features by pressing the re-instruction button 902 from the character display screen in Figure 9, illustrating an example of inputting features for the second time. At this time, based on the feature data 4012 that was instructed the first time and the generated and displayed character data, the corresponding feature information is automatically entered into the input field 504b of the feature item 504a on the screen.
[0333] A concrete example is as follows: The name "Koto-chan" has already been specified by user 230 from the first image 715, and is set in the character 900 data as a name recognized by LLM. Based on that data, "Koto-chan" is entered into the name input field 504b. Similarly, information such as gender "female", hair color "red", personality "cheerful, smiling", and other information such as "floating hair, long hair, ribbon, ends sentences with 'nyan'" is automatically entered and displayed.
[0334] The feature items such as "age," "blood type," "birthday," "height," "weight," "first-person pronoun," "image color," "likes," and "dislikes" were not specified by user 230 in the initial image 715. Based on the instruction text 41 that instructed the LLM to complete these feature items, the LLM automatically completes them during inference and sets them in the character data of response 42. Based on this data, the information is automatically entered into the input fields 504b for those feature items on the screen. For example, "age 18 years old" is entered.
[0335] As shown in the diagram, when the user returns to the screen for re-instruction, the input field 504b for each feature item 504a in the feature item list 504 (character sheet) is automatically populated with the feature information contained in the first image data (feature data 4012). In other words, this feature information is the feature information set for the character 900 generated by LLM in response to the features specified by user 230. User 230 does not need to re-enter the same feature information as the first time. If there are feature items that were not completed by LLM, the input field 504b for those feature items may be left blank.
[0336] User 230 checks the feature information in input field 504b of feature item 504a on the screen and, if necessary, modifies the feature information in some of the input fields 504b as desired. Modification of features can be done, for example, by re-entering the text representing the feature described in input field 504b of feature item 504a. This modification and re-entry can be done using the same operation method as, for example, the input mode in Figure 6 described above.
[0337] Figure 11, following Figure 10, shows an example of modifying some features in the feature item list 504 on the screen, and an example of the instruction statement 41 created in response to that modification. In this embodiment, the user 230 modified the "first person" feature item 504a in the input field 504b from "watashi" to "atai". As a result, the newly created instruction statement 41, i.e., the feature data 4012 (character sheet) attached to the instruction statement for re-instruction, will include a note indicating that the first person should be modified to "atai".
[0338] Furthermore, user 230 can also modify the features instructed to the AI by directly modifying the natural language entered in the instruction input field 503. For example, user 230 may want to change the "clothing" as a feature not specified in image 715 and not available in the feature item list 504. For example, user 230 may want to change the "dress with ribbon" in Figures 8 and 9 to "off-shoulder". In the natural language entered in the instruction input field 503, user 230 would input a modification instruction such as, "Please return a character that reflects the features. However, please change the clothing to 'off-shoulder'." In other words, the natural language is modified. As a result, the newly created instruction 41 will contain an instruction to change the clothing to "off-shoulder".
[0339] Furthermore, when re-instructing, if user 230 has characteristics of the generated and displayed character that they do not want to change, in other words, characteristics they want to maintain, they can specify those characteristics. That is, this embodiment also has a function to maintain some of the character's characteristics when re-instructing. For example, the characteristics to be maintained can be specified in natural language in the instruction input field 503. User 230 enters, for example, "Also, please do not change the characteristics entered in "Name," "Gender," "Hair Color," "Personality," and "Other," as well as the "Voice." As a result, the newly created instruction 41 will include instructions to maintain the specified characteristics. If you want to maintain the voice (range, tone, etc.) of the output character but change its actions, you can enter, "Please do not change the voice. Please change the actions," etc.
[0340] Furthermore, it may be cumbersome for user 230 to have to manually input instructions or feature items each time to maintain or change the various features described above. Therefore, instructions to maintain some of the character's features may be implemented using the GUI of the character sheet (list of feature items 504).
[0341] Figure 12 shows an example of providing a GUI in the feature item list 504 of the screen in Figure 11 that allows specifying the maintenance of some features. In this GUI, a checkbox 504d is provided on the right side for each feature item 504a and input field 504b. Checking the checkbox 504d (check mark) indicates that the features of the feature item 504a will be maintained when re-outputting, in other words, it is tentatively confirmed, while unchecking it indicates that changes to the features of the feature item 504a will be allowed when re-outputting.
[0342] User 230 sets the checkbox 504d of the desired feature item 504a to on / off. As a result, the newly created instruction statement 41 will include instructions to maintain or allow changes to each feature as specified by the checkbox 504d. In the illustrated example, feature items 504a such as "Name" are checked, so this instructs them to be maintained, and in the next (second) character output by a re-instruction (next instruction statement), the features such as "Name" from the first time will be maintained. Feature items 504a such as "Age" are unchecked, so this instructs them to allow changes to the feature, and in the second character, the settings such as age may be changed (though they may not be changed). Also, for the "Other" feature item 504a, several features were specified in the first time ("floating hair, long hair, ribbon, ending sentences with 'nyan'"), it is possible to specify that some of these be maintained and some be allowed to be changed. For example, as shown in the diagram, if you want to maintain "long hair" and "ending sentences with 'nyan'", one of the "Other" items will be checked, while "float hair, ribbon" will be unchecked in the other "Other" item.
[0343] Alternatively, instead of the format shown in the example above, a GUI (such as a checkbox) may be provided to instruct the user to change the features. In this case, checking the checkbox would indicate an instruction to change the features.
[0344] Figure 13 shows a specific example of generating and outputting a character again by re-instructing the system with some of the features changed as described above. Figure 13(A) shows the first character 900-1 before the feature change (same as Figure 9), and (B) shows the second character 900-2 after the feature change according to the feature modification instructions mentioned above.
[0345] As for the corrections, in the first instance, character 900-1 uses "watashi" (I) as the first-person pronoun in speech voice 1301, while in the second instance, character 900-2 uses "atai" (I) as the first-person pronoun in speech voice 1302. Also, in the first instance, character 900-1 is dressed in a "dress with a ribbon" (automatically completed due to no clothing specification), while in the second instance, character 900-2 is dressed in an "off-shoulder" top. For each feature that was specified to be maintained (name, gender, hair color, other hairstyles, ribbons, sentence endings, etc.), the output is fixed (unchanged). Conversely, some features that were not instructed to be corrected, such as pants and shoes, have been changed in the re-output.
[0346] As shown in the example above, user 230 can easily modify the desired features while maintaining them, and then re-output the character. By repeating this process several times, user 230 can create a character with the desired features.
[0347] [Exclusion Function] When using image data as feature input as shown in Figure 7 above, if there are features depicted in, for example, the image 715 (illustration 702) of paper 701 that you do not want to be reflected in the output character, you can input and instruct the system to exclude those features. In other words, you can create an instruction statement that includes instructions to not reflect some of the features in the image of the feature data in the character generation by LLM. This function is referred to as the exclusion function.
[0348] Figure 14 is an explanatory diagram regarding the exclusion function. Similar to Figure 7, it is an example in which a character drawing (illustration 702) previously created by user 230 is loaded as image data, and a character is generated and output. For example, user 230 wants to remove the "float hair" and "ribbon" depicted in the character illustration 702 in image 715 and not reflect them in the output character. In that case, user 230 enters natural language in the instruction input field 503 of Figure 14, for example, "However, remove the 'float hair' and do not add the 'ribbon'." This creates a new instruction 41 that includes those instructions. LLM generates the character without reflecting, i.e., excluding, the specified features.
[0349] Figure 14(A) shows character 1400-1 generated and output based on the image data in Figure 7 without any exclusion instructions, and Figure 14(B) shows character 1400-2 generated and output based on the same image data with exclusion instructions. The "floating hair" 1401 and "ribbon" 1402 in character 1400-1 (without exclusion instructions) are excluded in character 1400-2 (with exclusion instructions). The exclusion function makes it easier to generate the desired character when using image data for feature input.
[0350] As shown in the examples in Figures 10 to 14 above, the user 230 can confirm the character generated and output by LLM after specifying features with instruction statement 41, and then re-instruct with instruction statement 41 to maintain some features and change others, thereby re-outputting the character. Maintaining or changing features can be done by specifying them in feature item 504a or in instruction statement input field 503.
[0351] [Example of reloading] If user 230 is dissatisfied with the character (image, sound, actions, etc.) generated and output by LLM based on the features specified by user 230, it is possible to reload and re-output based on the same feature information as before. In other words, it is possible to reset the current character display and display a new character.
[0352] Figure 15 shows an example of character display when the character display screen in Figure 9 is reloaded, or in other words, reset, reloaded, or variation generated using the reload button 903.
[0353] The example in Figure 15(A) shows a screen / state displaying the character 1500A generated for the first time, based on the first feature input and instruction statement 41. If the user 230 wants to regenerate the character with the same features as this character 1500A, that is, with the feature conditions specified in the first instruction statement 41, the user presses the reload button 903. At this time, the reload button 903 is highlighted, for example, by changing color.
[0354] In response to this operation, the state / screen transitions from (A) to (B) and then to (C) in Figure 15. The response output device 10010 creates an instruction message 41 for reloading and regenerating in response to this operation and sends it to the LLM. The feature information contained in this second instruction message 41 is the same as the feature information when the instruction message 41 instructed the generation of the first character 1500A in Figure 15 (A). Alternatively, this second instruction message 41 may contain the content, "Please return a different character with the same features as those included in the previous instruction."
[0355] Figure 15(B) shows an intermediate display screen / state indicating that the page is being reloaded / refreshed / character display is being updated. This screen displays a reload icon 1502 to indicate that the page is being reloaded. The reload icon 1502 is, for example, a circular arrow. In this case, the indicator display shows, for example, "Reloading".
[0356] Upon receiving the second instruction 41, the LLM regenerates the character, reflecting the same characteristics as the first instruction 41 in Figure 15 (A), in other words, generates a variation of the character, and returns response 42. An intermediate display screen is shown depending on the time required for the reloading process. The generation of character variations by the LLM can be performed using random number generation techniques, such as seed values.
[0357] Figure 15(C) shows the screen / state where, after reloading and based on response 42, the regenerated character 1500B is displayed after the preparation for displaying a new character is complete. Character 1500B is a regenerated and re-outputted character. Character 1500B is generated as a character / variation with features that do not deviate from the features specified before reloading (first time). In this embodiment, compared to character 1500A, character 1500B maintains and displays previously specified features such as "floating hair," "long hair," and "ribbon" (Figure 7, etc.). Furthermore, for features not specified by user 230, such as clothing, the clothing portion that was held and displayed by the LLM side is updated to something different from the first "dress with ribbon," such as "off-shoulder."
[0358] Reloading is possible in the same way for the third time and beyond. In this way, user 230 can obtain variations of characters that meet the desired characteristics with simple operations.
[0359] [Example of Multiple Outputs] The above example describes a case where one character is generated and displayed with one response 42 for one instruction 41, but it is not limited to this. It is also possible to generate and display multiple characters with one response 42 for one instruction 41.
[0360] Figure 16 shows an example of a screen display where multiple characters are generated and displayed simultaneously by specifying a certain feature once. These multiple characters correspond to variations that satisfy the feature specified by user 230.
[0361] As described above, it is possible to repeatedly reload and update the display, allowing users 230 to view multiple characters / variations that satisfy the desired characteristics on a timeline. However, the operation may be perceived as cumbersome by the user 230. Therefore, the system has a function to simultaneously display multiple characters / variations that satisfy the specified characteristics on the screen. The instruction 41 given from the response output device 10010 to the LLM may be an instruction to generate multiple characters / variations based on the same characteristic information. Based on the instruction 41, the LLM generates multiple characters that reflect the characteristics input and specified by the user 230. Based on the response 42, the response output device 10010 displays these multiple characters simultaneously on the screen. The user 230 can view these multiple characters on the screen and select and adopt the desired character. An example of the instruction 41 for generating multiple characters may be, "Please generate multiple characters that reflect the characteristics of..."
[0362] Figure 16(A) shows an example of displaying multiple generated characters on the screen of the display unit 10011. On this screen (character selection screen), multiple characters such as characters 1600A, 1600B, 1600C, and 1600D are displayed in parallel. These characters are variations that have been generated to share the same characteristics specified in instruction 41. Each character has variations in appearance, voice, and behavior, within the limits of not deviating from the input / instructed characteristics. The parallel display of multiple characters may also be represented by displaying thumbnails or snapshots of each character.
[0363] If each displayed character does not fit on the screen, the excess portion can be displayed by using a scroll bar or similar operation. The indicator in the mode field will display a word such as "Select" to indicate that it is a screen / mode in which the user 230 can choose their favorite character from among the multiple characters that have been generated. In addition, this screen may also display a guide / message to the user 230, such as "Please select a character," to indicate that it is a screen / mode in which the user 230 can choose their favorite character from among the multiple characters that have been generated.
[0364] Also, as mentioned above, pressing the reload button 903 may reset all currently displayed characters and allow for the generation and display of new characters. Alternatively, pressing the same button may allow for the generation and display of new characters one by one in addition to the currently displayed characters.
[0365] In the screen shown in Figure 16(A), user 230 performs a selection operation for a desired character, for example, by touching character 1600D. At that time, the selected character / character area is highlighted, for example, by changing color. Then, the screen transitions to a screen like the one shown in Figure 16(B). The screen in Figure 16(B) is a screen that displays the details of the selected character in a larger size (character details screen). On the screen in Figure 16(B), user 230 can confirm the details of the selected character. On the screen in Figure 16(B), the indicator display shows, for example, the word "Details" to indicate that it is a screen / mode to confirm the details of the selected character.
[0366] In the screen shown in Figure 16(B), similar to Figure 9 mentioned above, the details of the displayed character can be checked by swiping or pinching in / out. In addition, this screen may also include a confirmation button 1611, a back button 1612, an operation confirmation button 1613, an audio confirmation button 1614, etc.
[0367] The Confirm button 1611 is used to confirm the character selection. The Back button 1612 is used to return to the character selection screen shown in Figure 16 (A). When the Confirm button 1611 is pressed, the character selection from the multiple outputs, i.e., the adoption, is confirmed, and the screen transitions to a character display screen, such as the one shown in Figure 9 above.
[0368] The action confirmation button 1613 allows you to check what actions (including gestures, etc.) the selected character will perform. When this button is pressed, a video of the character will be played. The character's video may also be played automatically in a loop. The voice confirmation button 1614 allows you to check the voice of the character if the character speaks or converses. When this button is pressed, the character's voice will be played. In addition to the voice, the text of the speech may also be displayed. The speech voice 1615 in Figure 16 (B) is an example.
[0369] In the illustrated example, the selected character 1600D is a variation that shares features such as "long hair" with the others, but has a different laughing gesture and speech voice 1615.
[0370] [Example of modifying features based on output] As a variation, the following functions may be provided. In the above example, the case in which the features of a character are modified based on the feature data in instruction 41 was shown. In contrast, in the following example of functions, once a character / avatar is generated and output with specified features, it is possible to perform modification instructions, regeneration, and re-output in order to maintain or modify a part of the output based on the state of that output (e.g., the displayed image). For example, the following GUI may be provided.
[0371] Figure 21 shows an example of the GUI in a modified example. For example, when re-instructing (correcting the output) from a screen like the one in Figure 9, the screen transitions to the one in Figure 21. This screen displays a character 2101 that was generated based on an instruction statement 41 that specified a certain feature. The screen in Figure 21 is in "output correction" mode, which indicates that the output state of the output character can be corrected. In the screen in Figure 21, along with the display and output of the character 2101, a guide / message "Please specify the part of the feature you want to correct" is displayed.
[0372] On this screen, user 230 can specify the part of character 2101 (this output state) that they want to modify by touch operation or other means. User 230 can specify the desired part of character 2101 by enclosing it in a frame or the like. Then, they can issue a re-instruction to maintain or change the specified part. For example, by pressing the "Regenerate" button, an instruction message 41 for output modification (re-instruction) can be sent to LLM.
[0373] In this embodiment, similar to the example in Figure 14, user 230 wants to change the "float hair" and "ribbon" of character 2101, and these parts are specified, with oval frames 2102 and 2103 displayed. When "Regenerate" is executed with these specifications, an instruction statement 41 is created that includes instructions to modify the features (part of the output) of the specified parts. This instruction statement 41 may contain instructions such as, "Modify the parts of the previously generated character enclosed in the image frame, and try not to change other parts as much as possible." This instruction statement 41 is accompanied by image data of the parts to be modified on this screen. Based on this instruction statement 41, LLM regenerates the character by modifying the features of the parts of the previously generated character enclosed in the image frame, and returns a response 42.
[0374] As a result of the regeneration, the characteristics of the character specified in frames 2102 and 2103 will be changed to other characteristics. For example, "ribbon" may be changed to another ornament (necklace, etc.). This change includes cases where the specified characteristics ("float" or "ribbon") are removed. Another format is also possible, in which the locations of the characteristics to be removed from the output state are specified.
[0375] Conversely to the above example, the screen could allow users to specify which parts of the character's output state they want to retain. In this case, LLM would not change the features of the specified parts, but would appropriately modify the features of other parts and regenerate the character. Figure 22 shows an example of the screen in this format. In the screen of Figure 22, the guide / message "Please specify the parts you want to retain" is displayed.
[0376] In this embodiment, user 230 wants to maintain (tentatively confirm) the upper body portion of character 2101, and that portion is specified, with the frame 2104 displayed. By executing "Regenerate" with this specification, an instruction statement 41 is created that includes instructions to maintain the characteristics (part of the output) of the specified portion. This instruction statement 41 may contain instructions such as, "Of the previously generated character, maintain the portion enclosed by the image frame, and modify the other portions as appropriate."
[0377] As a result of regeneration, the character will retain the features of the area specified in frame 2104, while the features of other areas may be changed. For example, the clothing on the lower body may be changed to other clothing (such as a skirt).
[0378] In this way, based on the output state of the character that has been generated, the user 230 can easily instruct the system to change or maintain some of the characteristics and re-output the character.
[0379] [Display control on an aerial floating image display device] When the response output device 10010 is an aerial floating image display device 185 (Figure 1C), the following display control related to black display may be applied.
[0380] Figure 17 is an explanatory diagram regarding display control in the aerial floating image display device 185. In the aerial floating image display device 185, the display content of the aerial floating image 195 has the characteristic that a greater sense of floating can be achieved by, for example, making the display background black. On the other hand, when the display content includes black, the black content part blends in with the black display background and appears transparent, which presents the problem that the display content becomes difficult to recognize.
[0381] Figure 17(A) shows an example in which a user 230 specifies character characteristics to the interface of the aerial floating image display device 185. An instruction statement 41 containing these characteristics is sent to the LLM server. Examples of characteristics to be specified include (1) cat, (2) entire body color is black, etc.
[0382] Figure 17(B) is an example of a character generated by LLM based on the instruction statement 41 with the features of Figure 17(A). Character 1701 in Figure 17(B) is a black cat character that reflects the specified features. The data 1702 for character 1701 is, for example, a 3D model or rendered video data based thereon.
[0383] Figure 17(C) shows an example in which the data 1702 from Figure 17(B) is displayed on the screen of the floating image 195, which corresponds to the display unit 10011 of the floating image display device 185. Specifically, an image is displayed on the screen of the display device 193, for example, a liquid crystal panel, and the floating image (optical image) 195 is formed through the imaging optical plate 194 based on the image light. In the optical image of Figure 17(C), the black part of the generated display character 1701 blends in with the black background, and the outline of the character 1701 is not visible. If left as is, the user 230 will not be able to see the character 1701 clearly. Figure 17(C) shows the display character 1701B that has blended in with the background.
[0384] In this embodiment, as in the example above, when the generated character displays black (for example, when it contains a large amount of black display), the aerial floating image display device 185 performs display control, for example, as described below, using a processor or the like.
[0385] The aerial floating image display device 185, which is the response output device 10010, determines whether the input characteristics of the instruction statement 41 include the case where "black (R0G0B0)" is specified for the display color of the display character portion adjacent to the background (black background) of the screen (optical image 195), in other words, the character outline portion. This determination may be made based on the characteristic information of the instruction statement 41, or based on the data of the response 42.
[0386] In this case, the aerial floating image display device 185 performs the following display control. Figures 17(D) to (F) show examples of this display control.
[0387] In the display control shown in Figure 17(D), the aerial floating image display device 185 changes the color of the background area adjacent to the character's black display area (outline) in the displayed image data from black to another color. The color of character 1701 is not changed. In the illustrated example, the background area 1704 adjacent to the outside of the outline (black) of character 1701 is changed to white. As a result, when the user 230 views the optical image, character 1701 appears to be outlined in white, allowing the user to recognize character 1701.
[0388] In the display control shown in Figure 17(E), the aerial floating image display device 185 increases the brightness of the black display area of character 1701 in the displayed image data. In other words, it changes to a color other than black. The black background remains unchanged. The illustrated character 1701E is gray. The brightness of the black display area of the character increases, creating a contrast difference with the black background. This allows the user 230 to recognize the character.
[0389] In the display control shown in Figure 17(F), the aerial floating image display device 185 changes the background color of the optical image 195 (screen) to a color other than black. In other words, it increases the brightness. The color of the character 1701 is not changed. The background 1706 shown is gray. By adding a color other than black to the background, the user 230 can recognize the black character.
[0390] As shown in the example above, in the case of the aerial floating image display device 185, the above display control enables a more suitable display for the character generated by LLM.
[0391] <Example 5> Example 5 will now be described. This example will mainly describe the differences from Example 4. Example 5 realizes a function in an artificial intelligence response output device such as an aerial floating image display device that sets, generates, and outputs a user's digital human (in other words, an avatar, etc.) using the response of AI, for example, a multimodal LLM.
[0392] A digital human, as defined by computer graphics and AI technology, refers to a three-dimensional model whose appearance and movements resemble those of a real person. It is sometimes also called a virtual human or 3D avatar. Digital humans have various potential applications, such as customer service and customer support. By closely reproducing the facial features, body movements, and expressions of real people, digital humans enable communication that is similar to or more natural than that between real people, fostering a sense of familiarity and connection with the user.
[0393] In this embodiment, the user creates a digital human avatar representing themselves. The artificial intelligence response output device in this embodiment uses AI responses to assist in the creation of this digital human, helping the user easily create a suitable avatar.
[0394] While Example 4 described above involved the generation of a non-realistic character, Example 5 differs significantly in that it generates a realistic digital human (avatar) 2300. Its functional features include the following:
[0395] After inputting personal characteristics information (e.g., a photograph), the user can specify whether or not to reflect, exclude, or modify each characteristic / element / item such as eyes, nose, and mouth when generating the avatar. Switching between reflecting, excluding, and modifying each characteristic item is easy.
[0396] Furthermore, for certain characteristic items among the features that constitute the digital human (avatar) 2300, predetermined exclusion or modification processes can be automatically applied by default to make them less likely to be used for biometric authentication. This ensures consideration for the protection of personal information and prevents third parties from misusing avatar data.
[0397] Furthermore, users can specify a level of stylization for each feature, reducing realism rather than making it realistic. In other words, it's possible to create an avatar that combines a realistic base with stylized representations for specific feature elements. This reduces the realism of the digital human (avatar)'s appearance, thus avoiding the so-called "uncanny valley" phenomenon.
[0398] [Configuration Overview] Figure 23 shows the overall system configuration and overview of the 5th embodiment, including the response output device 10010, illustrating the flow from user operation to avatar display. The 5th embodiment's response output device 10010 is a display device, for example, an aerial floating image display device 185 (Figure 1C).
[0399] In this embodiment, the person that user 230 specifies the characteristics of to be turned into a digital human (avatar) 2300 is user 230 himself. The LLM of the LLM application 8020 generates the user's own digital human (avatar) 2300, which is then output by the response output device 10010.
[0400] In Figure 23, the user 230 can easily create and display a digital human (avatar) 2300 by following the procedure below. The explanations 2301 to 2304 shown in the callouts describe the main steps (1) to (4).
[0401] (1) Description 2301 is a description of user operation and input. The first step is user operation input of information representing the characteristics of the person to be made into an avatar (i.e., the user 230 himself) (referred to as person characteristic information). The user inputs image data such as a photograph of himself as person characteristic information 2310 to the response output device 10010. This input can be done in the same way as in Embodiment 4.
[0402] In this case, the response output device 10010 may input image data obtained by imaging the target person (user 230 himself) using the imaging unit 1180 or the like. User 230 may also have the response output device 10010 read such image data. User 230 may also input image data, which is person characteristic information 2310, to the response output device 10010 via communication from the mobile terminal 20010.
[0403] In this case, the data of the person characteristic information 2310 to be entered is not limited to image data, but may also be video data including the person's recorded voice, for example.
[0404] In addition to the above-mentioned person characteristic information 2310, the user 230 can input, for example in natural language, 1. characteristics that they do not want to include in the avatar (in other words, characteristics they want to exclude), or 2. characteristics that they want to change when creating the avatar. Here, characteristics refer to the external features that make up the digital human (avatar) 2300, such as elements / items like eyes, nose, and mouth.
[0405] (2) Explanation 2302 is an explanation of the instruction text to the AI, LLM. The second procedure is to send the instruction text (in other words, a prompt) 2320, which was created based on the person characteristic information 2310, to the LLM of the LLM application 8020. This instruction text 2320 is an instruction text to cause the LLM to generate a digital human (avatar) 2300.
[0406] The response output device 10010 creates an instruction message 2320 that includes input information from the user 230, including person characteristic information from the first step, and transmits it to the LLM application 8020. The LLM application 8020 and LLM are provided, for example, on an LLM server 20001 on a communication network (Figure 1E). This instruction message 2320 includes data such as natural language, images, audio, or video as input information from the user.
[0407] An example of instruction statement 2320 is a combination of image data such as a photograph of a person (user 230) as person characteristic information 2310, and natural language (in other words, text) as follows.
[0408] Example 1 of natural language: "Please create a digital human (realistic avatar) of the person in this photo." Example 2 of natural language: "However, please remove the mole under the right eye." Example 3 of natural language: "Also, please change the hair color to blonde."
[0409] Example 1 corresponds to a basic instruction. Example 2 corresponds to an exclusion instruction, which excludes certain features of a photographic image from being reflected in avatar generation. Example 3 corresponds to a modification instruction, which modifies certain features of a photographic image.
[0410] (3) Explanation 2303 is an explanation of the response output from LLM. The third step is a procedure in which, based on the instruction 2320 and the control of the LLM application 8020, the LLM application 8020 transmits the data of the digital human (avatar) 2300 generated by LLM as a response 2330 to the response output device 10010, and the response output device 10010 acquires it. The data of this response 2330 includes the following data 2340 (digital human data / avatar data) of the digital human (avatar) 2300. That is, this data 2340 includes data (raw data) of the 3D model of the digital human (avatar) 2300, which was generated by reflecting and reproducing the person feature information 2310 (image data such as a photograph) input by the user 230 in the first step. This data 2340 also includes image / video and audio information data (in other words, rendering data) generated by rendering processing based on the 3D model. These data 2340 may be data from which some features instructed by the user 230 in the first and second steps have been excluded, or data from which some features have been modified.
[0411] (4) Description 2304 describes the processing of response output in the response output device 10010. The fourth step is the procedure by which the response output device 10010 outputs a response to the user in the user interface as processing in response to the data 2340 of the response 2330 from the LLM. This output includes the display of the digital human (avatar) 2300 and may be accompanied by the output of the movements of the digital human (avatar) 2300 or the output of conversational audio. Based on the data 2340 obtained in the third step, the digital human (avatar) 2300 is displayed on the screen of the display unit 10011 of the response output device 10010. The appearance of the displayed digital human (avatar) 2300 will reflect the features that were input and instructed in the first and second steps. In addition, if there are exclusion instructions or modification instructions, some features corresponding to those will be excluded or modified.
[0412] By following the above procedure, user 230 can easily create and verify a three-dimensional digital human (avatar) 2300 that reflects the characteristics of a specified person (user 230 himself). The response output device 10010 stores at least the data 2340 of the created digital human (avatar) 2300 in its memory resources and provides it to user 230 himself. The digital human (avatar) 2300 created above can be used in each case, as in Example 4.
[0413] [Input Interface Example 1] Figure 24 shows an example of a user interface (UI) when a user 230 inputs person characteristic information 2310 of the person (themselves) to be used as an avatar, and the "Image Capture" screen is shown as an example of an input interface.
[0414] The left side of Figure 24 shows an example in which user 230 inputs a photograph (image) 2401 showing the face (for example, a front view) of the person (themselves) to be displayed as an avatar to the response output device 10010, as person characteristic information 2310 of the person (themselves). In other words, the photograph 2401 is a captured image, image data. The photograph 2401 shows, for example, a person figure 2402 including the face and part of the upper body (roughly up to the chest) of the person (themselves).
[0415] Such a photograph (image) 2401 can be obtained by taking a picture of the subject person, user 230, or a paper document containing a photograph, etc., with the camera (imaging unit 1180) attached to the response output device 10010, or by transmitting an image taken by user 230's mobile terminal 20010, etc., to the response output device 10010. However, the photograph 2401 may be input by other means.
[0416] Furthermore, the person characteristic information 2310 may include multiple photographs of the same person taken from different angles or with different objects.
[0417] On the other hand, the right side of Figure 24 shows an example of the input interface on the display unit 10011 screen, or in other words, an example of a GUI screen, when the user 230 has the response output device 10010 load image data of photograph 2401. This screen is provided with, for example, an image data display field 2415. The image data display field 2415 displays the image of photograph 2401 (in other words, an image representing the characteristics of the person) 2416 after it has been captured using the method described above. The image data display field 2415 before capture may be left blank, or a message such as "Please capture an image" may be displayed to suggest to the user 230 that image data be entered.
[0418] This screen may also display other indicators, such as an indicator 2411 that shows the current display status. In the example shown in the figure, the indicator 2411 may display something like "Image Capture" to indicate that the system is waiting to capture an image or that it is displaying an image that has been captured.
[0419] This screen may also include other displays, such as a confirmation button 2413 for confirming the contents of the image after the image data has been imported. It may also include a back button 2412 for canceling, restarting, or ending the image import process. These buttons are GUI components that can be operated by, for example, a user 230 touching and pressing them with their finger 231.
[0420] In addition, this screen may display a message 2414, for example, "Is the following image OK?", above the image data display area 2415, with the intention of confirming with the user 230 whether to confirm or cancel the image import after it has been imported.
[0421] [Input Interface Example 2] Figure 25 shows an example of an input interface for when a user 230 specifies which features to exclude or modify from the person feature information 2310 (image 2416 of photograph 2401 in Figure 24). In other words, Figure 25 shows the "Input (Exclude / Modify)" screen as an example of a screen for specifying the exclusion and modification of some person features.
[0422] After inputting the aforementioned photograph 2401 and capturing the image, the user 230 then, on a screen like the one shown in Figure 25, specifies and inputs either or both of the features from the person feature information 2310 that they do not want to include when creating an avatar (in other words, features they want to exclude) and features that they want to change. These features (in other words, elements or items) are, for example, the external features of a person's image (image 2416), such as hair, eyes, and nose, which are identifiable and distinguishable parts or units.
[0423] In this case, user 230, while viewing the captured image 2508 (similar to image 2416) in the image data display field 2507 on this screen, enters instructions in natural language in the instruction input field 2505 to specify which external features to exclude or change.
[0424] This screen may also include an image data display area 2507 that displays the captured image 2508 (similar to image 2416), similar to the "Image Capture" screen in Figure 24 described above. Above the image data display area 2507, for example, a label 2506 such as "↓Preview" may be displayed to indicate that the image displayed in the area is the captured image data.
[0425] This screen includes an instruction input field 2505 in which the content of exclusions or changes can be entered in natural language, and the entered natural language is displayed. Above the instruction input field 2505, a notation 2504 may be displayed to instruct and guide the user 230 to enter the content of exclusions or changes, for example, "Please enter the content you wish to exclude or change." The instruction input field 2505 can be operated by the user 230, for example, by touching it with their finger 231, which will switch the screen to a user interface for natural language (text) input (for example, a software keyboard), or the interface will pop up. The user can then input in natural language using that interface. The switched screen is provided with a complete instruction input button and an interrupt button, which the user 230 can use to complete or interrupt the input and return to the instruction input field 2505 on the "Input (Exclusion / Change)" screen in Figure 25.
[0426] This screen may include an indicator 2501 that shows the current display status, a back button 2502, a confirmation button 2503, etc. The indicator 2501 may be labeled, for example, "Input (Exclude / Change)" to indicate that this is a screen for inputting instructions to exclude or change appearance features. The back button 2502 is a button to cancel the input of instructions and transition to the "Image Capture" screen in Figure 24. The confirmation button 2503 is a button to confirm the input of instructions and instruct the LLM to create the avatar.
[0427] Instructions regarding the exclusion or modification of features (exclusion instructions, modification instructions) can be entered in natural language in the instruction input field 2505, for example, as shown in the following example sentence.
[0428] Example 1: Avatar creation instruction (basic instruction) 2505A: "This image is a photograph of a person. Please return the digital human of the person in the photograph." Example 2: Feature removal instruction 2505B: "Please remove the beard." Example 3: Feature modification instruction 2505C: "Please change the hair color to white."
[0429] The input and display of natural language in the instruction input field 2505 may include basic instructions such as creating an avatar based on the captured image, as in Example 1.
[0430] Exclusion instruction 2505B is an exclusion instruction that prevents certain features, such as "beard," from being included in the avatar (the generated and outputted avatar) from the image 2508 of the person feature information 2310. In other words, it is an exclusion instruction that prevents those features from being reflected in the avatar generation.
[0431] Change instruction 2505C is a change instruction to modify some of the features of the image 2508 of the person feature information 2310, for example, "hair," from the original feature content of the image 2508, for example, changing it from black to white, and reflecting this change in avatar generation.
[0432] Through the means described above, user 230 can input exclusion instructions or modification instructions, and the content of these instructions can be reflected in the avatar generation by LLM as an instruction statement 2320. That is, LLM generates an avatar according to the content of the instructions, and response output device 10010 can output the digital human (avatar) 2300.
[0433] However, if user 230 has multiple features to exclude or modify, the process of inputting exclusion and modification instructions may become cumbersome. Therefore, as an alternative, exclusion and modification instructions may be made possible using a feature list or image selection method, as described later. In that case, user 230 can use the feature list, etc., to confirm whether to reflect, exclude, or modify each feature that makes up the avatar, and can issue exclusion and modification instructions for multiple features at once.
[0434] The screen in Figure 25 may include buttons for transitioning to screens for inputting exclusion / modification instructions using other methods, such as the feature list mentioned above. For example, a "Input (Select from Image)" button 2509 and an "Input (Select from List)" button 2510 may be provided. The "Input (Select from Image)" button 2509 is a button that transitions to the screen (Figure 31) when exclusion / modification instructions are given using the image selection method. The "Input (Select from List)" button 2510 is a button that transitions to the screen (Figure 26) when exclusion / modification instructions are given using the feature list method. The user 230 can transition screens by operating these buttons.
[0435] [Input Interface Example 3] Figure 26 shows the "Input (Select from List)" screen as an example of an input interface for when a user 230 specifies features to exclude or modify, using a feature list method. On this screen, the user can select features to exclude or modify from the feature list.
[0436] For example, from the "Input (Exclude / Change)" screen in Figure 25, user 230 presses the "Input (Select from List)" button 2510 to transition to the screen in Figure 26. This screen may also include an indicator 2601, a back button 2602, a confirmation button 2603, etc., as described above. The indicator 2601 may display, for example, "Input (Select from List)". The back button 2602 is a button to return to the "Image Capture" screen or the "Input (Instruction / Change)" screen. The confirmation button 2603 is a button to confirm the exclusion / change instruction.
[0437] This screen includes a list display 2605. Above the list display 2605, there is a column name 2606. The column name 2606 includes columns and notations such as "Features," "Reflect," "Exclude," "Change," and "Change Details."
[0438] The "Features" column 2607 displays a list (feature list) 2608 in which a person's physical characteristics (names, phrases, and feature items that describe them) are arranged vertically on each row, such as "skin," "forehead," "hair," etc.
[0439] The feature list 2608, or in other words, the multiple feature items, described in these "features" column 2607 may display a preset of representative physical features in advance within this system. Alternatively, as described later (Figure 30), the list 2608 may display physical features extracted by referencing and analyzing the human image recorded in the captured image through software processing.
[0440] To the right of the "Features" column 2607 are the "Reflect" column 2609, the "Exclude" column 2610, and the "Change" column 2611. These columns have corresponding checkboxes for each feature in the list 2608, for example. These checkboxes are GUI components that can be turned on (checked) or off (unchecked, blank) by the user.
[0441] This checkbox, for example, is initially in an off state (blank) and can be turned on when user 230 clicks it. When the checkbox is on, a check mark (✓) is displayed inside the box. The operation / action of displaying a check mark on a blank checkbox (which is in an off state) is sometimes referred to as "checking," "putting a check," or "performing a check."
[0442] Furthermore, checkboxes that are already checked (indicated by a checkmark) can be turned off by user 230 pressing them once or again. When a checkbox is off, it does not display a checkmark (✓). The operation / action of removing the checkmark from a checked checkbox is sometimes referred to as "unchecking" or "turning off the checkmark."
[0443] If user 230 wants to reflect each feature in the feature list 2608 in the list display 2605 onto their avatar, in other words, if they want to reproduce that feature, they check the checkbox corresponding to that feature in the "Reflect" column 2609. This allows them to send instructions for reflecting that feature to the LLM by writing them in the instruction text 2320.
[0444] Furthermore, if user 230 wants to exclude a particular feature from their avatar, they check the corresponding checkbox in the "Exclude" column 2610. This allows them to specify the exclusion instruction for that feature in the instruction statement 2320 and communicate it to the LLM.
[0445] Furthermore, if user 230 includes each feature in the avatar, but makes changes to it, or in other words, if there are changes to be made, they check the checkbox corresponding to that feature in the "Change" column 2611. This allows them to write the change instructions for that feature in the instruction statement 2320 and communicate it to LLM.
[0446] In this embodiment, the GUI screen format shown in Figure 26 ensures that for the same feature, it is not possible to have two or more checkboxes checked simultaneously: "Apply," "Exclude," and "Change." In other words, it is possible to select only one of these three checkboxes. For example, the checkbox that was last pressed by the user will be checked. If the system is not in a state of selection when the confirmation button is pressed, an alert or guide may be displayed to the user to enable selection.
[0447] As an alternative (other form) variation, when a change is instructed by checking the "Change" column 2611 for the same feature, both the "Apply" and "Change" checkboxes may be checked. For example, when the "Change" checkbox is clicked, the corresponding "Apply" checkbox may also be checked, regardless of whether a checkmark is already displayed in "Apply". Furthermore, if both the "Apply" and "Change" checkboxes are already checked, and the "Change" checkbox is clicked, the "Apply" checkbox may not be unchecked.
[0448] The "Change Details" column 2612 is a GUI that allows the user 230 to input the desired change details in natural language when the checkbox in the "Change" column 2611 is in the checked state. In the "Change Details" column 2612, for example, input fields 2613 for change details corresponding to each feature are provided. In the input fields 2613 for change details, the user 230 can input the change de...
Claims
1. A response output device that outputs a response from artificial intelligence, comprising an interface to the artificial intelligence, the interface receiving input from a user specifying the characteristics of a character to be generated by the artificial intelligence, creating a prompt for the artificial intelligence to generate the character reflecting the characteristics based on the input information and transmitting it to the artificial intelligence, receiving data of the character reflecting the characteristics generated by the artificial intelligence, and outputting the character at the interface based on the received data.
2. The response output device according to claim 1, wherein the artificial intelligence is a multimodal large-scale language model, and the input specifying the features in the interface can be text, images, or audio.
3. A response output device according to claim 1, wherein the interface displays a list of feature items constituting the character's features for input specifying the features, and accepts input of a value for each feature item.
4. A response output device according to claim 1, wherein the interface receives input for correcting previously input features after the user has confirmed the outputted character, creates a prompt to regenerate the character reflecting the corrected features based on the input corrected feature information and sends it to the artificial intelligence, receives data of the character reflecting the corrected features regenerated by the artificial intelligence, and outputs the regenerated character at the interface based on the received data.
5. The response output device according to claim 4, wherein the interface allows specifying at least one of a feature to be maintained and a feature to be changed for the input for modifying the previously input feature.
6. A response output device according to claim 1, wherein the interface receives input to modify some of the characteristics of the output state of the character after the user has confirmed the outputted character, creates a prompt to regenerate the character based on the information of the input to modify some of the characteristics of the output state and transmits it to the artificial intelligence, receives data of the character with some of the characteristics of the output state modified that has been regenerated by the artificial intelligence, and outputs the regenerated character at the interface based on the received data.
7. A response output device according to claim 6, wherein the interface allows specifying at least one of the features to be maintained and the features to be changed for an input to modify some of the features of the output state.
8. A response output device according to claim 1, wherein the interface allows input for specifying the features to be an image, and the interface allows exclusion of some feature information contained in the image so as not to be reflected in the generation of the character.
9. A response output device according to claim 1, wherein, after the user confirms the outputted character in the interface, when the user performs a reload operation, the interface creates a prompt to generate the character reflecting the previously input features based on the feature information and sends it to the artificial intelligence, receives data of the character reflecting the features that has been generated again by the artificial intelligence, and outputs the character in the interface based on the received data.
10. A response output device according to claim 1, wherein the interface creates a prompt for generating a plurality of characters that reflect the input features based on the input feature information and transmits it to the artificial intelligence; receives data of the plurality of characters that reflect the features generated by the artificial intelligence; and outputs the plurality of characters at the interface based on the received data.
11. The response output device according to claim 1, wherein the response output device is an aerial floating image display device, and the response output device displays the character in the aerial floating image.
12. The response output device according to claim 11, wherein the aerial floating image display device displays the character in the aerial floating image based on the background being black, and when the character has a black portion, it performs one of the following display controls: (A) Change the black portion of the character and the adjacent black background portion to a color other than black, (B) Change the black portion of the character to a color other than black, (C) Change the black background portion to a color other than black, the response output device.
13. A response output system comprising the response output device according to claim 1 and the server equipped with artificial intelligence.
14. A response output method performed in a response output device that outputs a response from artificial intelligence, the response output device comprising: an interface for the artificial intelligence; the response output device receiving input from a user specifying the characteristics of a character to be generated by the artificial intelligence at the interface; the response output device creating a prompt for the artificial intelligence to generate the character reflecting the characteristics based on the input information and transmitting it to the artificial intelligence; the response output device receiving data of the character reflecting the characteristics generated by the artificial intelligence; and the response output device outputting the character at the interface based on the received data.
15. A response output device that outputs a response from artificial intelligence, comprising an interface to the artificial intelligence, the interface receiving input from a user specifying the characteristics of a digital human to be generated by the artificial intelligence, creating a prompt to generate the digital human reflecting the characteristics based on the input information and transmitting it to the artificial intelligence, receiving data of the digital human reflecting the characteristics generated by the artificial intelligence, and outputting the digital human at the interface based on the received data.
16. A response output device according to claim 15, wherein the artificial intelligence is a multimodal large-scale language model, and in the interface, the input specifying the features can be text, images, or audio.
17. A response output device according to claim 15, wherein the interface displays a list of feature items constituting the characteristics of the digital human for input to specify the characteristics, and accepts input of a value for each of the feature items.
18. A response output device according to claim 15, wherein the interface receives input for correcting previously input characteristics after the user has confirmed the outputted digital human, creates a prompt to regenerate the digital human reflecting the corrected characteristics based on the input corrected characteristic information and transmits it to the artificial intelligence, receives data of the digital human reflecting the corrected characteristics regenerated by the artificial intelligence, and outputs the regenerated digital human at the interface based on the received data.
19. A response output device according to claim 18, wherein the interface allows specifying at least one of a feature to be maintained and a feature to be changed for the input for modifying the previously input feature.
20. A response output device according to claim 15, wherein, after the user confirms the outputted digital human, the interface receives input to modify some of the characteristics of the output state of the digital human; based on the information of the input to modify some of the characteristics of the output state, creates a prompt to regenerate the digital human and transmits it to the artificial intelligence; receives data of the digital human with some of the characteristics of the output state modified, which has been regenerated by the artificial intelligence; and outputs the regenerated digital human at the interface based on the received data.
21. A response output device according to claim 20, wherein the interface allows specifying at least one of a feature to be maintained and a feature to be changed for an input to modify some of the features of the output state.
22. A response output device according to claim 15, wherein the interface allows input for specifying the features to be an image, and the interface allows exclusion of some feature information contained in the image so as not to be reflected in the generation of the digital human.
23. A response output device according to claim 15, wherein, after the user confirms the outputted digital human in the interface, when the user performs a reload operation, the response output device creates a prompt to regenerate the digital human reflecting the previously input characteristic information and transmits it to the artificial intelligence, receives data of the digital human reflecting the characteristic regenerated by the artificial intelligence, and outputs the digital human in the interface based on the received data.
24. A response output device according to claim 15, wherein the interface creates a prompt for generating a plurality of digital humans that reflect the input features based on the input feature information and transmits it to the artificial intelligence; receives data of the plurality of digital humans that reflect the features generated by the artificial intelligence; and outputs the plurality of digital humans at the interface based on the received data.
25. The response output device according to claim 15, wherein the response output device is an aerial floating image display device, and the response output device displays the digital human in the aerial floating image.
26. A response output device according to claim 15, wherein the interface allows input for specifying the features to be an image, extracts feature information from the image by dividing it into parts, obtains the name of the extracted feature, an image of the extracted feature, and association information thereof, and based on the obtained information, displays a list of feature items constituting the features of the digital human for input for specifying the features to the interface, and accepts input of a value for each feature item.
27. A response output device according to claim 15, wherein the device controls the system to exclude or modify certain characteristic items among the characteristics constituting the digital human so that they are less likely to be used for biometric authentication.
28. A response output device according to claim 15, wherein the response output device controls the user to perform deformation processing on the feature items among the features constituting the digital human so that they are expressed in a less realistic manner.
29. A response output device according to claim 28, wherein the response output device displays a plurality of images that are candidates for the representation after deformation for the feature item specified by the user, and controls the device to apply the image selected by the user as the representation after deformation.
30. A response output device according to claim 15, wherein the device uploads the data of the digital human to a server on a communication network, controls the server to issue a code to access the data of the digital human to the user's device, and accesses and obtains the data of the digital human on the server from the user's device based on the code.
31. A response output system comprising the response output device according to claim 15 and the server equipped with artificial intelligence.
32. A response output method performed in a response output device that outputs a response from artificial intelligence, the response output device comprising: an interface for the artificial intelligence; the response output device receiving input at the interface specifying the characteristics of a digital human to be generated by a user; the response output device creating a prompt based on the input information to generate the digital human reflecting the characteristics and transmitting it to the artificial intelligence; the response output device receiving data of the digital human reflecting the characteristics generated by the artificial intelligence; and the response output device outputting the digital human at the interface based on the received data.