system
The system addresses limitations in existing information processing technologies by allowing users to interact with character traits through stereoscopic displays and audio, enhancing user experience and flexibility.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing information processing technologies are limited to ordinary text and voice interactions, making it difficult for users to experience familiarity and flexibility, especially in fields like education and entertainment, where individual needs are not adequately addressed.
A system comprising input, model selection, response generation, and output means that allows users to interact with information processing models having character traits, utilizing stereoscopic displays and audio devices for richer experiences.
Enables intuitive and user-friendly interactions by providing responses through stereoscopic displays and audio, enhancing user experience and flexibility in meeting individual needs.
Smart Images

Figure 2026074980000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Existing information processing technologies are usually limited to ordinary text and voice for interaction with users, making it difficult for users to use intuitively with a sense of familiarity. Also, especially in the fields of education and entertainment, it is difficult to provide flexible responses and experiences according to individual needs. In such a situation, there is a need for a new system to improve the user experience and make AI technology feel closer.
Means for Solving the Problems
[0005] The present invention is a system comprising an input means for receiving input from a user, a model selection means for selecting an information processing model having the appropriate characteristics, a response generation means for generating a response based on the user input, and an output means for outputting the generated response. This system allows users to interact with information processing models that have character traits, enabling them to use the system intuitively and with a sense of familiarity. Furthermore, by providing responses to the user using a stereoscopic display device or sound device, a richer and more user-friendly experience can be realized.
[0006] "Input means" refers to a device or function that receives information such as voice or text from the user.
[0007] "Model selection means" refers to a device or function that selects an appropriate information processing model based on user input and utilizes its functions.
[0008] "Response generation means" refers to a device or function that generates an appropriate response corresponding to user input by utilizing a selected information processing model.
[0009] "Output means" refers to a device or function for presenting the generated response to the user.
[0010] An "information processing model" is a data structure or algorithm designed to process data based on a specific task or condition and produce a specific output.
[0011] A "stereoscopic display device" is a device used to project characters and other visual information into a three-dimensional space.
[0012] An "audio device" is a device that reproduces audio signals and transmits them to the user in a format that can be heard. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, a numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system for realizing interaction between a user and an information processing model, aiming to make AI technology more accessible using various devices. This system selects an appropriate information processing model according to the user's needs and generates and outputs a response based on it.
[0035] Specifically, users access the system using devices such as smartphones, tablets, and PCs. These devices are equipped with input methods for voice recognition and text input, and receive instructions from the user. The received instructions are then sent from the device to the server.
[0036] The server analyzes the user's input and selects an appropriate information processing model based on its content. This model selection is based on the user's requested problem or theme, and then a specific response is generated using the selected model. In response generation, natural language processing techniques are used to create information that is easy for the user to understand and has a distinct character.
[0037] The generated response is sent from the server to the terminal, which then presents it to the user. For example, if a 3D hologram monitor is used, the terminal displays the character in 3D, enabling interaction with the user. Alternatively, if a smart speaker is used, information is provided via voice. This allows the user to experience interaction visually or aurally.
[0038] As a concrete example, consider a case where a user wishes to learn a language. The user uses the device's voice input function to send an instruction to the system to begin language learning. The server analyzes this instruction and selects the appropriate information processing model for language learning. The server then generates a language lesson plan and sends its contents to the device. Since the device delivers the lesson through a character, the user can proceed with their learning under the guidance of the character.
[0039] In this way, this system enhances the quality of interaction and provides convenience tailored to the user's purpose.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The user enters their request using voice or text via the device. This request is recognized or received by the device's input method.
[0043] Step 2:
[0044] The terminal sends user input to the server. During this process, the input data may be converted to an appropriate format.
[0045] Step 3:
[0046] The server analyzes the user's input and identifies a task based on its content. For example, if a user asks a question about language, the server identifies that category.
[0047] Step 4:
[0048] The server selects the information processing model best suited to the specified task. This model selection mechanism makes it possible to create a response that meets the user's objectives.
[0049] Step 5:
[0050] The server generates a user-appropriate response using the selected model. This response is written in natural language and includes elements that give it a character.
[0051] Step 6:
[0052] The server sends the generated response to the terminal. The response may include text, audio, and possibly visual data.
[0053] Step 7:
[0054] The terminal displays the received response to the user. For example, if a 3D hologram monitor is used, a character is projected in three dimensions, and dialogue takes place via voice.
[0055] Step 8:
[0056] The user can then interact further based on the information presented through the device. In this case, the process is repeated from step 1.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] With the evolution of information technology, there is a growing demand for information provision that meets diverse user needs. However, conventional systems have limitations in terms of user experience, making it difficult to respond quickly and accurately to individual needs. To address these challenges, there is a need for systems that provide intuitive and user-friendly interactions.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes terminal means for receiving information from the user, server means for selecting a model having the appropriate characteristics, and generation means for generating a response based on the information received from the user. This enables a rapid and accurate response to the diverse needs of the user.
[0062] A "terminal device" is a device that receives information from a user in the form of voice or text and transmits that information to a server.
[0063] A "server device" is a device that analyzes the received user information and selects a model with the corresponding characteristics based on the analysis results.
[0064] A "generation means" is a device that uses a selected model to generate an appropriate response based on the user's request.
[0065] A "presentation means" is a device for providing the generated response to the user visually or audibly.
[0066] A "three-dimensional display device" is a device that displays generated information in three dimensions, allowing users to experience visual interaction.
[0067] A "speech output device" is a device that converts the generated response into speech and provides it to the user audibly.
[0068] This invention relates to an interactive system for users to input information and receive appropriate responses. The system consists of terminal means, server means, generation means, and presentation means. The aim is to provide an interactive environment that can meet diverse needs.
[0069] Users access the system using terminal devices such as smartphones, tablets, or PCs. These terminal devices are equipped with voice recognition technology and text input capabilities, and can receive user input information in digital format. The terminal devices then transmit this information to the server devices via the internet.
[0070] The server uses artificial intelligence and natural language processing algorithms to analyze information received from the user. This analysis process helps understand the user's request and selects an appropriate information processing model using a generative AI model based on that understanding. The selected model is then used by the generative means to generate a specific response to the user's request.
[0071] The generated response is provided to the user via a presentation device. In the case of a three-dimensional display device, the response is displayed as a three-dimensional character, allowing the user to enjoy visual interaction. If an audio output device is used, the response is converted into audio, allowing the user to receive the information audibly.
[0072] As a concrete example, if a user wishes to learn a language, they send the instruction "I want to learn everyday conversation in English" to the system using the voice input function of their device. The server analyzes this instruction and selects the appropriate language learning information processing model. This generating AI model creates an appropriate lesson plan and provides it to the user via a presentation device. The user can then proceed with their learning in an interactive format through a three-dimensional character.
[0073] Example prompt: "If the user wants to learn English, generate a plan that provides lesson content tailored to the user's level."
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] Users access the system using devices such as smartphones and tablets, and input information via voice or text. For example, a user might input a specific request such as "I want to start an English lesson" into their device. The input instructions become input data for the system.
[0077] Step 2:
[0078] If voice input is used, the terminal uses speech recognition technology to convert the voice data into text data. If text input is used, it is retained in its original format. This text data is output as data to be sent from the terminal to the server. It is then transmitted to the server via the network.
[0079] Step 3:
[0080] The server analyzes the text data received from the terminal. Using natural language processing (NLP) techniques, it understands the user's request and identifies its intent. The user's request content obtained as a result of the analysis becomes the processing data within the server.
[0081] Step 4:
[0082] Based on the analysis results, the server selects a suitable generation AI model for the request. The selection process searches the database for the model that best matches the user's instructions and confirms it. This selected model is then used as input for the next processing step.
[0083] Step 5:
[0084] The server uses a selected generative AI model to generate a response to the user's request. For example, it can automatically construct an English lesson plan. The generated response becomes output data within the server and is formatted as data to be sent to the terminal.
[0085] Step 6:
[0086] The server sends the formatted response data to the terminal. Real-time communication protocols may be used here. The response data is converted into a format that can be presented to the user visually or audibly.
[0087] Step 7:
[0088] The terminal presents the received response data to the user. If the terminal is equipped with a 3D display, it uses characters to display the response in 3D. If a smart speaker is used, it provides information to the user via voice. The user can interact through this presentation and receive feedback on their requests.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] Traditional virtual stores have struggled to respond flexibly and interactively to individual user needs, resulting in decreased customer satisfaction and a lack of repeat store visits. Furthermore, there was a lack of efficient and effective technological means to provide visually and audibly rich interactions.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes a receiving device that receives information from the user, a model selection device that selects and uses a data processing model with the appropriate characteristics, and a response generation device that generates a response based on the information received from the user. This enables personalized interaction and provides a rich customer experience utilizing visual and auditory elements.
[0094] A "receiving device" is a device that receives information from a user as audio or text and converts it into a format that can be processed within the system.
[0095] A "model selection device" is a device used to select the optimal data processing model based on received information and generate a response.
[0096] A "reaction generation device" is a device that uses a selected data processing model to generate appropriate reactions according to the user's needs.
[0097] A "presentation device" is a device that effectively presents the generated response to the user, and uses visual and auditory means.
[0098] A "3D display system" is a technology that displays generated reactions in three dimensions, providing a visual impact to the user.
[0099] An "acoustic system" is a system that reproduces the generated response as sound, appealing to the user's hearing.
[0100] A "visual device" is an auxiliary device that utilizes the user's visual information to further enhance the user experience.
[0101] To implement this invention, the receiving device first acquires information from the user in the form of voice or text. The user can then initiate visual and auditory interaction by entering a virtual store through smart glasses or other visual devices. For example, the user might ask, "Which coat is suitable for winter?" within the virtual store.
[0102] This information is analyzed by a model selection device, and an appropriate data processing model is selected. The model used here is a generative AI model equipped with advanced natural language processing technology. Once this model is selected, the response generator produces responses tailored to the user's questions, creating personalized information.
[0103] The server provides this generated information to the user through a display device. The display device uses a stereoscopic display system to show three-dimensional images of related products and an audio system to provide detailed information via voice. The visual device can also highlight specific products based on the direction the user is looking. This allows for a visually and aurally enriching customer experience.
[0104] As a concrete example, the system translates a user's question into a prompt such as "Please recommend a coat for winter," and based on this prompt, a generating AI model constructs appropriate product information and recommendations, generating a response for the user. This enables flexible and interactive responses that meet the user's needs.
[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0106] Step 1:
[0107] The user accesses a virtual store through smart glasses. The device collects the user's voice commands and text messages via a receiving device and sends them to a server. The input here includes the user's interests and questions, such as "What products do you recommend for winter?"
[0108] Step 2:
[0109] The server uses a model selection device to select an appropriate data processing model based on the user information received. The input is a question from the user, which is then analyzed to perform data calculations to select the most relevant processing model. The output is information about the selected model.
[0110] Step 3:
[0111] The server uses a response generator to produce appropriate responses for the user from a selected data processing model. In this process, a generation AI model is used, and input is given in the form of a prompt such as "Please recommend some products for winter." As a result, detailed product information and recommendation information is generated. The output is a specific response to be presented to the user.
[0112] Step 4:
[0113] The server transmits the generated response to the terminal using a display device. The terminal presents the response to the user visually and audibly using a stereoscopic display system and an audio system. The input is the response data generated from the server, and the output is a visual and audible display that the user can confirm. When a visual device is used, information is also highlighted based on the direction the user is looking.
[0114] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0115] This invention is an information processing system that combines emotion recognition technology to improve user interaction. The system aims to achieve more natural and engaging dialogue by understanding the user's emotional state and providing appropriate responses.
[0116] This system is equipped with input mechanisms for receiving voice and text input from users. Users can access the system using a smartphone, PC, or dedicated device. The input mechanisms are connected to a server and are responsible for receiving instructions and inquiries from users.
[0117] The server transmits the received data to an emotion recognition system. This system analyzes the user's voice tone, facial expressions, and linguistic characteristics to evaluate their current emotional state. This information is fed back into the information processing model selection process, and the selected model and the generated response are adjusted based on the user's emotions.
[0118] The server selects an appropriate information processing model and generates an emotion-sensitive response. This response is generated using natural language processing to make it easy for the user to understand and relate to. The generated response is then sent to the terminal.
[0119] The terminal presents the received response to the user as audio or visual. In terminals equipped with a 3D display, the character is projected three-dimensionally, as if speaking directly to the user. In terminals with sound equipment, the character's voice and sounds are played in real time.
[0120] For example, if a user voice-inputs "I'm feeling a little down today" into the device, the emotion recognition system will determine the emotional state to be "sadness" based on the user's tone of voice and words. Based on this information, the server will select an appropriate comforting character, generate a gentle message to encourage the user, and send it to the device. The device will then display the character and deliver the comforting message to the user via voice.
[0121] In this way, systems that incorporate emotion engines can provide users with a more human-like and interactive experience.
[0122] The following describes the processing flow.
[0123] Step 1:
[0124] The user enters a request via voice or text through the terminal. The terminal receives this input and prepares to send the data to the server.
[0125] Step 2:
[0126] The device sends user input data to the server. This data uses a format designed to be sensitive to emotions.
[0127] Step 3:
[0128] The server analyzes the received data and uses emotion recognition to identify the user's emotions. Emotions are detected from voice tone and word choice.
[0129] Step 4:
[0130] The server decides which information processing model to use based on the emotion recognition results. For example, if the user appears happy, it will select a model that generates a cheerful response.
[0131] Step 5:
[0132] The server generates a user-specific response based on the selected information processing model. This response includes emotionally appropriate language and character tone.
[0133] Step 6:
[0134] The server sends the generated response to the terminal. The data is optimized to make the user feel more emotionally satisfied.
[0135] Step 7:
[0136] The terminal presents the received response to the user either audibly or visually. If necessary, it projects a character using a 3D display to express emotions.
[0137] Step 8:
[0138] The user can further interact based on the device's response. Emotionally responsive dialogue can continue, and the process restarts from step 1.
[0139] (Example 2)
[0140] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0141] Conventional dialogue systems have been unable to generate responses that adequately consider the user's emotional state, making it difficult to provide natural and affiliative interactions. Therefore, there is a need to improve the user experience.
[0142] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0143] In this invention, the server includes means for acquiring emotional information from the user, selection means for selecting and using the appropriate information processing technology, and response creation means for generating prompt sentences and creating responses based on the user's emotions. This makes it possible to provide natural and engaging dialogue that responds to the user's emotional state and improve the user experience.
[0144] "Emotional information" refers to data that indicates the user's emotional and mood state, and includes information such as voice characteristics and text characteristics.
[0145] "Information processing technology" is a collection of algorithms and models used for data analysis and response generation.
[0146] A "selection mechanism" is a component that has the function of selecting the optimal technology or method based on specific conditions or input data.
[0147] A "prompt" is text given to a generative model as instructions or hints to generate a target response.
[0148] A "response generation means" is a component that processes information to generate an appropriate response based on user input and emotional information.
[0149] "Presentation means" refers to a device or function for providing the generated content to the user through sight or hearing.
[0150] This invention relates to an information processing device for realizing natural and effective interaction with users. The system aims to improve the user experience by analyzing user input, recognizing emotions, and generating appropriate responses that are in line with those emotions.
[0151] Users access this system using devices such as smartphones and personal computers. These devices incorporate speech recognition software and input devices, enabling the acquisition of emotional information from speech and text. The system converts speech input into text using Google's (registered trademark) speech-to-text API, among others. Furthermore, it utilizes natural language processing technology to analyze the emotions expressed in the user's text input.
[0152] After receiving emotional information from the terminal, the server selects the optimal information processing technology through a pre-configured selection mechanism. The selected technology includes generative AI models, and in particular, natural language processing models utilize technologies based on open-source AI frameworks. This model generates responses to the user by leveraging prompts. An example of a prompt used is, "The user says, 'I'm feeling a little down today.' Please think of an encouraging message for them." This makes it possible to generate flexible and personalized responses that respond to the user's emotions.
[0153] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology, such as Google's text-to-speech technology, to convey the generated response to the user verbally. If visual elements are included, the content of the response can also be visually displayed on the terminal's display device. This provides the user with a natural and user-friendly interaction.
[0154] For example, if a user voice-inputs "I'm a little tired today," the system converts the voice into text and recognizes the emotion as "fatigue." The server then uses a generative AI model to generate a response such as "You did a great job today, please get some good rest tonight." The terminal reads this message aloud, speaking warmly to the user. In this way, the system responds sensitively to the user's emotions and enables appropriate dialogue.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] Users input information via voice or text using a smartphone or PC. The device uses speech recognition software to convert the voice into text. The raw data (voice or text) is received as input, and a process is performed to convert that data into text information, resulting in the output being the converted text data.
[0158] Step 2:
[0159] The terminal sends the converted text data to the server. The server receives this text data as input and sends it to an emotion recognition module. This module uses natural language processing technology to perform emotion analysis on the text and outputs emotion labels (e.g., "joy," "sadness," "surprise," etc.).
[0160] Step 3:
[0161] The server selects the appropriate information processing technology based on the obtained emotion labels. This involves a technology selection process that includes a generative AI model. The emotion labels are used as input, and an AI model is selected. The selected AI model is obtained as output.
[0162] Step 4:
[0163] The server generates prompts using a selected generative AI model. For example, it might generate the prompt, "The user says 'I'm feeling a little down today.' Please come up with an encouraging message." The input consists of a sentiment label and a standard prompt template, and the output is a specific prompt.
[0164] Step 5:
[0165] Based on the generated prompt text, the AI model generates a response. The server manages this process and inputs the prompt text into the AI model to generate an appropriate response for the user. As output, a customized response text for the user is obtained.
[0166] Step 6:
[0167] The server sends the generated response to the terminal. The terminal receives this response and presents it to the user as speech using a speech synthesis system. It also displays it visually using a display device if necessary. It receives the generated response text as input and provides the user with speech or visual content as output.
[0168] In this way, the entire system operates in a chain reaction, providing flexible responses tailored to the user's emotions.
[0169] (Application Example 2)
[0170] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0171] In recent years, there has been a growing demand for methods to improve the user experience, but current systems struggle to accurately grasp users' emotional states and generate individualized responses accordingly. In particular, there are difficulties in real-time emotion recognition and feedback, resulting in a challenge in providing natural and user-friendly interactions.
[0172] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0173] In this invention, the server includes an information processing device that receives input from a user, a pattern selection device that selects and uses an information processing pattern having the appropriate characteristics, and a response generation device that generates a response based on the input received from the user. This makes it possible to recognize the user's emotional state and generate and present an appropriate response in real time.
[0174] An "information processing device" is a device that receives input from a user and analyzes that data.
[0175] A "pattern selection device" is a device that selects and uses an information processing pattern with the appropriate characteristics based on the received input data.
[0176] A "reaction generation device" is a device that generates an appropriate reaction based on input received from a user.
[0177] An "output device" is a device that presents the generated reaction to the user.
[0178] The term "device" refers to a machine or instrument configured to perform a specific function within a system.
[0179] An "emotion recognition device" is a device that accurately recognizes the user's emotional state and adjusts its response based on that information.
[0180] A system for implementing this invention is configured primarily to provide emotion recognition and interaction based thereon. The system includes an information processing device, a pattern selection device, a response generation device, and an output device.
[0181] The server receives audio and text input from the user via an information processing device. The received data is sent to an emotion recognition device that analyzes voice tone, facial expressions, and word characteristics. The emotion recognition device evaluates the user's emotional state based on the analysis results and provides this information to a pattern selection device. The pattern selection device then selects a specific information processing pattern, and a response generation device uses natural language processing technology to generate a clear and user-friendly response for the user.
[0182] This system also uses a three-dimensional display device and an audio device via an output device to present the generated response to the user. For example, on a terminal equipped with a three-dimensional display device, a friendly character is displayed three-dimensionally, as if speaking directly to the user.
[0183] As a concrete example, when this system is used in a physical store, service staff wear smart glasses. They can instantly recognize the customer's emotional state and provide real-time instructions on appropriate customer service methods. For example, if a customer is nervous, the system might suggest, "Please use a relaxing customer service phrase." An example of a prompt sentence for the generating AI model would be, "The customer seems anxious, how can you reassure them?"
[0184] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0185] Step 1:
[0186] The user provides input. Voice and text data from the user are transmitted to the information processing device via smart glasses or other devices. The input data reflects the user's real-time emotions.
[0187] Step 2:
[0188] The server transmits the received data to the emotion recognition device. The emotion recognition device analyzes voice tone, linguistic features, and facial expression data to evaluate the user's emotional state. The analysis results are output as data indicating the user's specific emotion.
[0189] Step 3:
[0190] The server selects an appropriate information processing pattern in the pattern selection device based on the evaluated emotional state. In the selection process, the pattern most relevant to the emotional state is selected, and the selection result is retained.
[0191] Step 4:
[0192] The server generates a response based on the selected information processing pattern. Using natural language processing techniques, it creates a user-friendly and approachable response. The generated response takes the user's emotional state into consideration.
[0193] Step 5:
[0194] The terminal presents the generated response to the user via an output device. Three-dimensional display devices and sound devices are used to present the response using three-dimensional characters and voices. The output is perceived by the user as a familiar interaction.
[0195] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0196] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0197] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0198] [Second Embodiment]
[0199] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0200] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0201] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0202] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0203] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0205] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0206] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0207] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0208] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0209] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0210] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0211] This invention is a system for realizing interaction between a user and an information processing model, aiming to make AI technology more accessible using various devices. This system selects an appropriate information processing model according to the user's needs and generates and outputs a response based on it.
[0212] Specifically, users access the system using devices such as smartphones, tablets, and PCs. These devices are equipped with input methods for voice recognition and text input, and receive instructions from the user. The received instructions are then sent from the device to the server.
[0213] The server analyzes the user's input and selects an appropriate information processing model based on its content. This model selection is based on the user's requested problem or theme, and then a specific response is generated using the selected model. In response generation, natural language processing techniques are used to create information that is easy for the user to understand and has a distinct character.
[0214] The generated response is sent from the server to the terminal, which then presents it to the user. For example, if a 3D hologram monitor is used, the terminal displays the character in 3D, enabling interaction with the user. Alternatively, if a smart speaker is used, information is provided via voice. This allows the user to experience interaction visually or aurally.
[0215] As a concrete example, consider a case where a user wishes to learn a language. The user uses the device's voice input function to send an instruction to the system to begin language learning. The server analyzes this instruction and selects the appropriate information processing model for language learning. The server then generates a language lesson plan and sends its contents to the device. Since the device delivers the lesson through a character, the user can proceed with their learning under the guidance of the character.
[0216] In this way, this system enhances the quality of interaction and provides convenience tailored to the user's purpose.
[0217] The following describes the processing flow.
[0218] Step 1:
[0219] The user enters their request using voice or text via the device. This request is recognized or received by the device's input method.
[0220] Step 2:
[0221] The terminal sends user input to the server. During this process, the input data may be converted to an appropriate format.
[0222] Step 3:
[0223] The server analyzes the user's input and identifies a task based on its content. For example, if a user asks a question about language, the server identifies that category.
[0224] Step 4:
[0225] The server selects the information processing model best suited to the specified task. This model selection mechanism makes it possible to create a response that meets the user's objectives.
[0226] Step 5:
[0227] The server generates a user-appropriate response using the selected model. This response is written in natural language and includes elements that give it a character.
[0228] Step 6:
[0229] The server sends the generated response to the terminal. The response may include text, audio, and possibly visual data.
[0230] Step 7:
[0231] The terminal displays the received response to the user. For example, if a 3D hologram monitor is used, a character is projected in three dimensions, and dialogue takes place via voice.
[0232] Step 8:
[0233] The user can then interact further based on the information presented through the device. In this case, the process is repeated from step 1.
[0234] (Example 1)
[0235] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0236] With the evolution of information technology, there is a growing demand for information provision that meets diverse user needs. However, conventional systems have limitations in terms of user experience, making it difficult to respond quickly and accurately to individual needs. To address these challenges, there is a need for systems that provide intuitive and user-friendly interactions.
[0237] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0238] In this invention, the server includes terminal means for receiving information from the user, server means for selecting a model having the appropriate characteristics, and generation means for generating a response based on the information received from the user. This enables a rapid and accurate response to the diverse needs of the user.
[0239] A "terminal device" is a device that receives information from a user in the form of voice or text and transmits that information to a server.
[0240] A "server device" is a device that analyzes the received user information and selects a model with the corresponding characteristics based on the analysis results.
[0241] A "generation means" is a device that uses a selected model to generate an appropriate response based on the user's request.
[0242] A "presentation means" is a device for providing the generated response to the user visually or audibly.
[0243] A "three-dimensional display device" is a device that displays generated information in three dimensions, allowing users to experience visual interaction.
[0244] A "speech output device" is a device that converts the generated response into speech and provides it to the user audibly.
[0245] This invention relates to an interactive system for users to input information and receive appropriate responses. The system consists of terminal means, server means, generation means, and presentation means. The aim is to provide an interactive environment that can meet diverse needs.
[0246] Users access the system using terminal devices such as smartphones, tablets, or PCs. These terminal devices are equipped with voice recognition technology and text input capabilities, and can receive user input information in digital format. The terminal devices then transmit this information to the server devices via the internet.
[0247] The server uses artificial intelligence and natural language processing algorithms to analyze information received from the user. This analysis process helps understand the user's request and selects an appropriate information processing model using a generative AI model based on that understanding. The selected model is then used by the generative means to generate a specific response to the user's request.
[0248] The generated response is provided to the user via a presentation device. In the case of a three-dimensional display device, the response is displayed as a three-dimensional character, allowing the user to enjoy visual interaction. If an audio output device is used, the response is converted into audio, allowing the user to receive the information audibly.
[0249] As a concrete example, if a user wishes to learn a language, they send the instruction "I want to learn everyday conversation in English" to the system using the voice input function of their device. The server analyzes this instruction and selects the appropriate language learning information processing model. This generating AI model creates an appropriate lesson plan and provides it to the user via a presentation device. The user can then proceed with their learning in an interactive format through a three-dimensional character.
[0250] Example prompt: "If the user wants to learn English, generate a plan that provides lesson content tailored to the user's level."
[0251] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0252] Step 1:
[0253] Users access the system using devices such as smartphones and tablets, and input information via voice or text. For example, a user might input a specific request such as "I want to start an English lesson" into their device. The input instructions become input data for the system.
[0254] Step 2:
[0255] If voice input is used, the terminal uses speech recognition technology to convert the voice data into text data. If text input is used, it is retained in its original format. This text data is output as data to be sent from the terminal to the server. It is then transmitted to the server via the network.
[0256] Step 3:
[0257] The server analyzes the text data received from the terminal. Using natural language processing (NLP) techniques, it understands the user's request and identifies its intent. The user's request content obtained as a result of the analysis becomes the processing data within the server.
[0258] Step 4:
[0259] Based on the analysis results, the server selects a suitable AI model for the request. The selection process searches the database for the model that best matches the user's instructions and confirms it. This selected model is then used as input for the next processing step.
[0260] Step 5:
[0261] The server uses a selected generative AI model to generate a response to the user's request. For example, it can automatically construct an English lesson plan. The generated response becomes output data within the server and is formatted as data to be sent to the terminal.
[0262] Step 6:
[0263] The server sends the formatted response data to the terminal. Real-time communication protocols may be used here. The response data is converted into a format that can be presented to the user visually or audibly.
[0264] Step 7:
[0265] The terminal presents the received response data to the user. If the terminal is equipped with a 3D display, it uses characters to display the response in 3D. If a smart speaker is used, it provides information to the user via voice. The user can interact through this presentation and receive feedback on their requests.
[0266] (Application Example 1)
[0267] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0268] Traditional virtual stores have struggled to respond flexibly and interactively to individual user needs, resulting in decreased customer satisfaction and a lack of repeat store visits. Furthermore, there was a lack of efficient and effective technological means to provide visually and audibly rich interactions.
[0269] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0270] In this invention, the server includes a receiving device that receives information from the user, a model selection device that selects and uses a data processing model with the appropriate characteristics, and a response generation device that generates a response based on the information received from the user. This enables personalized interaction and provides a rich customer experience utilizing visual and auditory elements.
[0271] A "receiving device" is a device that receives information from a user as audio or text and converts it into a format that can be processed within the system.
[0272] A "model selection device" is a device used to select the optimal data processing model based on received information and generate a response.
[0273] A "reaction generation device" is a device that uses a selected data processing model to generate appropriate reactions according to the user's needs.
[0274] A "presentation device" is a device that effectively presents the generated response to the user, and uses visual and auditory means.
[0275] A "3D display system" is a technology that displays generated reactions in three dimensions, providing a visual impact to the user.
[0276] An "acoustic system" is a system that reproduces the generated response as sound, appealing to the user's hearing.
[0277] A "visual device" is an auxiliary device that utilizes the user's visual information to further enhance the user experience.
[0278] To implement this invention, the receiving device first acquires information from the user in the form of voice or text. The user can then initiate visual and auditory interaction by entering a virtual store through smart glasses or other visual devices. For example, the user might ask, "Which coat is suitable for winter?" within the virtual store.
[0279] This information is analyzed by a model selection device, and an appropriate data processing model is selected. The model used here is a generative AI model equipped with advanced natural language processing technology. Once this model is selected, the response generator produces responses tailored to the user's questions, creating personalized information.
[0280] The server provides this generated information to the user through a display device. The display device uses a stereoscopic display system to show three-dimensional images of related products and an audio system to provide detailed information via voice. The visual device can also highlight specific products based on the direction the user is looking. This allows for a visually and aurally enriching customer experience.
[0281] As a concrete example, the system translates a user's question into a prompt such as "Please recommend a coat for winter," and based on this prompt, a generating AI model constructs appropriate product information and recommendations, generating a response for the user. This enables flexible and interactive responses that meet the user's needs.
[0282] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0283] Step 1:
[0284] The user accesses the virtual store through the smart glasses. The terminal collects the user's voice instructions and text through the receiving device and sends them to the server. What is input here are the user's interests and questions, including, for example, "What are the recommended products for winter?"
[0285] Step 2:
[0286] Based on the received user information, the server uses the model selection device to select an appropriate data processing model. The input is the question from the user, and data operations are performed to analyze it and select a highly relevant processing model. What is output is the information about the selected model.
[0287] Step 3: [[ID=I8]]
[0288] The server uses the response generation device to generate an appropriate response for the user from the selected data processing model. In this process, a generation AI model is used, and the input is given in the form of "Please tell me the recommended products for winter" as a prompt sentence, and as a result, detailed product information and recommendation information are generated. The output is a specific response for presenting to the user.
[0289] Step 4:
[0290] The server sends the generated response to the terminal using the presentation device. The terminal uses a stereoscopic display system and an acoustic system to present the response to the user visually and audibly. The input is the generated response data from the server, and the output is a visual and audible display that the user can confirm. When a visual device is used, information is also highlighted based on the direction the user is looking.
[0291] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0292] This invention is an information processing system that combines emotion recognition technology to improve user interaction. The system aims to achieve more natural and engaging dialogue by understanding the user's emotional state and providing appropriate responses.
[0293] This system is equipped with input mechanisms for receiving voice and text input from users. Users can access the system using a smartphone, PC, or dedicated device. The input mechanisms are connected to a server and are responsible for receiving instructions and inquiries from users.
[0294] The server transmits the received data to an emotion recognition system. This system analyzes the user's voice tone, facial expressions, and linguistic characteristics to evaluate their current emotional state. This information is fed back into the information processing model selection process, and the selected model and the generated response are adjusted based on the user's emotions.
[0295] The server selects an appropriate information processing model and generates an emotion-sensitive response. This response is generated using natural language processing to make it easy for the user to understand and relate to. The generated response is then sent to the terminal.
[0296] The terminal presents the received response to the user as audio or visual. In terminals equipped with a 3D display, the character is projected three-dimensionally, as if speaking directly to the user. In terminals with sound equipment, the character's voice and sounds are played in real time.
[0297] As a specific example, when a user inputs "I'm feeling a bit down today" into the terminal by voice, the emotion recognition means determines the emotional state as "sadness" from the user's voice tone and words. Based on this information, the server selects an appropriate soothing character, generates a kind message to encourage the user, and transmits it to the terminal. The terminal displays the character and conveys the message that soothes the user by voice.
[0298] In this way, a system combined with an emotion engine can provide a more human and interactive experience to the user.
[0299] The following describes the processing flow.
[0300] Step 1:
[0301] The user inputs a request by voice or text via the terminal. The terminal receives this through the input means and prepares to transmit the data to the server.
[0302] Step 2:
[0303] The terminal transmits the user's input data to the server. This data uses a format designed to be affected by emotions.
[0304] Step 3:
[0305] The server analyzes the received data and identifies the user's emotion using the emotion recognition means. Emotions are detected from the voice tone and the choice of words.
[0306] Step 4:
[0307] Based on the result of emotion recognition, the server determines which information processing model to use. For example, if the user seems happy, it selects a model that creates a response with enjoyable content.
[0308] Step 5:
[0309] The server generates a user-specific response based on the selected information processing model. This response includes emotionally appropriate language and character tone.
[0310] Step 6:
[0311] The server sends the generated response to the terminal. The data is optimized to make the user feel more emotionally satisfied.
[0312] Step 7:
[0313] The terminal presents the received response to the user either audibly or visually. If necessary, it projects a character using a 3D display to express emotions.
[0314] Step 8:
[0315] The user can further interact based on the device's response. Emotionally responsive dialogue can continue, and the process restarts from step 1.
[0316] (Example 2)
[0317] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0318] Conventional dialogue systems have been unable to generate responses that adequately consider the user's emotional state, making it difficult to provide natural and affiliative interactions. Therefore, there is a need to improve the user experience.
[0319] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0320] In this invention, the server includes means for acquiring emotional information from the user, selection means for selecting and using the appropriate information processing technology, and response creation means for generating prompt sentences and creating responses based on the user's emotions. This makes it possible to provide natural and engaging dialogue that responds to the user's emotional state and improve the user experience.
[0321] "Emotional information" refers to data that indicates the user's emotional and mood state, and includes information such as voice characteristics and text characteristics.
[0322] "Information processing technology" is a collection of algorithms and models used for data analysis and response generation.
[0323] A "selection mechanism" is a component that has the function of selecting the optimal technology or method based on specific conditions or input data.
[0324] A "prompt" is text given to a generative model as instructions or hints to generate a target response.
[0325] A "response generation means" is a component that processes information to generate an appropriate response based on user input and emotional information.
[0326] "Presentation means" refers to a device or function for providing the generated content to the user through sight or hearing.
[0327] This invention relates to an information processing device for realizing natural and effective interaction with users. The system aims to improve the user experience by analyzing user input, recognizing emotions, and generating appropriate responses that are in line with those emotions.
[0328] Users access this system using devices such as smartphones and personal computers. These devices incorporate speech recognition software and input devices, enabling the acquisition of emotional information from speech and text. The system converts speech input into text using Google's speech-to-text API, among others. Furthermore, it utilizes natural language processing techniques to analyze the emotions expressed in the user's text input.
[0329] After receiving emotional information from the terminal, the server selects the optimal information processing technology through a pre-configured selection mechanism. The selected technology includes generative AI models, and in particular, natural language processing models utilize technologies based on open-source AI frameworks. This model generates responses to the user by leveraging prompts. An example of a prompt used is, "The user says, 'I'm feeling a little down today.' Please think of an encouraging message for them." This makes it possible to generate flexible and personalized responses that respond to the user's emotions.
[0330] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology, such as Google's text-to-speech technology, to convey the generated response to the user verbally. If visual elements are included, the content of the response can also be visually displayed on the terminal's display device. This provides the user with a natural and user-friendly interaction.
[0331] For example, if a user voice-inputs "I'm a little tired today," the system converts the voice into text and recognizes the emotion as "fatigue." The server then uses a generative AI model to generate a response such as "You did a great job today, please get some good rest tonight." The terminal reads this message aloud, speaking warmly to the user. In this way, the system responds sensitively to the user's emotions and enables appropriate dialogue.
[0332] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0333] Step 1:
[0334] Users input information via voice or text using a smartphone or PC. The device uses speech recognition software to convert the voice into text. The raw data (voice or text) is received as input, and a process is performed to convert that data into text information, resulting in the output being the converted text data.
[0335] Step 2:
[0336] The terminal sends the converted text data to the server. The server receives this text data as input and sends it to an emotion recognition module. This module uses natural language processing technology to perform emotion analysis on the text and outputs emotion labels (e.g., "joy," "sadness," "surprise," etc.).
[0337] Step 3:
[0338] The server selects the appropriate information processing technology based on the obtained emotion labels. This involves a technology selection process that includes a generative AI model. The emotion labels are used as input, and an AI model is selected. The selected AI model is obtained as output.
[0339] Step 4:
[0340] The server generates prompts using a selected generative AI model. For example, it might generate the prompt, "The user says 'I'm feeling a little down today.' Please come up with an encouraging message." The input consists of a sentiment label and a standard prompt template, and the output is a specific prompt.
[0341] Step 5:
[0342] Based on the generated prompt text, the AI model generates a response. The server manages this process, inputting the prompt text into the AI model to generate an appropriate response for the user. The output is a customized response text for the user.
[0343] Step 6:
[0344] The server sends the generated response to the terminal. The terminal receives this response and presents it to the user as speech using a speech synthesis system. It also displays it visually using a display device if necessary. It receives the generated response text as input and provides the user with speech or visual content as output.
[0345] In this way, the entire system operates in a chain reaction, providing flexible responses tailored to the user's emotions.
[0346] (Application Example 2)
[0347] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0348] In recent years, there has been a growing demand for methods to improve the user experience, but current systems struggle to accurately grasp users' emotional states and generate individualized responses accordingly. In particular, there are difficulties in real-time emotion recognition and feedback, resulting in a challenge in providing natural and user-friendly interactions.
[0349] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0350] In this invention, the server includes an information processing device that receives input from a user, a pattern selection device that selects and uses an information processing pattern having the appropriate characteristics, and a response generation device that generates a response based on the input received from the user. This makes it possible to recognize the user's emotional state and generate and present an appropriate response in real time.
[0351] An "information processing device" is a device that receives input from a user and analyzes that data.
[0352] A "pattern selection device" is a device that selects and uses an information processing pattern with the appropriate characteristics based on the received input data.
[0353] A "reaction generation device" is a device that generates an appropriate reaction based on input received from a user.
[0354] An "output device" is a device that presents the generated reaction to the user.
[0355] The term "device" refers to a machine or instrument configured to perform a specific function within a system.
[0356] An "emotion recognition device" is a device that accurately recognizes the user's emotional state and adjusts its response based on that information.
[0357] A system for implementing this invention is configured primarily to provide emotion recognition and interaction based thereon. The system includes an information processing device, a pattern selection device, a response generation device, and an output device.
[0358] The server receives audio and text input from the user via an information processing device. The received data is sent to an emotion recognition device that analyzes voice tone, facial expressions, and word characteristics. The emotion recognition device evaluates the user's emotional state based on the analysis results and provides this information to a pattern selection device. The pattern selection device then selects a specific information processing pattern, and a response generation device uses natural language processing technology to generate a clear and user-friendly response for the user.
[0359] This system also uses a three-dimensional display device and an audio device via an output device to present the generated response to the user. For example, on a terminal equipped with a three-dimensional display device, a friendly character is displayed three-dimensionally, as if speaking directly to the user.
[0360] As a concrete example, when this system is used in a physical store, service staff wear smart glasses. They can instantly recognize the customer's emotional state and provide real-time instructions on appropriate customer service methods. For example, if a customer is nervous, the system might suggest, "Please use a relaxing customer service phrase." An example of a prompt sentence for the generating AI model would be, "The customer seems anxious, how can you reassure them?"
[0361] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0362] Step 1:
[0363] The user provides input. Voice and text data from the user are transmitted to the information processing device via smart glasses or other devices. The input data reflects the user's real-time emotions.
[0364] Step 2:
[0365] The server transmits the received data to the emotion recognition device. The emotion recognition device analyzes voice tone, linguistic features, and facial expression data to evaluate the user's emotional state. The analysis results are output as data indicating the user's specific emotion.
[0366] Step 3:
[0367] The server selects an appropriate information processing pattern in the pattern selection device based on the evaluated emotional state. In the selection process, the pattern most relevant to the emotional state is selected, and the selection result is retained.
[0368] Step 4:
[0369] The server generates a response based on the selected information processing pattern. Using natural language processing techniques, it creates a user-friendly and approachable response. The generated response takes the user's emotional state into consideration.
[0370] Step 5:
[0371] The terminal presents the generated response to the user via an output device. Three-dimensional display devices and sound devices are used to present the response using three-dimensional characters and voices. The output is perceived by the user as a familiar interaction.
[0372] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0373] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0374] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0375] [Third Embodiment]
[0376] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0377] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0378] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0379] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0380] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0381] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0382] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0383] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0384] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0385] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0386] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0387] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0388] This invention is a system for realizing interaction between a user and an information processing model, aiming to make AI technology more accessible using various devices. This system selects an appropriate information processing model according to the user's needs and generates and outputs a response based on it.
[0389] Specifically, users access the system using devices such as smartphones, tablets, and PCs. These devices are equipped with input methods for voice recognition and text input, and receive instructions from the user. The received instructions are then sent from the device to the server.
[0390] The server analyzes the user's input and selects an appropriate information processing model based on its content. This model selection is based on the user's requested problem or theme, and then a specific response is generated using the selected model. In response generation, natural language processing techniques are used to create information that is easy for the user to understand and has a distinct character.
[0391] The generated response is sent from the server to the terminal, which then presents it to the user. For example, if a 3D hologram monitor is used, the terminal displays the character in 3D, enabling interaction with the user. Alternatively, if a smart speaker is used, information is provided via voice. This allows the user to experience interaction visually or aurally.
[0392] As a concrete example, consider a case where a user wishes to learn a language. The user uses the device's voice input function to send an instruction to the system to begin language learning. The server analyzes this instruction and selects the appropriate information processing model for language learning. The server then generates a language lesson plan and sends its contents to the device. Since the device delivers the lesson through a character, the user can proceed with their learning under the guidance of the character.
[0393] In this way, this system enhances the quality of interaction and provides convenience tailored to the user's purpose.
[0394] The following describes the processing flow.
[0395] Step 1:
[0396] The user enters their request using voice or text via the device. This request is recognized or received by the device's input method.
[0397] Step 2:
[0398] The terminal sends user input to the server. During this process, the input data may be converted to an appropriate format.
[0399] Step 3:
[0400] The server analyzes the user's input and identifies a task based on its content. For example, if a user asks a question about language, the server identifies that category.
[0401] Step 4:
[0402] The server selects the information processing model best suited to the specified task. This model selection mechanism makes it possible to create a response that meets the user's objectives.
[0403] Step 5:
[0404] The server generates a user-appropriate response using the selected model. This response is written in natural language and includes elements that give it a character.
[0405] Step 6:
[0406] The server sends the generated response to the terminal. The response may include text, audio, and possibly visual data.
[0407] Step 7:
[0408] The terminal displays the received response to the user. For example, if a 3D hologram monitor is used, a character is projected in three dimensions, and dialogue takes place via voice.
[0409] Step 8:
[0410] The user can then interact further based on the information presented through the device. In this case, the process is repeated from step 1.
[0411] (Example 1)
[0412] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0413] With the evolution of information technology, there is a growing demand for information provision that meets diverse user needs. However, conventional systems have limitations in terms of user experience, making it difficult to respond quickly and accurately to individual needs. To address these challenges, there is a need for systems that provide intuitive and user-friendly interactions.
[0414] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0415] In this invention, the server includes terminal means for receiving information from the user, server means for selecting a model having the appropriate characteristics, and generation means for generating a response based on the information received from the user. This enables a rapid and accurate response to the diverse needs of the user.
[0416] A "terminal device" is a device that receives information from a user in the form of voice or text and transmits that information to a server.
[0417] A "server device" is a device that analyzes the received user information and selects a model with the corresponding characteristics based on the analysis results.
[0418] A "generation means" is a device that uses a selected model to generate an appropriate response based on the user's request.
[0419] A "presentation means" is a device for providing the generated response to the user visually or audibly.
[0420] A "three-dimensional display device" is a device that displays generated information in three dimensions, allowing users to experience visual interaction.
[0421] A "speech output device" is a device that converts the generated response into speech and provides it to the user audibly.
[0422] This invention relates to an interactive system for users to input information and receive appropriate responses. The system consists of terminal means, server means, generation means, and presentation means. The aim is to provide an interactive environment that can meet diverse needs.
[0423] Users access the system using terminal devices such as smartphones, tablets, or PCs. These terminal devices are equipped with voice recognition technology and text input capabilities, and can receive user input information in digital format. The terminal devices then transmit this information to the server devices via the internet.
[0424] The server uses artificial intelligence and natural language processing algorithms to analyze information received from the user. This analysis process helps understand the user's request and selects an appropriate information processing model using a generative AI model based on that understanding. The selected model is then used by the generative means to generate a specific response to the user's request.
[0425] The generated response is provided to the user via a presentation device. In the case of a three-dimensional display device, the response is displayed as a three-dimensional character, allowing the user to enjoy visual interaction. If an audio output device is used, the response is converted into audio, allowing the user to receive the information audibly.
[0426] As a concrete example, if a user wishes to learn a language, they send the instruction "I want to learn everyday conversation in English" to the system using the voice input function of their device. The server analyzes this instruction and selects the appropriate language learning information processing model. This generating AI model creates an appropriate lesson plan and provides it to the user via a presentation device. The user can then proceed with their learning in an interactive format through a three-dimensional character.
[0427] Example prompt: "If the user wants to learn English, generate a plan that provides lesson content tailored to the user's level."
[0428] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0429] Step 1:
[0430] Users access the system using devices such as smartphones and tablets, and input information via voice or text. For example, a user might input a specific request such as "I want to start an English lesson" into their device. The input instructions become input data for the system.
[0431] Step 2:
[0432] If voice input is used, the terminal uses speech recognition technology to convert the voice data into text data. If text input is used, it is retained in its original format. This text data is output as data to be sent from the terminal to the server. It is then transmitted to the server via the network.
[0433] Step 3:
[0434] The server analyzes the text data received from the terminal. Using natural language processing (NLP) techniques, it understands the user's request and identifies its intent. The user's request content obtained as a result of the analysis becomes the processing data within the server.
[0435] Step 4:
[0436] Based on the analysis results, the server selects a suitable AI model for the request. The selection process searches the database for the model that best matches the user's instructions and confirms it. This selected model is then used as input for the next processing step.
[0437] Step 5:
[0438] The server uses a selected generative AI model to generate a response to the user's request. For example, it can automatically construct an English lesson plan. The generated response becomes output data within the server and is formatted as data to be sent to the terminal.
[0439] Step 6:
[0440] The server sends the formatted response data to the terminal. Real-time communication protocols may be used here. The response data is converted into a format that can be presented to the user visually or audibly.
[0441] Step 7:
[0442] The terminal presents the received response data to the user. If the terminal is equipped with a 3D display, it uses characters to display the response in 3D. If a smart speaker is used, it provides information to the user via voice. The user can interact through this presentation and receive feedback on their requests.
[0443] (Application Example 1)
[0444] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0445] Traditional virtual stores have struggled to respond flexibly and interactively to individual user needs, resulting in decreased customer satisfaction and a lack of repeat store visits. Furthermore, there was a lack of efficient and effective technological means to provide visually and audibly rich interactions.
[0446] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0447] In this invention, the server includes a receiving device that receives information from the user, a model selection device that selects and uses a data processing model with the appropriate characteristics, and a response generation device that generates a response based on the information received from the user. This enables personalized interaction and provides a rich customer experience utilizing visual and auditory elements.
[0448] A "receiving device" is a device that receives information from a user as audio or text and converts it into a format that can be processed within the system.
[0449] A "model selection device" is a device used to select the optimal data processing model based on received information and generate a response.
[0450] A "reaction generation device" is a device that uses a selected data processing model to generate appropriate reactions according to the user's needs.
[0451] A "presentation device" is a device that effectively presents the generated response to the user, and uses visual and auditory means.
[0452] A "3D display system" is a technology that displays generated reactions in three dimensions, providing a visual impact to the user.
[0453] An "acoustic system" is a system that reproduces the generated response as sound, appealing to the user's hearing.
[0454] A "visual device" is an auxiliary device that utilizes the user's visual information to further enhance the user experience.
[0455] To implement this invention, the receiving device first acquires information from the user in the form of voice or text. The user can then initiate visual and auditory interaction by entering a virtual store through smart glasses or other visual devices. For example, the user might ask, "Which coat is suitable for winter?" within the virtual store.
[0456] This information is analyzed by a model selection device, and an appropriate data processing model is selected. The model used here is a generative AI model equipped with advanced natural language processing technology. Once this model is selected, the response generator produces responses tailored to the user's questions, creating personalized information.
[0457] The server provides this generated information to the user through a display device. The display device uses a stereoscopic display system to show three-dimensional images of related products and an audio system to provide detailed information via voice. The visual device can also highlight specific products based on the direction the user is looking. This allows for a visually and aurally enriching customer experience.
[0458] As a concrete example, the system translates a user's question into a prompt such as "Please recommend a coat for winter," and based on this prompt, a generating AI model constructs appropriate product information and recommendations, generating a response for the user. This enables flexible and interactive responses that meet the user's needs.
[0459] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0460] Step 1:
[0461] The user accesses a virtual store through smart glasses. The device collects the user's voice commands and text messages via a receiving device and sends them to a server. The input here includes the user's interests and questions, such as "What products do you recommend for winter?"
[0462] Step 2:
[0463] The server uses a model selection device to select an appropriate data processing model based on the user information received. The input is a question from the user, which is then analyzed to perform data calculations to select the most relevant processing model. The output is information about the selected model.
[0464] Step 3:
[0465] The server uses a response generator to produce appropriate responses for the user from a selected data processing model. In this process, a generation AI model is used, and input is given in the form of a prompt such as "Please recommend some products for winter." As a result, detailed product information and recommendation information is generated. The output is a specific response to be presented to the user.
[0466] Step 4:
[0467] The server transmits the generated response to the terminal using a display device. The terminal presents the response to the user visually and audibly using a stereoscopic display system and an audio system. The input is the response data generated from the server, and the output is a visual and audible display that the user can confirm. When a visual device is used, information is also highlighted based on the direction the user is looking.
[0468] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0469] This invention is an information processing system that combines emotion recognition technology to improve user interaction. The system aims to achieve more natural and engaging dialogue by understanding the user's emotional state and providing appropriate responses.
[0470] This system is equipped with input mechanisms for receiving voice and text input from users. Users can access the system using a smartphone, PC, or dedicated device. The input mechanisms are connected to a server and are responsible for receiving instructions and inquiries from users.
[0471] The server transmits the received data to an emotion recognition system. This system analyzes the user's voice tone, facial expressions, and linguistic characteristics to evaluate their current emotional state. This information is fed back into the information processing model selection process, and the selected model and the generated response are adjusted based on the user's emotions.
[0472] The server selects an appropriate information processing model and generates an emotion-sensitive response. This response is generated using natural language processing to make it easy for the user to understand and relate to. The generated response is then sent to the terminal.
[0473] The terminal presents the received response to the user as audio or visual. In terminals equipped with a 3D display, the character is projected three-dimensionally, as if speaking directly to the user. In terminals with sound equipment, the character's voice and sounds are played in real time.
[0474] For example, if a user voice-inputs "I'm feeling a little down today" into the device, the emotion recognition system will determine the emotional state to be "sadness" based on the user's tone of voice and words. Based on this information, the server will select an appropriate comforting character, generate a gentle message to encourage the user, and send it to the device. The device will then display the character and deliver the comforting message to the user via voice.
[0475] In this way, systems that incorporate emotion engines can provide users with a more human-like and interactive experience.
[0476] The following describes the processing flow.
[0477] Step 1:
[0478] The user enters a request via voice or text through the terminal. The terminal receives this input and prepares to send the data to the server.
[0479] Step 2:
[0480] The device sends user input data to the server. This data uses a format designed to be sensitive to emotions.
[0481] Step 3:
[0482] The server analyzes the received data and uses emotion recognition to identify the user's emotions. Emotions are detected from voice tone and word choice.
[0483] Step 4:
[0484] The server decides which information processing model to use based on the emotion recognition results. For example, if the user appears happy, it will select a model that generates a cheerful response.
[0485] Step 5:
[0486] The server generates a user-specific response based on the selected information processing model. This response includes emotionally appropriate language and character tone.
[0487] Step 6:
[0488] The server sends the generated response to the terminal. The data is optimized to make the user feel more emotionally satisfied.
[0489] Step 7:
[0490] The terminal presents the received response to the user either audibly or visually. If necessary, it projects a character using a 3D display to express emotions.
[0491] Step 8:
[0492] The user can further interact based on the device's response. Emotionally responsive dialogue can continue, and the process restarts from step 1.
[0493] (Example 2)
[0494] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0495] Conventional dialogue systems have been unable to generate responses that adequately consider the user's emotional state, making it difficult to provide natural and affiliative interactions. Therefore, there is a need to improve the user experience.
[0496] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0497] In this invention, the server includes means for acquiring emotional information from the user, selection means for selecting and using the appropriate information processing technology, and response creation means for generating prompt sentences and creating responses based on the user's emotions. This makes it possible to provide natural and engaging dialogue that responds to the user's emotional state and improve the user experience.
[0498] "Emotional information" refers to data that indicates the user's emotional and mood state, and includes information such as voice characteristics and text characteristics.
[0499] "Information processing technology" is a collection of algorithms and models used for data analysis and response generation.
[0500] A "selection mechanism" is a component that has the function of selecting the optimal technology or method based on specific conditions or input data.
[0501] A "prompt" is text given to a generative model as instructions or hints to generate a target response.
[0502] A "response generation means" is a component that processes information to generate an appropriate response based on user input and emotional information.
[0503] "Presentation means" refers to a device or function for providing the generated content to the user through sight or hearing.
[0504] This invention relates to an information processing device for realizing natural and effective interaction with users. The system aims to improve the user experience by analyzing user input, recognizing emotions, and generating appropriate responses that are in line with those emotions.
[0505] Users access this system using devices such as smartphones and personal computers. These devices incorporate speech recognition software and input devices, enabling the acquisition of emotional information from speech and text. The system converts speech input into text using Google's speech-to-text API, among others. Furthermore, it utilizes natural language processing techniques to analyze the emotions expressed in the user's text input.
[0506] After receiving emotional information from the terminal, the server selects the optimal information processing technology through a pre-configured selection mechanism. The selected technology includes generative AI models, and in particular, natural language processing models utilize technologies based on open-source AI frameworks. This model generates responses to the user by leveraging prompts. An example of a prompt used is, "The user says, 'I'm feeling a little down today.' Please think of an encouraging message for them." This makes it possible to generate flexible and personalized responses that respond to the user's emotions.
[0507] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology, such as Google's text-to-speech technology, to convey the generated response to the user verbally. If visual elements are included, the content of the response can also be visually displayed on the terminal's display device. This provides the user with a natural and user-friendly interaction.
[0508] For example, if a user voice-inputs "I'm a little tired today," the system converts the voice into text and recognizes the emotion as "fatigue." The server then uses a generative AI model to generate a response such as "You did a great job today, please get some good rest tonight." The terminal reads this message aloud, speaking warmly to the user. In this way, the system responds sensitively to the user's emotions and enables appropriate dialogue.
[0509] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0510] Step 1:
[0511] Users input information via voice or text using a smartphone or PC. The device uses speech recognition software to convert the voice into text. The raw data (voice or text) is received as input, and a process is performed to convert that data into text information, resulting in the output being the converted text data.
[0512] Step 2:
[0513] The terminal sends the converted text data to the server. The server receives this text data as input and sends it to an emotion recognition module. This module uses natural language processing technology to perform emotion analysis on the text and outputs emotion labels (e.g., "joy," "sadness," "surprise," etc.).
[0514] Step 3:
[0515] The server selects the appropriate information processing technology based on the obtained emotion labels. This involves a technology selection process that includes a generative AI model. The emotion labels are used as input, and an AI model is selected. The selected AI model is obtained as output.
[0516] Step 4:
[0517] The server generates prompts using a selected generative AI model. For example, it might generate the prompt, "The user says 'I'm feeling a little down today.' Please come up with an encouraging message." The input consists of a sentiment label and a standard prompt template, and the output is a specific prompt.
[0518] Step 5:
[0519] Based on the generated prompt text, the AI model generates a response. The server manages this process, inputting the prompt text into the AI model to generate an appropriate response for the user. The output is a customized response text for the user.
[0520] Step 6:
[0521] The server sends the generated response to the terminal. The terminal receives this response and presents it to the user as speech using a speech synthesis system. It also displays it visually using a display device if necessary. It receives the generated response text as input and provides the user with speech or visual content as output.
[0522] In this way, the entire system operates in a chain reaction, providing flexible responses tailored to the user's emotions.
[0523] (Application Example 2)
[0524] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0525] In recent years, there has been a growing demand for methods to improve the user experience, but current systems struggle to accurately grasp users' emotional states and generate individualized responses accordingly. In particular, there are difficulties in real-time emotion recognition and feedback, resulting in a challenge in providing natural and user-friendly interactions.
[0526] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0527] In this invention, the server includes an information processing device that receives input from a user, a pattern selection device that selects and uses an information processing pattern having the appropriate characteristics, and a response generation device that generates a response based on the input received from the user. This makes it possible to recognize the user's emotional state and generate and present an appropriate response in real time.
[0528] An "information processing device" is a device that receives input from a user and analyzes that data.
[0529] A "pattern selection device" is a device that selects and uses an information processing pattern with the appropriate characteristics based on the received input data.
[0530] A "reaction generation device" is a device that generates an appropriate reaction based on input received from a user.
[0531] An "output device" is a device that presents the generated reaction to the user.
[0532] The term "device" refers to a machine or instrument configured to perform a specific function within a system.
[0533] An "emotion recognition device" is a device that accurately recognizes the user's emotional state and adjusts its response based on that information.
[0534] A system for implementing this invention is configured primarily to provide emotion recognition and interaction based thereon. The system includes an information processing device, a pattern selection device, a response generation device, and an output device.
[0535] The server receives audio and text input from the user via an information processing device. The received data is sent to an emotion recognition device that analyzes voice tone, facial expressions, and word characteristics. The emotion recognition device evaluates the user's emotional state based on the analysis results and provides this information to a pattern selection device. The pattern selection device then selects a specific information processing pattern, and a response generation device uses natural language processing technology to generate a clear and user-friendly response for the user.
[0536] This system also uses a three-dimensional display device and an audio device via an output device to present the generated response to the user. For example, on a terminal equipped with a three-dimensional display device, a friendly character is displayed three-dimensionally, as if speaking directly to the user.
[0537] As a concrete example, when this system is used in a physical store, service staff wear smart glasses. They can instantly recognize the customer's emotional state and provide real-time instructions on appropriate customer service methods. For example, if a customer is nervous, the system might suggest, "Please use a relaxing customer service phrase." An example of a prompt sentence for the generating AI model would be, "The customer seems anxious, how can you reassure them?"
[0538] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0539] Step 1:
[0540] The user provides input. Voice and text data from the user are transmitted to the information processing device via smart glasses or other devices. The input data reflects the user's real-time emotions.
[0541] Step 2:
[0542] The server transmits the received data to the emotion recognition device. The emotion recognition device analyzes voice tone, linguistic features, and facial expression data to evaluate the user's emotional state. The analysis results are output as data indicating the user's specific emotion.
[0543] Step 3:
[0544] The server selects an appropriate information processing pattern in the pattern selection device based on the evaluated emotional state. In the selection process, the pattern most relevant to the emotional state is selected, and the selection result is retained.
[0545] Step 4:
[0546] The server generates a response based on the selected information processing pattern. Using natural language processing techniques, it creates a user-friendly and approachable response. The generated response takes the user's emotional state into consideration.
[0547] Step 5:
[0548] The terminal presents the generated response to the user via an output device. Three-dimensional display devices and sound devices are used to present the response using three-dimensional characters and voices. The output is perceived by the user as a familiar interaction.
[0549] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0550] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0551] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0552] [Fourth Embodiment]
[0553] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0554] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0555] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0556] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0557] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0558] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0559] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0560] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0561] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0562] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0563] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0564] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0565] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0566] This invention is a system for realizing interaction between a user and an information processing model, aiming to make AI technology more accessible using various devices. This system selects an appropriate information processing model according to the user's needs and generates and outputs a response based on it.
[0567] Specifically, users access the system using devices such as smartphones, tablets, and PCs. These devices are equipped with input methods for voice recognition and text input, and receive instructions from the user. The received instructions are then sent from the device to the server.
[0568] The server analyzes the user's input and selects an appropriate information processing model based on its content. This model selection is based on the user's requested problem or theme, and then a specific response is generated using the selected model. In response generation, natural language processing techniques are used to create information that is easy for the user to understand and has a distinct character.
[0569] The generated response is sent from the server to the terminal, which then presents it to the user. For example, if a 3D hologram monitor is used, the terminal displays the character in 3D, enabling interaction with the user. Alternatively, if a smart speaker is used, information is provided via voice. This allows the user to experience interaction visually or aurally.
[0570] As a concrete example, consider a case where a user wishes to learn a language. The user uses the device's voice input function to send an instruction to the system to begin language learning. The server analyzes this instruction and selects the appropriate information processing model for language learning. The server then generates a language lesson plan and sends its contents to the device. Since the device delivers the lesson through a character, the user can proceed with their learning under the guidance of the character.
[0571] In this way, this system enhances the quality of interaction and provides convenience tailored to the user's purpose.
[0572] The following describes the processing flow.
[0573] Step 1:
[0574] The user enters their request using voice or text via the device. This request is recognized or received by the device's input method.
[0575] Step 2:
[0576] The terminal sends user input to the server. During this process, the input data may be converted to an appropriate format.
[0577] Step 3:
[0578] The server analyzes the user's input and identifies a task based on its content. For example, if a user asks a question about language, the server identifies that category.
[0579] Step 4:
[0580] The server selects the information processing model best suited to the specified task. This model selection mechanism makes it possible to create a response that meets the user's objectives.
[0581] Step 5:
[0582] The server generates a user-appropriate response using the selected model. This response is written in natural language and includes elements that give it a character.
[0583] Step 6:
[0584] The server sends the generated response to the terminal. The response may include text, audio, and possibly visual data.
[0585] Step 7:
[0586] The terminal displays the received response to the user. For example, if a 3D hologram monitor is used, a character is projected in three dimensions, and dialogue takes place via voice.
[0587] Step 8:
[0588] The user can then interact further based on the information presented through the device. In this case, the process is repeated from step 1.
[0589] (Example 1)
[0590] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0591] With the evolution of information technology, there is a growing demand for information provision that meets diverse user needs. However, conventional systems have limitations in terms of user experience, making it difficult to respond quickly and accurately to individual needs. To address these challenges, there is a need for systems that provide intuitive and user-friendly interactions.
[0592] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0593] In this invention, the server includes terminal means for receiving information from the user, server means for selecting a model having the appropriate characteristics, and generation means for generating a response based on the information received from the user. This enables a rapid and accurate response to the diverse needs of the user.
[0594] A "terminal device" is a device that receives information from a user in the form of voice or text and transmits that information to a server.
[0595] A "server device" is a device that analyzes the received user information and selects a model with the corresponding characteristics based on the analysis results.
[0596] A "generation means" is a device that uses a selected model to generate an appropriate response based on the user's request.
[0597] A "presentation means" is a device for providing the generated response to the user visually or audibly.
[0598] A "three-dimensional display device" is a device that displays generated information in three dimensions, allowing users to experience visual interaction.
[0599] A "speech output device" is a device that converts the generated response into speech and provides it to the user audibly.
[0600] This invention relates to an interactive system for users to input information and receive appropriate responses. The system consists of terminal means, server means, generation means, and presentation means. The aim is to provide an interactive environment that can meet diverse needs.
[0601] Users access the system using terminal devices such as smartphones, tablets, or PCs. These terminal devices are equipped with voice recognition technology and text input capabilities, and can receive user input information in digital format. The terminal devices then transmit this information to the server devices via the internet.
[0602] The server uses artificial intelligence and natural language processing algorithms to analyze information received from the user. This analysis process helps understand the user's request and selects an appropriate information processing model using a generative AI model based on that understanding. The selected model is then used by the generative means to generate a specific response to the user's request.
[0603] The generated response is provided to the user via a presentation device. In the case of a three-dimensional display device, the response is displayed as a three-dimensional character, allowing the user to enjoy visual interaction. If an audio output device is used, the response is converted into audio, allowing the user to receive the information audibly.
[0604] As a concrete example, if a user wishes to learn a language, they send the instruction "I want to learn everyday conversation in English" to the system using the voice input function of their device. The server analyzes this instruction and selects the appropriate language learning information processing model. This generating AI model creates an appropriate lesson plan and provides it to the user via a presentation device. The user can then proceed with their learning in an interactive format through a three-dimensional character.
[0605] Example prompt: "If the user wants to learn English, generate a plan that provides lesson content tailored to the user's level."
[0606] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0607] Step 1:
[0608] Users access the system using devices such as smartphones and tablets, and input information via voice or text. For example, a user might input a specific request such as "I want to start an English lesson" into their device. The input instructions become input data for the system.
[0609] Step 2:
[0610] If voice input is used, the terminal uses speech recognition technology to convert the voice data into text data. If text input is used, it is retained in its original format. This text data is output as data to be sent from the terminal to the server. It is then transmitted to the server via the network.
[0611] Step 3:
[0612] The server analyzes the text data received from the terminal. Using natural language processing (NLP) techniques, it understands the user's request and identifies its intent. The user's request content obtained as a result of the analysis becomes the processing data within the server.
[0613] Step 4:
[0614] Based on the analysis results, the server selects a suitable AI model for the request. The selection process searches the database for the model that best matches the user's instructions and confirms it. This selected model is then used as input for the next processing step.
[0615] Step 5:
[0616] The server uses a selected generative AI model to generate a response to the user's request. For example, it can automatically construct an English lesson plan. The generated response becomes output data within the server and is formatted as data to be sent to the terminal.
[0617] Step 6:
[0618] The server sends the formatted response data to the terminal. Real-time communication protocols may be used here. The response data is converted into a format that can be presented to the user visually or audibly.
[0619] Step 7:
[0620] The terminal presents the received response data to the user. If the terminal is equipped with a 3D display, it uses characters to display the response in 3D. If a smart speaker is used, it provides information to the user via voice. The user can interact through this presentation and receive feedback on their requests.
[0621] (Application Example 1)
[0622] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0623] Traditional virtual stores have struggled to respond flexibly and interactively to individual user needs, resulting in decreased customer satisfaction and a lack of repeat store visits. Furthermore, there was a lack of efficient and effective technological means to provide visually and audibly rich interactions.
[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0625] In this invention, the server includes a receiving device that receives information from the user, a model selection device that selects and uses a data processing model with the appropriate characteristics, and a response generation device that generates a response based on the information received from the user. This enables personalized interaction and provides a rich customer experience utilizing visual and auditory elements.
[0626] A "receiving device" is a device that receives information from a user as audio or text and converts it into a format that can be processed within the system.
[0627] A "model selection device" is a device used to select the optimal data processing model based on received information and generate a response.
[0628] A "reaction generation device" is a device that uses a selected data processing model to generate appropriate reactions according to the user's needs.
[0629] A "presentation device" is a device that effectively presents the generated response to the user, and uses visual and auditory means.
[0630] A "3D display system" is a technology that displays generated reactions in three dimensions, providing a visual impact to the user.
[0631] An "acoustic system" is a system that reproduces the generated response as sound, appealing to the user's hearing.
[0632] A "visual device" is an auxiliary device that utilizes the user's visual information to further enhance the user experience.
[0633] To implement this invention, the receiving device first acquires information from the user in the form of voice or text. The user can then initiate visual and auditory interaction by entering a virtual store through smart glasses or other visual devices. For example, the user might ask, "Which coat is suitable for winter?" within the virtual store.
[0634] This information is analyzed by a model selection device, and an appropriate data processing model is selected. The model used here is a generative AI model equipped with advanced natural language processing technology. Once this model is selected, the response generator produces responses tailored to the user's questions, creating personalized information.
[0635] The server provides this generated information to the user through a display device. The display device uses a stereoscopic display system to show three-dimensional images of related products and an audio system to provide detailed information via voice. The visual device can also highlight specific products based on the direction the user is looking. This allows for a visually and aurally enriching customer experience.
[0636] As a concrete example, the system translates a user's question into a prompt such as "Please recommend a coat for winter," and based on this prompt, a generating AI model constructs appropriate product information and recommendations, generating a response for the user. This enables flexible and interactive responses that meet the user's needs.
[0637] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0638] Step 1:
[0639] The user accesses a virtual store through smart glasses. The device collects the user's voice commands and text messages via a receiving device and sends them to a server. The input here includes the user's interests and questions, such as "What products do you recommend for winter?"
[0640] Step 2:
[0641] The server uses a model selection device to select an appropriate data processing model based on the user information received. The input is a question from the user, which is then analyzed to perform data calculations to select the most relevant processing model. The output is information about the selected model.
[0642] Step 3:
[0643] The server uses a response generator to produce appropriate responses for the user from a selected data processing model. In this process, a generation AI model is used, and input is given in the form of a prompt such as "Please recommend some products for winter." As a result, detailed product information and recommendation information is generated. The output is a specific response to be presented to the user.
[0644] Step 4:
[0645] The server transmits the generated response to the terminal using a display device. The terminal presents the response to the user visually and audibly using a stereoscopic display system and an audio system. The input is the response data generated from the server, and the output is a visual and audible display that the user can confirm. When a visual device is used, information is also highlighted based on the direction the user is looking.
[0646] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0647] This invention is an information processing system that combines emotion recognition technology to improve user interaction. The system aims to achieve more natural and engaging dialogue by understanding the user's emotional state and providing appropriate responses.
[0648] This system is equipped with input mechanisms for receiving voice and text input from users. Users can access the system using a smartphone, PC, or dedicated device. The input mechanisms are connected to a server and are responsible for receiving instructions and inquiries from users.
[0649] The server transmits the received data to an emotion recognition system. This system analyzes the user's voice tone, facial expressions, and linguistic characteristics to evaluate their current emotional state. This information is fed back into the information processing model selection process, and the selected model and the generated response are adjusted based on the user's emotions.
[0650] The server selects an appropriate information processing model and generates an emotion-sensitive response. This response is generated using natural language processing to make it easy for the user to understand and relate to. The generated response is then sent to the terminal.
[0651] The terminal presents the received response to the user as audio or visual. In terminals equipped with a 3D display, the character is projected three-dimensionally, as if speaking directly to the user. In terminals with sound equipment, the character's voice and sounds are played in real time.
[0652] For example, if a user voice-inputs "I'm feeling a little down today" into the device, the emotion recognition system will determine the emotional state to be "sadness" based on the user's tone of voice and words. Based on this information, the server will select an appropriate comforting character, generate a gentle message to encourage the user, and send it to the device. The device will then display the character and deliver the comforting message to the user via voice.
[0653] In this way, systems that incorporate emotion engines can provide users with a more human-like and interactive experience.
[0654] The following describes the processing flow.
[0655] Step 1:
[0656] The user enters a request via voice or text through the terminal. The terminal receives this input and prepares to send the data to the server.
[0657] Step 2:
[0658] The device sends user input data to the server. This data uses a format designed to be sensitive to emotions.
[0659] Step 3:
[0660] The server analyzes the received data and uses emotion recognition to identify the user's emotions. Emotions are detected from voice tone and word choice.
[0661] Step 4:
[0662] The server decides which information processing model to use based on the emotion recognition results. For example, if the user appears happy, it will select a model that generates a cheerful response.
[0663] Step 5:
[0664] The server generates a user-specific response based on the selected information processing model. This response includes emotionally appropriate language and character tone.
[0665] Step 6:
[0666] The server sends the generated response to the terminal. The data is optimized to make the user feel more emotionally satisfied.
[0667] Step 7:
[0668] The terminal presents the received response to the user either audibly or visually. If necessary, it projects a character using a 3D display to express emotions.
[0669] Step 8:
[0670] The user can further interact based on the device's response. Emotionally responsive dialogue can continue, and the process restarts from step 1.
[0671] (Example 2)
[0672] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0673] Conventional dialogue systems have been unable to generate responses that adequately consider the user's emotional state, making it difficult to provide natural and affiliative interactions. Therefore, there is a need to improve the user experience.
[0674] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0675] In this invention, the server includes means for acquiring emotional information from the user, selection means for selecting and using the appropriate information processing technology, and response creation means for generating prompt sentences and creating responses based on the user's emotions. This makes it possible to provide natural and engaging dialogue that responds to the user's emotional state and improve the user experience.
[0676] "Emotional information" refers to data that indicates the user's emotional and mood state, and includes information such as voice characteristics and text characteristics.
[0677] "Information processing technology" is a collection of algorithms and models used for data analysis and response generation.
[0678] A "selection mechanism" is a component that has the function of selecting the optimal technology or method based on specific conditions or input data.
[0679] A "prompt" is text given to a generative model as instructions or hints to generate a target response.
[0680] A "response generation means" is a component that processes information to generate an appropriate response based on user input and emotional information.
[0681] "Presentation means" refers to a device or function for providing the generated content to the user through sight or hearing.
[0682] This invention relates to an information processing device for realizing natural and effective interaction with users. The system aims to improve the user experience by analyzing user input, recognizing emotions, and generating appropriate responses that are in line with those emotions.
[0683] Users access this system using devices such as smartphones and personal computers. These devices incorporate speech recognition software and input devices, enabling the acquisition of emotional information from speech and text. The system converts speech input into text using Google's speech-to-text API, among others. Furthermore, it utilizes natural language processing techniques to analyze the emotions expressed in the user's text input.
[0684] After receiving emotional information from the terminal, the server selects the optimal information processing technology through a pre-configured selection mechanism. The selected technology includes generative AI models, and in particular, natural language processing models utilize technologies based on open-source AI frameworks. This model generates responses to the user by leveraging prompts. An example of a prompt used is, "The user says, 'I'm feeling a little down today.' Please think of an encouraging message for them." This makes it possible to generate flexible and personalized responses that respond to the user's emotions.
[0685] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology, such as Google's text-to-speech technology, to convey the generated response to the user verbally. If visual elements are included, the content of the response can also be visually displayed on the terminal's display device. This provides the user with a natural and user-friendly interaction.
[0686] For example, if a user voice-inputs "I'm a little tired today," the system converts the voice into text and recognizes the emotion as "fatigue." The server then uses a generative AI model to generate a response such as "You did a great job today, please get some good rest tonight." The terminal reads this message aloud, speaking warmly to the user. In this way, the system responds sensitively to the user's emotions and enables appropriate dialogue.
[0687] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0688] Step 1:
[0689] Users input information via voice or text using a smartphone or PC. The device uses speech recognition software to convert the voice into text. The raw data (voice or text) is received as input, and a process is performed to convert that data into text information, resulting in the output being the converted text data.
[0690] Step 2:
[0691] The terminal sends the converted text data to the server. The server receives this text data as input and sends it to an emotion recognition module. This module uses natural language processing technology to perform emotion analysis on the text and outputs emotion labels (e.g., "joy," "sadness," "surprise," etc.).
[0692] Step 3:
[0693] The server selects the appropriate information processing technology based on the obtained emotion labels. This involves a technology selection process that includes a generative AI model. The emotion labels are used as input, and an AI model is selected. The selected AI model is obtained as output.
[0694] Step 4:
[0695] The server generates prompts using a selected generative AI model. For example, it might generate the prompt, "The user says 'I'm feeling a little down today.' Please come up with an encouraging message." The input consists of a sentiment label and a standard prompt template, and the output is a specific prompt.
[0696] Step 5:
[0697] Based on the generated prompt text, the AI model generates a response. The server manages this process, inputting the prompt text into the AI model to generate an appropriate response for the user. The output is a customized response text for the user.
[0698] Step 6:
[0699] The server sends the generated response to the terminal. The terminal receives this response and presents it to the user as speech using a speech synthesis system. It also displays it visually using a display device if necessary. It receives the generated response text as input and provides the user with speech or visual content as output.
[0700] In this way, the entire system operates in a chain reaction, providing flexible responses tailored to the user's emotions.
[0701] (Application Example 2)
[0702] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0703] In recent years, there has been a growing demand for methods to improve the user experience, but current systems struggle to accurately grasp users' emotional states and generate individualized responses accordingly. In particular, there are difficulties in real-time emotion recognition and feedback, resulting in a challenge in providing natural and user-friendly interactions.
[0704] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0705] In this invention, the server includes an information processing device that receives input from a user, a pattern selection device that selects and uses an information processing pattern having the appropriate characteristics, and a response generation device that generates a response based on the input received from the user. This makes it possible to recognize the user's emotional state and generate and present an appropriate response in real time.
[0706] An "information processing device" is a device that receives input from a user and analyzes that data.
[0707] A "pattern selection device" is a device that selects and uses an information processing pattern with the appropriate characteristics based on the received input data.
[0708] A "reaction generation device" is a device that generates an appropriate reaction based on input received from a user.
[0709] An "output device" is a device that presents the generated reaction to the user.
[0710] The term "device" refers to a machine or instrument configured to perform a specific function within a system.
[0711] An "emotion recognition device" is a device that accurately recognizes the user's emotional state and adjusts its response based on that information.
[0712] A system for implementing this invention is configured primarily to provide emotion recognition and interaction based thereon. The system includes an information processing device, a pattern selection device, a response generation device, and an output device.
[0713] The server receives audio and text input from the user via an information processing device. The received data is sent to an emotion recognition device that analyzes voice tone, facial expressions, and word characteristics. The emotion recognition device evaluates the user's emotional state based on the analysis results and provides this information to a pattern selection device. The pattern selection device then selects a specific information processing pattern, and a response generation device uses natural language processing technology to generate a clear and user-friendly response for the user.
[0714] This system also uses a three-dimensional display device and an audio device via an output device to present the generated response to the user. For example, on a terminal equipped with a three-dimensional display device, a friendly character is displayed three-dimensionally, as if speaking directly to the user.
[0715] As a concrete example, when this system is used in a physical store, service staff wear smart glasses. They can instantly recognize the customer's emotional state and provide real-time instructions on appropriate customer service methods. For example, if a customer is nervous, the system might suggest, "Please use a relaxing customer service phrase." An example of a prompt sentence for the generating AI model would be, "The customer seems anxious, how can you reassure them?"
[0716] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0717] Step 1:
[0718] The user provides input. Voice and text data from the user are transmitted to the information processing device via smart glasses or other devices. The input data reflects the user's real-time emotions.
[0719] Step 2:
[0720] The server transmits the received data to the emotion recognition device. The emotion recognition device analyzes voice tone, linguistic features, and facial expression data to evaluate the user's emotional state. The analysis results are output as data indicating the user's specific emotion.
[0721] Step 3:
[0722] The server selects an appropriate information processing pattern in the pattern selection device based on the evaluated emotional state. In the selection process, the pattern most relevant to the emotional state is selected, and the selection result is retained.
[0723] Step 4:
[0724] The server generates a response based on the selected information processing pattern. Using natural language processing techniques, it creates a user-friendly and approachable response. The generated response takes the user's emotional state into consideration.
[0725] Step 5:
[0726] The terminal presents the generated response to the user via an output device. Three-dimensional display devices and sound devices are used to present the response using three-dimensional characters and voices. The output is perceived by the user as a familiar interaction.
[0727] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0728] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0729] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0730] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0731] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0732] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0733] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0734] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0735] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0736] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0737] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0738] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0739] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0740] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0741] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0742] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0743] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0744] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0745] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0746] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0747] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0748] The following is further disclosed regarding the embodiments described above.
[0749] (Claim 1)
[0750] An input means for receiving input from the user,
[0751] A model selection means for selecting an information processing model with the relevant characteristics and using it,
[0752] A response generation means that generates a response based on input received from the user,
[0753] An output means for outputting the generated response,
[0754] A system that includes this.
[0755] (Claim 2)
[0756] The system according to claim 1, wherein the input means is capable of receiving both voice and text.
[0757] (Claim 3)
[0758] The system according to claim 1, wherein the output means presents the response of the information processing model to the user using a stereoscopic display device or an acoustic device.
[0759] "Example 1"
[0760] (Claim 1)
[0761] A terminal means for receiving information from the user,
[0762] A server means for selecting a model with the relevant characteristics,
[0763] A generation means that generates a response based on information received from the user,
[0764] A presentation means for providing the generated response to the user,
[0765] A system that includes this.
[0766] (Claim 2)
[0767] The system according to claim 1, wherein the terminal means is capable of receiving at least one of voice input and text input.
[0768] (Claim 3)
[0769] The system according to claim 1, wherein the presentation means presents the response of the generation means to the user using a three-dimensional display device or an audio output device.
[0770] "Application Example 1"
[0771] (Claim 1)
[0772] A receiving device that receives information from the user,
[0773] Select a data processing model with the appropriate characteristics, and use the model selection device.
[0774] A reaction generating device that generates a reaction based on information received from the user,
[0775] A presentation device that presents the generated response through a perceptual device,
[0776] A system that includes this.
[0777] (Claim 2)
[0778] The system according to claim 1, wherein the receiving device is capable of receiving both voice and text.
[0779] (Claim 3)
[0780] The system according to claim 1, wherein the presentation device presents the response of the data processing model to the user using a stereoscopic display system or an acoustic system, and further utilizes information based on the user's vision using a visual device.
[0781] "Example 2 of combining an emotion engine"
[0782] (Claim 1)
[0783] Means for obtaining emotional information from users,
[0784] Select the appropriate information processing technology and the means of selection for its use.
[0785] A response generation means that generates prompt text and creates a response based on the user's emotions,
[0786] A means of presenting the created response,
[0787] Information processing device including
[0788] (Claim 2)
[0789] The information processing apparatus according to claim 1, wherein the emotion information acquisition means is capable of acquiring both voice characteristics and text characteristics.
[0790] (Claim 3)
[0791] The information processing apparatus according to claim 1, wherein the presentation means presents the response to the user using a stereoscopic image device or an audio playback device.
[0792] "Application example 2 when combining with an emotional engine"
[0793] (Claim 1)
[0794] An information processing device that receives input from a user,
[0795] A pattern selection device selects and uses an information processing pattern that has the appropriate characteristics,
[0796] A reaction generating device that generates a reaction based on input received from a user,
[0797] An output device that outputs the generated reaction,
[0798] An emotion recognition device that recognizes the user's emotional state and uses that information to generate an appropriate response,
[0799] A device that includes this.
[0800] (Claim 2)
[0801] The information processing device according to claim 1, which is capable of receiving both sound and text.
[0802] (Claim 3)
[0803] The apparatus according to claim 1, wherein the output device presents the response of the information processing pattern to the user using a three-dimensional display device or an acoustic device. [Explanation of symbols]
[0804] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An input means for receiving input from the user, A model selection means for selecting an information processing model with the relevant characteristics and using it, A response generation means that generates a response based on input received from the user, An output means for outputting the generated response, A system that includes this.
2. The system according to claim 1, wherein the input means is capable of receiving both voice and text.
3. The system according to claim 1, wherein the output means presents the response of the information processing model to the user using a stereoscopic display device or an acoustic device.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A