system

The system addresses the limitations of conventional robot demonstrations by integrating voice input, text conversion, and content generation to provide interactive and engaging product information, enhancing user experience and purchase intent.

JP2026060622APending Publication Date: 2026-04-08SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Conventional product demonstrations by robots are limited, failing to effectively convey product features and benefits, particularly when relying solely on voice, leading to insufficient user engagement and reduced purchase intent due to the lack of visual and auditory integration.

Method used

A system that captures user voice inputs, converts them to text, analyzes the text to generate relevant visual, audio, and video content, and presents this content visually and audibly through a robot, allowing for interactive and responsive demonstrations.

Benefits of technology

Enhances user experience by providing a rich, interactive demonstration that effectively communicates product features and benefits, thereby increasing user interest and purchase intent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026060622000001_ABST
    Figure 2026060622000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for inputting the user's voice, A means for converting the audio into text data, A means for analyzing the text data and generating related visual, audio, video, and other content, Means for displaying and playing the generated content, A system including means for presenting the generated content to a user visually and aurally.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional product demonstration by a robot, the method is limited, and there is a problem that it is difficult to effectively convey the features and advantages of the product to the user. In particular, when explaining only by voice, visual information is insufficient, and a rich user experience cannot be provided. As a result, it has been difficult to sufficiently arouse the user's interest and purchase desire.

Means for Solving the Problems

[0005] The present invention provides a robot system that includes means for inputting a user's voice, means for converting the voice into text data, means for analyzing the text data and generating related content such as visuals, audio, and video, means for displaying and playing the generated content, and means for presenting the generated content to the user visually and audibly.

[0006] By recognizing and analyzing user voice commands to generate relevant content and presenting it through the display and speakers, the product's features and benefits can be communicated visually and intuitively. Furthermore, by responding to user questions and requests in real time, a more interactive demonstration can be achieved, thereby increasing user purchasing intent.

[0007] A "voice input device" is a device that has the function of capturing voice commands from a user and processing them as electronic data.

[0008] "Text conversion means" refers to a device or software for converting audio data into text data.

[0009] "Data analysis means" refers to a device or software that analyzes received text data to understand the user's requests and intentions.

[0010] "Content generation means" refers to a device or software that generates content in various formats, such as visuals, audio, and video, based on analyzed information.

[0011] "Display means" refers to a display device for visually presenting generated visual content to the user.

[0012] A "playback device" is a speaker device that presents the generated audio content to the user audibly.

[0013] A "user" is an individual or group that inputs voice commands to the system and views the generated content. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice instructions via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those instructions.

[0036] System Configuration

[0037] Hardware configuration

[0038] This system consists of the following main components:

[0039] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[0040] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[0041] Software Configuration

[0042] The following software and algorithms operate between the terminal and the server:

[0043] Speech recognition software: Converts voice input from the user into text data.

[0044] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[0045] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[0046] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[0047] System operation

[0048] User voice input

[0049] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[0050] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[0051] 3. The terminal sends text data to the server.

[0052] Data analysis and content generation

[0053] 4. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[0054] 5. The server uses the generated AI model to produce relevant visual, audio, and video content.

[0055] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[0056] Content display and response

[0057] 6. The device receives the generated content.

[0058] 7. Display visual content on the device's screen and play audio content using the speaker.

[0059] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[0060] 8. If the user requests further information (e.g., "How long does the battery last?"), they input voice again, and the process is repeated.

[0061] This allows users to visually and audibly experience the features and benefits of a product through a robot, thereby increasing their desire to purchase. The present invention is a system that, through this interactive process, solves the problems of conventional, limited demonstration methods and enables the effective provision of product information.

[0062] The following describes the processing flow.

[0063] Step 1:

[0064] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[0065] Step 2:

[0066] The terminal (robot) uses its built-in microphone to capture the user's voice.

[0067] Step 3:

[0068] The device uses speech recognition software to convert the captured audio into text data.

[0069] Step 4:

[0070] The terminal sends the converted text data to the server.

[0071] Step 5:

[0072] The server runs a natural language processing (NLP) engine to analyze the text data it receives.

[0073] Step 6:

[0074] The server uses an NLP engine to understand the user's request and identify the necessary information (smartphone features).

[0075] Step 7:

[0076] The server uses a generated AI model to create content such as visuals, audio, and video based on identified information.

[0077] Step 8:

[0078] The server sends the generated content to the terminal.

[0079] Step 9:

[0080] The device displays the received content on its screen and plays the audio through its speaker.

[0081] Step 10:

[0082] The user visually and audibly confirms the content that has been displayed and played.

[0083] Step 11:

[0084] If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[0085] Step 12:

[0086] The device captures the user's voice again, converts it into text data, and sends it to the server.

[0087] Step 13:

[0088] The server analyzes the new text data, generates the necessary new content, and sends it to the terminal.

[0089] Step 14:

[0090] The device displays new content on its screen and plays audio through its speaker.

[0091] Step 15:

[0092] When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[0093] (Example 1)

[0094] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] Current product demonstration systems lack the interactivity necessary for users to effectively understand product information. Traditional systems only provide pre-prepared information unilaterally, lacking the flexibility to respond to specific user questions and requests. Furthermore, the lack of integration of visual and auditory content limits the user experience. This can lead to users not fully understanding the product's features and benefits, potentially reducing their purchase intent.

[0096] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0097] In this invention, the server includes means for analyzing voice data and understanding the user's intent, means for generating relevant content such as visuals, audio, and video using a generative AI model, and means for transmitting the generated content to a terminal. This makes it possible to generate and provide appropriate and interactive content in response to the user's specific questions and requests.

[0098] A "user" refers to a person who operates a system and obtains information.

[0099] "Means of inputting voice" refers to devices such as microphones that capture the user's voice and input it into the system.

[0100] "Means of converting to text data" refers to the process or equipment used to convert captured audio into text information.

[0101] "Methods for analyzing and understanding user intent" refers to the process of analyzing text data converted from speech to understand the information and questions the user is seeking.

[0102] A "generative AI model" refers to an artificial intelligence algorithm that learns from large amounts of data and generates new content.

[0103] "Means for generating visual, audio, and video content" refers to processes and devices that generate visual and auditory content for users based on analysis results.

[0104] "Means for sending generated content to a terminal" refers to communication means for transmitting content generated on a server to a terminal.

[0105] "Means of visual and auditory presentation" refers to devices and methods that provide generated content to users visually and aurally, using displays, speakers, etc.

[0106] A "prompt statement" refers to a sentence used as a specific input to a generative AI model.

[0107] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those commands.

[0108] Hardware configuration

[0109] This system consists of the following main components:

[0110] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[0111] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[0112] Software Configuration

[0113] The following software and algorithms operate between the terminal and the server:

[0114] Speech recognition software: Converts voice input from the user into text data.

[0115] Specific software used: For example, Google® Cloud Speech-to-Text API.

[0116] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[0117] Specific software used: For example, the Python®-based NLTK library and SpaCy.

[0118] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[0119] Specific models used: For example, OpenAI® GPT-3®.

[0120] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[0121] Specific protocols to use: HTTP, MQTT, etc.

[0122] System operation

[0123] User voice input

[0124] The user asks a question about product information, and the device's microphone captures the user's voice. Speech recognition software converts this into text data, and the device sends the text data to the server.

[0125] Data analysis and content generation

[0126] The server receives text data, and the NLP engine analyzes it. Based on the analysis, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.). Next, the server uses a generative AI model to generate relevant visual, audio, and video content.

[0127] Examples of generated content: photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[0128] Content display and response

[0129] The device receives the generated content, displays the visual content on the screen, and plays the audio content using the speaker. If the user requests further information, they can input voice again, and the process repeats.

[0130] Specific example

[0131] When a user asks, "What are the features of this smartphone?", the device's microphone captures the audio, and speech recognition software converts it into text data. Next, the text data is sent to a server and analyzed by an NLP engine. Based on the analysis, a generative AI model generates content with the prompt, "Please describe the smartphone's camera features." Then, a photo taken with the high-resolution camera is displayed on the device's screen, and explanatory audio is played through the speaker.

[0132] In this way, users can visually and audibly experience the features and benefits of a product through the robot, thereby increasing their desire to purchase. This system solves the problems of conventional, limited demonstration methods and enables the effective delivery of product information.

[0133] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0134] Step 1:

[0135] A user asks a question about product information. The user's voice is input to the microphone. Example of a user question: "What are the features of this smartphone?" The input is voice data. Specifically, the microphone captures the voice signal and temporarily stores it in internal memory.

[0136] Step 2:

[0137] The device's speech recognition software captures audio and converts it into text data. The audio data is then sent to the Google Cloud Speech-to-Text API to retrieve the text data. The input is audio data, and the output is text data. Specifically, the device sends audio data to an external speech recognition API and receives text data as a response from the API.

[0138] Step 3:

[0139] The device sends the generated text data to the server. The text data is sent to the server's API endpoint using the HTTP protocol. The input is text data, and the output is an HTTP request. Specifically, the device sends the text data to the server as the payload of the HTTP request.

[0140] Step 4:

[0141] The server receives text data and performs analysis using an NLP engine. It uses Python's NLTK and SpaCy libraries to analyze the text data and extract the user's intent. The input is text data, and the output is the analysis result (user intent). Specifically, the server inputs text data into the NLP library and retrieves the analysis result.

[0142] Step 5:

[0143] The server generates relevant visual, audio, and video content using a generative AI model based on the analysis results. A prompt is input to the generative AI model (e.g., OpenAI GPT-3) to generate the necessary content. The input is the analysis results and the prompt, and the output is the generated content. Specifically, a prompt is created and input to the generative AI model to obtain content. Example: Prompt: "Please explain the camera functions of a smartphone."

[0144] Step 6:

[0145] The server sends the generated content to the terminal. The generated visual, audio, and video files are sent to the terminal as an HTTP response. The input is the generated content, and the output is the HTTP response. Specifically, the server sends the generated content to the terminal as an HTTP response.

[0146] Step 7:

[0147] The device receives generated content, displays visual content on the screen, and plays audio content through the speaker. The input is the HTTP response (generated content), and the output is the display of visual and auditory content. Specifically, the display module displays images and videos, and the audio module plays audio. Example: The display shows a photo taken with a high-resolution camera and plays the audio, "This smartphone has a high-resolution camera and can take very clear photos."

[0148] Step 8:

[0149] If the user requests further information, they input voice again, and the process is repeated. The input is new voice data, and the output is updated text data. Specifically, the device captures the user's new question again, and the aforementioned processing steps are executed again. Example: "How long does the battery last?"

[0150] (Application Example 1)

[0151] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0152] In traditional brick-and-mortar stores, customers often relied on sales staff to obtain detailed product information. This method suffers from inconsistencies in information delivery due to variations in staff knowledge and explanation skills. Furthermore, during busy periods, adequate service may be difficult to provide. Additionally, insufficient visual and auditory information makes it difficult for customers to fully understand product features and benefits.

[0153] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0154] In this invention, the server includes means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for presenting the generated content to the user visually and aurally, means for providing relevant content based on questions from customers via a robot installed in a physical store, and means for processing the analyzed and generated content through a cloud or local server. This enables users to receive consistent, high-quality information and to better understand the features and benefits of products.

[0155] "Means of inputting voice" refers to devices or systems that effectively capture voice from users.

[0156] "Methods for converting audio to text data" refer to software or algorithms that extract text information from captured audio and convert it into text data.

[0157] "Means for analyzing text data and generating related visual, auditory, and video content" refers to technologies that understand user intent and questions based on transcribed data and create corresponding visual, auditory, and video content.

[0158] "Means for displaying and playing generated content" refers to devices and mechanisms for showing or letting users hear generated visual, auditory, and video content.

[0159] "Means for presenting the generated content to the user visually and audibly" refers to a system that transmits the generated content to the user through a screen, speakers, or the like.

[0160] "A means of providing relevant content based on customer questions via robots installed in physical stores" refers to a system in which robots placed in actual stores generate and provide various content in response to user questions.

[0161] "Means for processing the analyzed and generated content via a cloud or local server" refers to a technology for efficiently processing the generated and analyzed content on a cloud server or a server on an internal network.

[0162] This invention provides a system that allows customers to receive product information visually and aurally using a robot installed in a physical store. The system comprises means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating relevant visual, auditory, and video content, means for displaying and playing the generated content, means for presenting it visually and aurally, means for providing relevant content based on customer questions via a robot installed in the physical store, and means for processing the generated content through a cloud or local server.

[0163] Hardware configuration

[0164] Terminal (robot body):

[0165] Microphone: A high-sensitivity microphone is used to capture the voices of customers.

[0166] Display: Use a high-resolution display to show the generated visual content.

[0167] Speakers: High-quality speakers are used to play the generated audio content.

[0168] server:

[0169] Cloud servers or local servers: Used for data analysis and content generation.

[0170] Software Configuration

[0171] Software and algorithms that operate between terminals and servers:

[0172] Speech recognition software: Uses the Google Speech-to-Text API to convert user speech into text data.

[0173] Natural Language Processing (NLP) Engine: Uses SpaCy or Google Natural Language API to analyze text data and understand user intent.

[0174] Generative AI Model: Uses OpenAI GPT-4®, DALL-E, and a TTS engine to generate relevant visual, auditory, and video content.

[0175] Specific examples of actions

[0176] 1. A customer asks, "Could you tell me about the features of this TV?"

[0177] The voices of customers are captured using a microphone.

[0178] The captured audio is converted into text data using the Google Speech-to-Text API.

[0179] Using an NLP engine, text data is analyzed to understand the customer's intentions.

[0180] Relevant visual, auditory, and video content is generated using generative AI models (GPT-4, DALL-E, TTS engine).

[0181] 2. The generated content is displayed on the robot's screen and played back as audio through its speaker.

[0182] Example: A high-resolution television image is displayed on the screen, and a voice message plays from the speaker saying, "This television is 4K compatible, so you can enjoy high-resolution images."

[0183] A short demo video is played to visually demonstrate the video quality.

[0184] Example of a prompt:

[0185] Input: "Please tell me about the features of this television."

[0186] Prompt message:

[0187] Please provide the following information, including the features of your television:

[0188] 1. Resolution

[0189] 2. Special features (4K, HDR, smart features, etc.)

[0190] 3. Size

[0191] 4. Price range

[0192] 5. Buyer Reviews

[0193] Proposed output:

[0194] This TV is 4K compatible, allowing you to enjoy high-resolution images. It also features HDR functionality, providing vibrant colors. It's a 55-inch model and priced at approximately 100,000 yen. It has received high praise from many buyers.

[0195] In this way, customers can receive detailed product information through the robot. The entire system is managed via the cloud or a local server, enabling efficient and effective information delivery.

[0196] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0197] Step 1:

[0198] Process name: Voice input capture

[0199] Subject: User

[0200] Specific actions:

[0201] The user asks a question to the robot. An example message might be, "Please tell me about the features of this television." A high-sensitivity microphone built into the robot captures this audio.

[0202] Input: User's voice

[0203] Output: Audio data

[0204] Step 2:

[0205] Process name: Text conversion of audio data

[0206] Subject: terminal

[0207] Specific actions:

[0208] The captured audio data is converted into text data in real time by the device's speech recognition software (Google Speech-to-Text API).

[0209] Input: Audio data

[0210] Output: Text data

[0211] Step 3:

[0212] Process name: Sending text data

[0213] Subject: terminal

[0214] Specific actions:

[0215] Text data is sent from the terminal to the server. A secure communication protocol is used for this transmission.

[0216] Input: Text data

[0217] Output: Text data sent to the server

[0218] Step 4:

[0219] Process name: Text data analysis

[0220] Subject: Server

[0221] Specific actions:

[0222] The server analyzes the received text data using a natural language processing (NLP) engine (SpaCy or Google Natural Language API) to understand the user's intent and requests. Specifically, if the question is related to a product, it identifies information such as the product's features, benefits, price, and ratings.

[0223] Input: Text data sent to the server

[0224] Output: Analyzed text data (information such as product features and benefits)

[0225] Step 5:

[0226] Process name: Content generation

[0227] Subject: Server

[0228] Specific actions:

[0229] The server uses a generation AI model (GPT-4, DALL-E, TTS engine) to generate visual, auditory, and video content based on the analyzed text data. Examples include high-resolution images of the product, graphs and explanatory audio demonstrating product features, and demo videos.

[0230] Input: Analyzed text data

[0231] Output: Generated visual, auditory, and video content

[0232] Step 6:

[0233] Process name: Send content

[0234] Subject: Server

[0235] Specific actions:

[0236] The generated visual, auditory, and video content is transmitted from the server to the terminal. Protocols are used to maintain data integrity and security during this process.

[0237] Input: Generated visual, auditory, and video content

[0238] Output: Content sent to the device

[0239] Step 7:

[0240] Process name: Display and play content

[0241] Subject: Terminal (robot body)

[0242] Specific actions:

[0243] The device displays and plays received content. Images and graphs are shown on a high-resolution display, and explanatory audio is played using high-quality speakers. Demo videos are also played on the display.

[0244] For example, a voice explanation such as, "This TV is 4K compatible, allowing you to enjoy high-resolution images," is played while product images and demo videos are displayed.

[0245] Input: Content sent to the device

[0246] Output: Visual and auditory content displayed and played for the user.

[0247] This series of steps allows customers to interactively obtain product information.

[0248] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0249] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands through a robot, which then generates and presents various forms of data (images, audio, video, etc.) based on those commands. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it enables interactive responses that correspond to the user's emotional state.

[0250] System Configuration

[0251] Hardware configuration

[0252] This system consists of the following main components:

[0253] Terminal (robot body): Equipped with voice input means, display, speaker, and camera, it serves as the main component for interactive demonstrations.

[0254] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[0255] Software Configuration

[0256] The following software and algorithms operate between the terminal and the server:

[0257] Speech recognition software: Converts voice input from the user into text data.

[0258] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[0259] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[0260] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state.

[0261] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[0262] System operation

[0263] User voice input and emotion recognition

[0264] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[0265] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[0266] 3. The device uses its camera to capture the user's facial expressions, and the emotion engine analyzes this to recognize the emotional state (e.g., excitement, interest, questioning, etc.).

[0267] 4. The device sends text data and sentiment analysis results to the server.

[0268] Data analysis and content generation

[0269] 5. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[0270] 6. The server takes the sentiment analysis results into account and uses a generative AI model to generate relevant visual, audio, and video content.

[0271] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[0272] Content display and response

[0273] 7. The device receives the generated content.

[0274] 8. Display visual content on the device's screen and play audio content using the speaker.

[0275] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[0276] 9. The user visually and audibly confirms the content that has been displayed and played.

[0277] 10. The emotion engine continuously monitors user responses and updates the emotion state to the server as needed.

[0278] Providing additional information and continuing the interaction

[0279] 11. If the user requests more detailed information, additional questions are asked (e.g., "How long is the battery life?").

[0280] 12. The terminal captures the user's voice again, converts it into text data, and sends it to the server.

[0281] 13. The server analyzes the new text data and the latest sentiment analysis results, generates the necessary new content, and sends it to the terminal.

[0282] End of Interaction

[0283] 14. If the user instructs the robot to end the interaction, the terminal captures the voice instruction to end, and the system switches to the standby mode.

[0284] In this way, based on the user's voice input and emotional state, this system provides product details in a step-by-step and interactive manner. As a result, it is expected to arouse further interest from users and enhance their purchasing desire.

[0285] The following describes the processing flow.

[0286] Step 1:

[0287] The user addresses the robot and says, "Please tell me the features of this smartphone."

[0288] Step 2:

[0289] The microphone of the terminal captures the user's voice.

[0290] Step 3:

[0291] The terminal uses voice recognition software to convert the captured voice into text data.

[0292] Step 4:

[0293] The device's camera captures the user's facial expressions.

[0294] Step 5:

[0295] The device uses an emotion engine to analyze captured facial data and recognize emotional states (e.g., excitement, interest, questioning).

[0296] Step 6:

[0297] The terminal sends the converted text data and sentiment analysis results to the server.

[0298] Step 7:

[0299] The server receives text data and performs analysis using an NLP engine. Based on the analysis results, it identifies the information the user is requesting (e.g., smartphone camera functions, battery life, etc.).

[0300] Step 8:

[0301] The server uses a generative AI model, taking sentiment analysis results into account, to generate relevant visual, audio, and video content.

[0302] Step 9:

[0303] The server sends the generated content to the terminal.

[0304] Step 10:

[0305] The device receives the generated content.

[0306] Step 11:

[0307] Visual content is displayed on the terminal's display, and audio content is played through the speaker. As an example, a photo taken by a high-resolution camera is displayed on the display, and audio such as "This smartphone has a high-resolution camera and can take very clear photos" is played.

[0308] Step 12:

[0309] The user visually and auditorily confirms the displayed content and the played audio.

[0310] Step 13:

[0311] The emotion engine continuously monitors the user's reaction and determines the emotional state.

[0312] Step 14:

[0313] If the user requests more detailed information (for example, asks "How long is the battery life?"), additional voice instructions are given.

[0314] Step 15:

[0315] The terminal captures the user's voice again, converts it into text data using voice recognition software, and sends it to the server.

[0316] Step 16:

[0317] The server receives the new text data and the latest emotion analysis results, and analyzes them again using the NLP engine. Based on the analysis results, new necessary content is generated and sent to the terminal.

[0318] Step 17:

[0319] The device receives new content, displays it on the screen, and plays it through the speaker. For example, a graph showing battery life might be displayed, and a voice message might say, "This smartphone's battery lasts 24 hours with normal use."

[0320] Step 18:

[0321] If the user is satisfied, they instruct the robot to end the interaction.

[0322] Step 19:

[0323] The terminal captures the voice command to terminate, and the system enters standby mode.

[0324] (Example 2)

[0325] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0326] Conventional systems, when providing information based on user voice input, lacked the ability to respond in a way that took into account the user's emotional state, resulting in a uniform user experience and low satisfaction. Furthermore, they lacked the ability to integrate and deliver diverse content formats, limiting the effectiveness of information transmission.

[0327] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0328] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for capturing the user's facial expressions and analyzing their emotional state, means for transmitting the text data and emotional state, means for analyzing the text data and generating related content such as visuals, audio, and video, means for displaying and playing the generated content, and means for presenting the generated content to the user visually and aurally. This enables the provision of interactive content that reflects the user's emotional state, thereby improving the quality of the user experience.

[0329] A "user" is an entity that gives instructions or asks questions to a system using voice.

[0330] A "terminal" is a device that performs voice input, voice recognition, facial expression capture, result display, and playback.

[0331] A "server" is a device that receives text data and sentiment analysis results, and performs data analysis and content generation.

[0332] "Voice input means" refers to devices such as microphones that capture the user's voice.

[0333] "Speech recognition software" is a program that converts captured audio into text data.

[0334] "Text data" refers to character information converted by speech recognition software.

[0335] "Facial expression capture means" refers to devices such as cameras that capture the user's facial expressions.

[0336] An "emotion engine" is software that analyzes a user's emotional state from captured facial expressions.

[0337] A "communication protocol" is a set of communication rules for efficiently sending and receiving data between a terminal and a server.

[0338] A "natural language processing (NLP) engine" is a program that analyzes text data to understand the user's intent.

[0339] A "generative AI model" is an algorithm that generates relevant visual, audio, and video content based on data analysis results.

[0340] "Visual content" refers to visual information such as images and graphics.

[0341] "Audio content" refers to auditory information such as spoken language.

[0342] "Video content" refers to information that combines dynamic visual and audio information.

[0343] "Content presentation means" refers to a function that provides generated visual, audio, and video content to the user visually and aurally.

[0344] The system of the present invention enhances the user experience by dynamically generating and presenting relevant visual, audio, and video content based on the user's voice input and emotion analysis. The embodiments for carrying out the present invention will be described in detail below.

[0345] Hardware configuration

[0346] This system consists of the following main hardware components:

[0347] Terminal: Equipped with voice input, display, speaker, and camera, it serves as the main component for interactive demonstrations. Specifically, it includes a high-sensitivity microphone, high-resolution display, speaker, and high-resolution camera.

[0348] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation. Servers with high-performance computing resources are required.

[0349] Software Configuration

[0350] The following software and algorithms operate between the terminal and the server:

[0351] Speech recognition software: Converts voice input from a user into text data. For example, technologies such as Google Cloud Speech-to-Text are used for speech recognition.

[0352] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent. OpenAI GPT-3 is an example of this.

[0353] Generative AI models: These models generate relevant content such as visuals, audio, and video based on analysis results. For example, OpenAI DALL-E and GPT-3 are used.

[0354] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state. For example, the Microsoft® Azure® Emotion API is used.

[0355] Communication protocol: A protocol that efficiently sends and receives data between a terminal and a server. For example, HTTPS is used.

[0356] System operation

[0357] User voice input and emotion recognition

[0358] The user asks questions about product information using voice.

[0359] The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[0360] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state.

[0361] The device sends text data and sentiment analysis results to the server.

[0362] Data analysis and content generation

[0363] The server analyzes the received text data using a natural language processing engine to identify the user's intent.

[0364] The server uses a generation AI model that takes sentiment analysis results into account to generate content in various formats.

[0365] For example, if a user asks, "What are the features of this smartwatch?", the system will generate a demo video of the heart rate monitoring function and a graph showing battery life.

[0366] Content display and response

[0367] The device receives the generated content, displays the visual content on its screen, and plays the audio content using its speaker.

[0368] For example, a demo video of the heart rate measurement function is displayed on the screen, and a voice guide plays saying, "This smartwatch is capable of accurate heart rate measurement."

[0369] Check the content that users have viewed and played.

[0370] The emotion engine monitors the user's reactions and updates the emotion state to the server as needed.

[0371] Examples of specific cases and prompt statements

[0372] As a concrete example, consider a scenario where a user asks, "Tell me about the features of this smartwatch." In this case, the user's voice is captured by the device's microphone and converted into text data by speech recognition software. The server then analyzes this text data with a natural language processing engine, and a generative AI model generates a demo video of the heart rate measurement function and a graph of battery life. If the emotion engine analyzes the user's facial expressions and recognizes an excited emotional state, it continues with a more detailed explanation of the functions.

[0373] An example of a prompt might be: "Consider a scenario where a user asks, 'What are the features of this smartwatch?' and come up with prompts that would allow the generative AI model to create appropriate visual, audio, and video content."

[0374] As described above, the system of the present invention can provide detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state, thereby significantly improving the quality of the user experience.

[0375] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0376] Step 1:

[0377] The user asks a question about product information using voice. For example, they might say, "Tell me about the features of this smartwatch."

[0378] Input: User voice input

[0379] Output: Captured audio data

[0380] Step 2:

[0381] The device's microphone captures the user's voice, and speech recognition software converts this into text data. For example, Google Cloud Speech-to-Text can be used.

[0382] Input: Captured audio data

[0383] Output: Converted text data (e.g., "Tell me the features of this smartwatch")

[0384] Step 3:

[0385] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state. For example, the Microsoft Azure Emotion API can be used.

[0386] Input: Captured facial expression data

[0387] Output: Analyzed emotional state data (e.g., excitement, interest)

[0388] Step 4:

[0389] The device sends text data and sentiment analysis results to the server. For example, HTTPS is used to efficiently transmit the data.

[0390] Input: Text data and sentiment state data

[0391] Output: Sending data to the server

[0392] Step 5:

[0393] The server analyzes the received text data using a natural language processing (NLP) engine. For example, OpenAI GPT-3 can be used.

[0394] Input: Text data

[0395] Output: Analysis results of user intent (e.g., smartwatch heart rate monitoring function, battery life)

[0396] Step 6:

[0397] The server considers the sentiment analysis results and uses a generative AI model to generate relevant visual, audio, and video content. For example, it might use OpenAI DALL-E or GPT-3.

[0398] Input: Analysis results of user intent and emotional state data

[0399] Output: Generated visual, audio, and video content (e.g., a demo video of the heart rate measurement function, a graph showing battery life)

[0400] Step 7:

[0401] The device receives the generated content from the server.

[0402] Input: Content data sent from the server

[0403] Output: Content data stored on the device

[0404] Step 8:

[0405] The device displays visual content on its screen and plays audio content through its speaker. For example, it might show a demo video of the heart rate measurement function on the screen and play an audio guide through the speaker saying, "This smartwatch is capable of accurate heart rate measurement."

[0406] Input: Content data stored on the device

[0407] Output: Display of visual content and playback of audio content

[0408] Step 9:

[0409] Review the content displayed and played by the user. For example, review the display image and speaker description.

[0410] Input: Visual and audio content

[0411] Output: User understanding and response

[0412] Step 10:

[0413] The emotion engine continuously monitors the user's reactions and updates the emotional state to the server as needed. For example, if the user's facial expression changes to one of surprise, the emotion engine analyzes this and sends the information to the server.

[0414] Input: User's facial expression data

[0415] Output: Updated sentiment state data

[0416] Step 11:

[0417] The user asks additional questions by voice. For example, they might say, "How long does the battery last?"

[0418] Input: Additional voice questions

[0419] Output: Captured audio data

[0420] Step 12:

[0421] The device then captures the user's voice again and converts it into text data using speech recognition software.

[0422] Input: Captured audio data

[0423] Output: Converted text data

[0424] Step 13:

[0425] The server analyzes the new text data and the latest sentiment analysis results to generate the necessary new content.

[0426] Input: New text data and latest sentiment state data

[0427] Output: Newly generated visual, audio, and video content

[0428] Step 14:

[0429] The device receives newly generated content and delivers it to the user using its display and speaker. For example, it might display a graph showing battery life on the display and provide an explanation with voice guidance.

[0430] Input: Newly generated content data

[0431] Output: Display of visual content and playback of audio content

[0432] Step 15:

[0433] The user gives a voice command to end the interaction. For example, they might say, "I'm done."

[0434] Input: Voice command to end

[0435] Output: Captured audio data

[0436] Step 16:

[0437] The terminal captures the voice command to terminate the process, converts it into text data using speech recognition software, and then puts the system into standby mode.

[0438] Input: Voice command to end

[0439] Output: System transition to standby mode

[0440] (Application Example 2)

[0441] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0442] Conventional product information systems in physical stores have problems in providing quick and detailed responses to customer questions, and furthermore, in providing interactive responses that match the customer's emotions. In addition, there is a lack of means for customers to obtain specific product information visually and aurally, making it difficult to increase their purchasing intent. Therefore, the present invention aims to solve these problems and realize more effective and customized information provision to customers.

[0443] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0444] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for analyzing the user's facial expressions and recognizing their emotional state, and means for customizing the generated content according to the user's emotional state and presenting it visually and aurally. As a result, customers can obtain specific, visual, and auditory product information in real time within a physical store, and since that information is customized according to the customer's emotional state, an interactive experience that enhances their desire to purchase becomes possible.

[0445] "Means of inputting voice" refers to devices or software that receive voice signals from a user and process them as digital signals.

[0446] "Methods for converting speech to text data" refer to algorithms or software that analyze an input speech signal and convert it into corresponding text data.

[0447] "Means for analyzing text data and generating related visual, auditory, and video content" refers to a system that uses natural language processing and generative AI models based on text data to generate content in various formats.

[0448] "Means for displaying and playing generated content" refers to a system that outputs generated visual and auditory content to a display device or speaker so that the user can see and hear it.

[0449] "Methods for analyzing a user's facial expressions to recognize their emotional state" refer to software or algorithms that capture a user's facial expressions through cameras or sensors, analyze them, and estimate the user's emotional state.

[0450] "Means for customizing generated content according to the user's emotional state and presenting it visually and aurally" refers to a system that adjusts the format and content of the content considering the user's emotional state and provides information to the user in an appropriate manner.

[0451] One embodiment of the present invention relates to a system for effectively providing product information to customers in a physical store. This system uses voice input, speech recognition, natural language processing, sentiment recognition, and generative AI models to provide customized information in response to customer inquiries.

[0452] Configuration of the main components

[0453] Hardware configuration

[0454] Terminal: An interactive device equipped with voice input, a display, speakers, and a camera. It captures the user's voice and facial expressions and displays and plays the generated content.

[0455] Server: A central system located in the cloud or locally, which performs data analysis and content generation.

[0456] Software Configuration

[0457] The following software and algorithms operate between the terminal and the server:

[0458] Speech recognition software: Converts user speech into text data. Example: SpeechRecognition library.

[0459] Natural Language Processing Engine: Analyzes text data and generates the best possible answer to a question. Example: Transformers in Hugging Face.

[0460] Generative AI models: Generate relevant visual and auditory content. Example: A generative model for Hugging Face.

[0461] Emotion recognition engine: Analyzes the user's facial expressions to recognize their emotional state. Example: EmotionRecognizer.

[0462] Communication protocol: Enables efficient transmission and reception of data between terminals and servers.

[0463] Content generation process

[0464] 1. Voice Input: Users can ask questions about the product using voice. Example: "Please tell me about the features of this smartphone."

[0465] 2. Speech Recognition: The device's microphone captures speech and converts it to text using SpeechRecognition software.

[0466] 3. Text Analysis: The converted text data is sent to the server and analyzed by a natural language processing engine.

[0467] 4. Emotion Recognition: The device's camera captures the user's facial expressions, and the EmotionRecognizer recognizes their emotional state.

[0468] 5. Content Generation: Based on the analysis results and emotional state, the AI ​​generation model generates relevant content such as visuals, audio, and video.

[0469] 6. Content display and playback: The generated content is sent to the device, displayed on the screen, and the audio is played from the speaker.

[0470] Specific example

[0471] For example, if a user asks, "What are the features of this smartphone?", the system will operate as follows:

[0472] Speech recognition software converts the user's voice into text data.

[0473] A natural language processing engine analyzes text data to identify information about the smartphone's features (e.g., camera functions, battery life).

[0474] The generative AI model generates visual content such as high-resolution camera images and graphs showing battery life.

[0475] The emotion recognition engine analyzes the user's facial expressions, and if it determines that the user is in an excited state, it provides additional information such as a voice message saying, "You can take very clear photos."

[0476] Example of a prompt

[0477] "Please tell me about the features of this smartphone."

[0478] "How long does the battery last?"

[0479] This allows customers to receive specific and detailed information in real time, and since that information is customized according to the customer's emotional state, it enables an interactive experience that increases their desire to buy.

[0480] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0481] Step 1:

[0482] Users ask questions about the product using voice.

[0483] Input: User's voice.

[0484] Operation: The device's microphone captures the user's voice.

[0485] Output: Captured audio data.

[0486] Step 2:

[0487] The speech recognition software converts the captured audio data into text data.

[0488] Input: Captured audio data.

[0489] Operation: The device's speech recognition software (e.g., SpeechRecognition library) analyzes the speech data and converts it into corresponding text.

[0490] Output: Converted text data.

[0491] Step 3:

[0492] The terminal sends text data to the server.

[0493] Input: Converted text data.

[0494] Operation: Text data is sent to the server via a communication protocol.

[0495] Output: Text data received by the server.

[0496] Step 4:

[0497] The server's natural language processing engine analyzes the text data and generates the best possible answer to the question.

[0498] Input: Received text data.

[0499] Operation: The server's natural language processing engine (e.g., Hugging Face's Transformers) analyzes the text data, understands the user's intent in the question, and generates an appropriate answer.

[0500] Output: Generated answer text.

[0501] Step 5:

[0502] The device's camera captures the user's facial expressions, and the emotion recognition engine recognizes the user's emotional state.

[0503] Input: User's facial expression image.

[0504] Operation: The device's camera captures the user's facial expressions, and an emotion recognition engine (e.g., EmotionRecognizer) analyzes the facial data to estimate the emotional state.

[0505] Output: Estimated emotional state data.

[0506] Step 6:

[0507] The server uses a generation AI model to generate visual and auditory content based on response text and sentiment state data.

[0508] Input: Generated response text, estimated sentiment state data.

[0509] Operation: The server's generation AI model considers the response text and emotional state to generate relevant visual content (e.g., photos, graphs) and auditory content (e.g., voice guidance).

[0510] Output: Generated visual and auditory content.

[0511] Step 7:

[0512] The generated content is sent to the device, displayed on the screen, and played through the speaker.

[0513] Input: Generated visual and auditory content.

[0514] Operation: Content is sent from the server to the terminal, visual content is displayed on the screen, and audio content is played from the speaker.

[0515] Output: Users visually and aurally perceive the content.

[0516] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0517] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0518] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0519] [Second Embodiment]

[0520] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0521] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0522] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0523] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0524] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0525] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0526] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0527] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0528] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0529] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0530] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0531] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0532] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice instructions via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those instructions.

[0533] System Configuration

[0534] Hardware configuration

[0535] This system consists of the following main components:

[0536] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[0537] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[0538] Software Configuration

[0539] The following software and algorithms operate between the terminal and the server:

[0540] Speech recognition software: Converts voice input from the user into text data.

[0541] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[0542] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[0543] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[0544] System operation

[0545] User voice input

[0546] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[0547] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[0548] 3. The terminal sends text data to the server.

[0549] Data analysis and content generation

[0550] 4. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[0551] 5. The server uses the generated AI model to produce relevant visual, audio, and video content.

[0552] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[0553] Content display and response

[0554] 6. The device receives the generated content.

[0555] 7. Display visual content on the device's screen and play audio content using the speaker.

[0556] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[0557] 8. If the user requests further information (e.g., "How long does the battery last?"), they input voice again, and the process is repeated.

[0558] This allows users to visually and audibly experience the features and benefits of a product through a robot, thereby increasing their desire to purchase. The present invention is a system that, through this interactive process, solves the problems of conventional, limited demonstration methods and enables the effective provision of product information.

[0559] The following describes the processing flow.

[0560] Step 1:

[0561] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[0562] Step 2:

[0563] The terminal (robot) uses its built-in microphone to capture the user's voice.

[0564] Step 3:

[0565] The device uses speech recognition software to convert the captured audio into text data.

[0566] Step 4:

[0567] The terminal sends the converted text data to the server.

[0568] Step 5:

[0569] The server runs a natural language processing (NLP) engine to analyze the text data it receives.

[0570] Step 6:

[0571] The server uses an NLP engine to understand the user's request and identify the necessary information (smartphone features).

[0572] Step 7:

[0573] The server uses a generated AI model to create content such as visuals, audio, and video based on identified information.

[0574] Step 8:

[0575] The server sends the generated content to the terminal.

[0576] Step 9:

[0577] The device displays the received content on its screen and plays the audio through its speaker.

[0578] Step 10:

[0579] The user visually and audibly confirms the content that has been displayed and played.

[0580] Step 11:

[0581] If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[0582] Step 12:

[0583] The device captures the user's voice again, converts it into text data, and sends it to the server.

[0584] Step 13:

[0585] The server analyzes the new text data, generates the necessary new content, and sends it to the terminal.

[0586] Step 14:

[0587] The device displays new content on its screen and plays audio through its speaker.

[0588] Step 15:

[0589] When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[0590] (Example 1)

[0591] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0592] Current product demonstration systems lack the interactivity necessary for users to effectively understand product information. Traditional systems only provide pre-prepared information unilaterally, lacking the flexibility to respond to specific user questions and requests. Furthermore, the lack of integration of visual and auditory content limits the user experience. This can lead to users not fully understanding the product's features and benefits, potentially reducing their purchase intent.

[0593] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0594] In this invention, the server includes means for analyzing voice data and understanding the user's intent, means for generating relevant content such as visuals, audio, and video using a generative AI model, and means for transmitting the generated content to a terminal. This makes it possible to generate and provide appropriate and interactive content in response to the user's specific questions and requests.

[0595] A "user" refers to a person who operates a system and obtains information.

[0596] "Means of inputting voice" refers to devices such as microphones that capture the user's voice and input it into the system.

[0597] "Means of converting to text data" refers to the process or equipment used to convert captured audio into text information.

[0598] "Methods for analyzing and understanding user intent" refers to the process of analyzing text data converted from speech to understand the information and questions the user is seeking.

[0599] A "generative AI model" refers to an artificial intelligence algorithm that learns from large amounts of data and generates new content.

[0600] "Means for generating visual, audio, and video content" refers to processes and devices that generate visual and auditory content for users based on analysis results.

[0601] "Means for sending generated content to a terminal" refers to communication means for transmitting content generated on a server to a terminal.

[0602] "Means of visual and auditory presentation" refers to devices and methods that provide generated content to users visually and aurally, using displays, speakers, etc.

[0603] A "prompt statement" refers to a sentence used as a specific input to a generative AI model.

[0604] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those commands.

[0605] Hardware configuration

[0606] This system consists of the following main components:

[0607] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[0608] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[0609] Software Configuration

[0610] The following software and algorithms operate between the terminal and the server:

[0611] Speech recognition software: Converts voice input from the user into text data.

[0612] Specific software to use: For example, Google Cloud Speech-to-Text API.

[0613] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[0614] Specific software to be used: for example, the Python-based NLTK library or SpaCy.

[0615] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[0616] Specific models to use: For example, OpenAI GPT-3.

[0617] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[0618] Specific protocols to use: HTTP, MQTT, etc.

[0619] System operation

[0620] User voice input

[0621] The user asks a question about product information, and the device's microphone captures the user's voice. Speech recognition software converts this into text data, and the device sends the text data to the server.

[0622] Data analysis and content generation

[0623] The server receives text data, and the NLP engine analyzes it. Based on the analysis, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.). Next, the server uses a generative AI model to generate relevant visual, audio, and video content.

[0624] Examples of generated content: photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[0625] Content display and response

[0626] The device receives the generated content, displays the visual content on the screen, and plays the audio content using the speaker. If the user requests further information, they can input voice again, and the process repeats.

[0627] Specific example

[0628] When a user asks, "What are the features of this smartphone?", the device's microphone captures the audio, and speech recognition software converts it into text data. Next, the text data is sent to a server and analyzed by an NLP engine. Based on the analysis, a generative AI model generates content with the prompt, "Please describe the smartphone's camera features." Then, a photo taken with the high-resolution camera is displayed on the device's screen, and explanatory audio is played through the speaker.

[0629] In this way, users can visually and audibly experience the features and benefits of a product through the robot, thereby increasing their desire to purchase. This system solves the problems of conventional, limited demonstration methods and enables the effective delivery of product information.

[0630] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0631] Step 1:

[0632] A user asks a question about product information. The user's voice is input to the microphone. Example of a user question: "What are the features of this smartphone?" The input is voice data. Specifically, the microphone captures the voice signal and temporarily stores it in internal memory.

[0633] Step 2:

[0634] The device's speech recognition software captures audio and converts it into text data. The audio data is then sent to the Google Cloud Speech-to-Text API to retrieve the text data. The input is audio data, and the output is text data. Specifically, the device sends audio data to an external speech recognition API and receives text data as a response from the API.

[0635] Step 3:

[0636] The device sends the generated text data to the server. The text data is sent to the server's API endpoint using the HTTP protocol. The input is text data, and the output is an HTTP request. Specifically, the device sends the text data to the server as the payload of the HTTP request.

[0637] Step 4:

[0638] The server receives text data and performs analysis using an NLP engine. It uses Python's NLTK and SpaCy libraries to analyze the text data and extract the user's intent. The input is text data, and the output is the analysis result (user intent). Specifically, the server inputs text data into the NLP library and retrieves the analysis result.

[0639] Step 5:

[0640] The server generates relevant visual, audio, and video content using a generative AI model based on the analysis results. A prompt is input to the generative AI model (e.g., OpenAI GPT-3) to generate the necessary content. The input is the analysis results and the prompt, and the output is the generated content. Specifically, a prompt is created and input to the generative AI model to obtain content. Example: Prompt: "Please explain the camera functions of a smartphone."

[0641] Step 6:

[0642] The server sends the generated content to the terminal. The generated visual, audio, and video files are sent to the terminal as an HTTP response. The input is the generated content, and the output is the HTTP response. Specifically, the server sends the generated content to the terminal as an HTTP response.

[0643] Step 7:

[0644] The device receives generated content, displays visual content on the screen, and plays audio content through the speaker. The input is the HTTP response (generated content), and the output is the display of visual and auditory content. Specifically, the display module displays images and videos, and the audio module plays audio. Example: The display shows a photo taken with a high-resolution camera and plays the audio, "This smartphone has a high-resolution camera and can take very clear photos."

[0645] Step 8:

[0646] If the user requests further information, they input voice again, and the process is repeated. The input is new voice data, and the output is updated text data. Specifically, the device captures the user's new question again, and the aforementioned processing steps are executed again. Example: "How long does the battery last?"

[0647] (Application Example 1)

[0648] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0649] In traditional brick-and-mortar stores, customers often relied on sales staff to obtain detailed product information. This method suffers from inconsistencies in information delivery due to variations in staff knowledge and explanation skills. Furthermore, during busy periods, adequate service may be difficult to provide. Additionally, insufficient visual and auditory information makes it difficult for customers to fully understand product features and benefits.

[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0651] In this invention, the server includes means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for presenting the generated content to the user visually and aurally, means for providing relevant content based on questions from customers via a robot installed in a physical store, and means for processing the analyzed and generated content through a cloud or local server. This enables users to receive consistent, high-quality information and to better understand the features and benefits of products.

[0652] "Means of inputting voice" refers to devices or systems that effectively capture voice from users.

[0653] "Methods for converting audio to text data" refer to software or algorithms that extract text information from captured audio and convert it into text data.

[0654] "Means for analyzing text data and generating related visual, auditory, and video content" refers to technologies that understand user intent and questions based on transcribed data and create corresponding visual, auditory, and video content.

[0655] "Means for displaying and playing generated content" refers to devices and mechanisms for showing or letting users hear generated visual, auditory, and video content.

[0656] "Means for presenting the generated content to the user visually and audibly" refers to a system that transmits the generated content to the user through a screen, speakers, or the like.

[0657] "A means of providing relevant content based on customer questions via robots installed in physical stores" refers to a system in which robots placed in actual stores generate and provide various content in response to user questions.

[0658] "Means for processing the analyzed and generated content via a cloud or local server" refers to a technology for efficiently processing the generated and analyzed content on a cloud server or a server on an internal network.

[0659] This invention provides a system that allows customers to receive product information visually and aurally using a robot installed in a physical store. The system comprises means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating relevant visual, auditory, and video content, means for displaying and playing the generated content, means for presenting it visually and aurally, means for providing relevant content based on customer questions via a robot installed in the physical store, and means for processing the generated content through a cloud or local server.

[0660] Hardware configuration

[0661] Terminal (robot body):

[0662] Microphone: A high-sensitivity microphone is used to capture the voices of customers.

[0663] Display: Use a high-resolution display to show the generated visual content.

[0664] Speakers: High-quality speakers are used to play the generated audio content.

[0665] server:

[0666] Cloud servers or local servers: Used for data analysis and content generation.

[0667] Software Configuration

[0668] Software and algorithms that operate between terminals and servers:

[0669] Speech recognition software: Uses the Google Speech-to-Text API to convert user speech into text data.

[0670] Natural Language Processing (NLP) Engine: Uses SpaCy or Google Natural Language API to analyze text data and understand user intent.

[0671] Generative AI Model: Generates relevant visual, auditory, and video content using OpenAI GPT-4, DALL-E, and a TTS engine.

[0672] Specific examples of actions

[0673] 1. A customer asks, "Could you tell me about the features of this TV?"

[0674] The voices of customers are captured using a microphone.

[0675] The captured audio is converted into text data using the Google Speech-to-Text API.

[0676] Using an NLP engine, text data is analyzed to understand the customer's intentions.

[0677] Relevant visual, auditory, and video content is generated using generative AI models (GPT-4, DALL-E, TTS engine).

[0678] 2. The generated content is displayed on the robot's screen and played back as audio through its speaker.

[0679] Example: A high-resolution television image is displayed on the screen, and a voice message plays from the speaker saying, "This television is 4K compatible, so you can enjoy high-resolution images."

[0680] A short demo video is played to visually demonstrate the video quality.

[0681] Example of a prompt:

[0682] Input: "Please tell me about the features of this television."

[0683] Prompt message:

[0684] Please provide the following information, including the features of your television:

[0685] 1. Resolution

[0686] 2. Special features (4K, HDR, smart features, etc.)

[0687] 3. Size

[0688] 4. Price range

[0689] 5. Buyer Reviews

[0690] Proposed output:

[0691] This TV is 4K compatible, allowing you to enjoy high-resolution images. It also features HDR functionality, providing vibrant colors. It's a 55-inch model and priced at approximately 100,000 yen. It has received high praise from many buyers.

[0692] In this way, customers can receive detailed product information through the robot. The entire system is managed via the cloud or a local server, enabling efficient and effective information delivery.

[0693] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0694] Step 1:

[0695] Process name: Voice input capture

[0696] Subject: User

[0697] Specific actions:

[0698] The user asks a question to the robot. An example message might be, "Please tell me about the features of this television." A high-sensitivity microphone built into the robot captures this audio.

[0699] Input: User's voice

[0700] Output: Audio data

[0701] Step 2:

[0702] Process name: Text conversion of audio data

[0703] Subject: terminal

[0704] Specific actions:

[0705] The captured audio data is converted into text data in real time by the device's speech recognition software (Google Speech-to-Text API).

[0706] Input: Audio data

[0707] Output: Text data

[0708] Step 3:

[0709] Process name: Sending text data

[0710] Subject: terminal

[0711] Specific actions:

[0712] Text data is sent from the terminal to the server. A secure communication protocol is used for this transmission.

[0713] Input: Text data

[0714] Output: Text data sent to the server

[0715] Step 4:

[0716] Process name: Text data analysis

[0717] Subject: Server

[0718] Specific actions:

[0719] The server analyzes the received text data using a natural language processing (NLP) engine (SpaCy or Google Natural Language API) to understand the user's intent and requests. Specifically, if the question is related to a product, it identifies information such as the product's features, benefits, price, and ratings.

[0720] Input: Text data sent to the server

[0721] Output: Analyzed text data (information such as product features and benefits)

[0722] Step 5:

[0723] Process name: Content generation

[0724] Subject: Server

[0725] Specific actions:

[0726] The server uses a generation AI model (GPT-4, DALL-E, TTS engine) to generate visual, auditory, and video content based on the analyzed text data. Examples include high-resolution images of the product, graphs and explanatory audio demonstrating product features, and demo videos.

[0727] Input: Analyzed text data

[0728] Output: Generated visual, auditory, and video content

[0729] Step 6:

[0730] Process name: Send content

[0731] Subject: Server

[0732] Specific actions:

[0733] The generated visual, auditory, and video content is transmitted from the server to the terminal. Protocols are used to maintain data integrity and security during this process.

[0734] Input: Generated visual, auditory, and video content

[0735] Output: Content sent to the device

[0736] Step 7:

[0737] Process name: Display and play content

[0738] Subject: Terminal (robot body)

[0739] Specific actions:

[0740] The device displays and plays received content. Images and graphs are shown on a high-resolution display, and explanatory audio is played using high-quality speakers. Demo videos are also played on the display.

[0741] For example, a voice explanation such as, "This TV is 4K compatible, allowing you to enjoy high-resolution images," is played while product images and demo videos are displayed.

[0742] Input: Content sent to the device

[0743] Output: Visual and auditory content displayed and played for the user.

[0744] This series of steps allows customers to interactively obtain product information.

[0745] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0746] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands through a robot, which then generates and presents various forms of data (images, audio, video, etc.) based on those commands. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it enables interactive responses that correspond to the user's emotional state.

[0747] System Configuration

[0748] Hardware configuration

[0749] This system consists of the following main components:

[0750] Terminal (robot body): Equipped with voice input means, display, speaker, and camera, it serves as the main component for interactive demonstrations.

[0751] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[0752] Software Configuration

[0753] The following software and algorithms operate between the terminal and the server:

[0754] Speech recognition software: Converts voice input from the user into text data.

[0755] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[0756] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[0757] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state.

[0758] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[0759] System operation

[0760] User voice input and emotion recognition

[0761] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[0762] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[0763] 3. The device uses its camera to capture the user's facial expressions, and the emotion engine analyzes this to recognize the emotional state (e.g., excitement, interest, questioning, etc.).

[0764] 4. The device sends text data and sentiment analysis results to the server.

[0765] Data analysis and content generation

[0766] 5. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[0767] 6. The server takes the sentiment analysis results into account and uses a generative AI model to generate relevant visual, audio, and video content.

[0768] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[0769] Content display and response

[0770] 7. The device receives the generated content.

[0771] 8. Display visual content on the device's screen and play audio content using the speaker.

[0772] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[0773] 9. The user visually and audibly confirms the content that has been displayed and played.

[0774] 10. The emotion engine continuously monitors user responses and updates the emotion state to the server as needed.

[0775] Providing additional information and continuing the interaction

[0776] 11. If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[0777] 12. The device captures the user's voice again, converts it to text data, and sends it to the server.

[0778] 13. The server analyzes the new text data and the latest sentiment analysis results, generates the necessary new content, and sends it to the terminal.

[0779] End of interaction

[0780] 14. When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[0781] In this way, the system provides detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state. This is expected to further pique the user's interest and increase their desire to purchase.

[0782] The following describes the processing flow.

[0783] Step 1:

[0784] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[0785] Step 2:

[0786] The device's microphone captures the user's voice.

[0787] Step 3:

[0788] The device uses speech recognition software to convert the captured audio into text data.

[0789] Step 4:

[0790] The device's camera captures the user's facial expressions.

[0791] Step 5:

[0792] The device uses an emotion engine to analyze captured facial data and recognize emotional states (e.g., excitement, interest, questioning).

[0793] Step 6:

[0794] The terminal sends the converted text data and sentiment analysis results to the server.

[0795] Step 7:

[0796] The server receives text data and performs analysis using an NLP engine. Based on the analysis results, it identifies the information the user is requesting (e.g., smartphone camera functions, battery life, etc.).

[0797] Step 8:

[0798] The server uses a generative AI model, taking sentiment analysis results into account, to generate relevant visual, audio, and video content.

[0799] Step 9:

[0800] The server sends the generated content to the terminal.

[0801] Step 10:

[0802] The device receives the generated content.

[0803] Step 11:

[0804] Visual content is displayed on the device's screen, and audio content is played through the speaker. For example, a photo taken with a high-resolution camera is displayed on the screen, and an audio message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[0805] Step 12:

[0806] Users visually and aurally confirm the displayed content and played audio.

[0807] Step 13:

[0808] The emotion engine continuously monitors the user's reactions and determines their emotional state.

[0809] Step 14:

[0810] If the user requests more detailed information (for example, "How long does the battery last?"), additional voice instructions will be provided.

[0811] Step 15:

[0812] The device captures the user's voice again, converts it into text data using speech recognition software, and sends it to the server.

[0813] Step 16:

[0814] The server receives new text data and the latest sentiment analysis results, and analyzes them again using the NLP engine. Based on the analysis results, it generates the necessary new content and sends it to the terminal.

[0815] Step 17:

[0816] The device receives new content, displays it on the screen, and plays it through the speaker. For example, a graph showing battery life might be displayed, and a voice message might say, "This smartphone's battery lasts 24 hours with normal use."

[0817] Step 18:

[0818] If the user is satisfied, they instruct the robot to end the interaction.

[0819] Step 19:

[0820] The terminal captures the voice command to terminate, and the system enters standby mode.

[0821] (Example 2)

[0822] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0823] Conventional systems, when providing information based on user voice input, lacked the ability to respond in a way that took into account the user's emotional state, resulting in a uniform user experience and low satisfaction. Furthermore, they lacked the ability to integrate and deliver diverse content formats, limiting the effectiveness of information transmission.

[0824] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0825] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for capturing the user's facial expressions and analyzing their emotional state, means for transmitting the text data and emotional state, means for analyzing the text data and generating related content such as visuals, audio, and video, means for displaying and playing the generated content, and means for presenting the generated content to the user visually and aurally. This enables the provision of interactive content that reflects the user's emotional state, thereby improving the quality of the user experience.

[0826] A "user" is an entity that gives instructions or asks questions to a system using voice.

[0827] A "terminal" is a device that performs voice input, voice recognition, facial expression capture, result display, and playback.

[0828] A "server" is a device that receives text data and sentiment analysis results, and performs data analysis and content generation.

[0829] "Voice input means" refers to devices such as microphones that capture the user's voice.

[0830] "Speech recognition software" is a program that converts captured audio into text data.

[0831] "Text data" refers to character information converted by speech recognition software.

[0832] "Facial expression capture means" refers to devices such as cameras that capture the user's facial expressions.

[0833] An "emotion engine" is software that analyzes a user's emotional state from captured facial expressions.

[0834] A "communication protocol" is a set of communication rules for efficiently sending and receiving data between a terminal and a server.

[0835] A "natural language processing (NLP) engine" is a program that analyzes text data to understand the user's intent.

[0836] A "generative AI model" is an algorithm that generates relevant visual, audio, and video content based on data analysis results.

[0837] "Visual content" refers to visual information such as images and graphics.

[0838] "Audio content" refers to auditory information such as spoken language.

[0839] "Video content" refers to information that combines dynamic visual and audio information.

[0840] "Content presentation means" refers to a function that provides generated visual, audio, and video content to the user visually and aurally.

[0841] The system of the present invention enhances the user experience by dynamically generating and presenting relevant visual, audio, and video content based on the user's voice input and emotion analysis. The embodiments for carrying out the present invention will be described in detail below.

[0842] Hardware configuration

[0843] This system consists of the following main hardware components:

[0844] Terminal: Equipped with voice input, display, speaker, and camera, it serves as the main component for interactive demonstrations. Specifically, it includes a high-sensitivity microphone, high-resolution display, speaker, and high-resolution camera.

[0845] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation. Servers with high-performance computing resources are required.

[0846] Software Configuration

[0847] The following software and algorithms operate between the terminal and the server:

[0848] Speech recognition software: Converts voice input from a user into text data. For example, technologies such as Google Cloud Speech-to-Text are used for speech recognition.

[0849] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent. OpenAI GPT-3 is an example of this.

[0850] Generative AI models: These models generate relevant content such as visuals, audio, and video based on analysis results. For example, OpenAI DALL-E and GPT-3 are used.

[0851] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state. For example, the Microsoft Azure Emotion API is used.

[0852] Communication protocol: A protocol that efficiently sends and receives data between a terminal and a server. For example, HTTPS is used.

[0853] System operation

[0854] User voice input and emotion recognition

[0855] The user asks questions about product information using voice.

[0856] The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[0857] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state.

[0858] The device sends text data and sentiment analysis results to the server.

[0859] Data analysis and content generation

[0860] The server analyzes the received text data using a natural language processing engine to identify the user's intent.

[0861] The server uses a generation AI model that takes sentiment analysis results into account to generate content in various formats.

[0862] For example, if a user asks, "What are the features of this smartwatch?", the system will generate a demo video of the heart rate monitoring function and a graph showing battery life.

[0863] Content display and response

[0864] The device receives the generated content, displays the visual content on its screen, and plays the audio content using its speaker.

[0865] For example, a demo video of the heart rate measurement function is displayed on the screen, and a voice guide plays saying, "This smartwatch is capable of accurate heart rate measurement."

[0866] Check the content that users have viewed and played.

[0867] The emotion engine monitors the user's reactions and updates the emotion state to the server as needed.

[0868] Examples of specific cases and prompt statements

[0869] As a concrete example, consider a scenario where a user asks, "Tell me about the features of this smartwatch." In this case, the user's voice is captured by the device's microphone and converted into text data by speech recognition software. The server then analyzes this text data with a natural language processing engine, and a generative AI model generates a demo video of the heart rate measurement function and a graph of battery life. If the emotion engine analyzes the user's facial expressions and recognizes an excited emotional state, it continues with a more detailed explanation of the functions.

[0870] An example of a prompt might be: "Consider a scenario where a user asks, 'What are the features of this smartwatch?' and come up with prompts that would allow the generative AI model to create appropriate visual, audio, and video content."

[0871] As described above, the system of the present invention can provide detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state, thereby significantly improving the quality of the user experience.

[0872] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0873] Step 1:

[0874] The user asks a question about product information using voice. For example, they might say, "Tell me about the features of this smartwatch."

[0875] Input: User voice input

[0876] Output: Captured audio data

[0877] Step 2:

[0878] The device's microphone captures the user's voice, and speech recognition software converts this into text data. For example, Google Cloud Speech-to-Text can be used.

[0879] Input: Captured audio data

[0880] Output: Converted text data (e.g., "Tell me the features of this smartwatch")

[0881] Step 3:

[0882] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state. For example, the Microsoft Azure Emotion API can be used.

[0883] Input: Captured facial expression data

[0884] Output: Analyzed emotional state data (e.g., excitement, interest)

[0885] Step 4:

[0886] The device sends text data and sentiment analysis results to the server. For example, HTTPS is used to efficiently transmit the data.

[0887] Input: Text data and sentiment state data

[0888] Output: Sending data to the server

[0889] Step 5:

[0890] The server analyzes the received text data using a natural language processing (NLP) engine. For example, OpenAI GPT-3 can be used.

[0891] Input: Text data

[0892] Output: Analysis results of user intent (e.g., smartwatch heart rate monitoring function, battery life)

[0893] Step 6:

[0894] The server considers the sentiment analysis results and uses a generative AI model to generate relevant visual, audio, and video content. For example, it might use OpenAI DALL-E or GPT-3.

[0895] Input: Analysis results of user intent and emotional state data

[0896] Output: Generated visual, audio, and video content (e.g., a demo video of the heart rate measurement function, a graph showing battery life)

[0897] Step 7:

[0898] The device receives the generated content from the server.

[0899] Input: Content data sent from the server

[0900] Output: Content data stored on the device

[0901] Step 8:

[0902] The device displays visual content on its screen and plays audio content through its speaker. For example, it might show a demo video of the heart rate measurement function on the screen and play an audio guide through the speaker saying, "This smartwatch is capable of accurate heart rate measurement."

[0903] Input: Content data stored on the device

[0904] Output: Display of visual content and playback of audio content

[0905] Step 9:

[0906] Review the content displayed and played by the user. For example, review the display image and speaker description.

[0907] Input: Visual and audio content

[0908] Output: User understanding and response

[0909] Step 10:

[0910] The emotion engine continuously monitors the user's reactions and updates the emotional state to the server as needed. For example, if the user's facial expression changes to one of surprise, the emotion engine analyzes this and sends the information to the server.

[0911] Input: User's facial expression data

[0912] Output: Updated sentiment state data

[0913] Step 11:

[0914] The user asks additional questions by voice. For example, they might say, "How long does the battery last?"

[0915] Input: Additional voice questions

[0916] Output: Captured audio data

[0917] Step 12:

[0918] The device then captures the user's voice again and converts it into text data using speech recognition software.

[0919] Input: Captured audio data

[0920] Output: Converted text data

[0921] Step 13:

[0922] The server analyzes the new text data and the latest sentiment analysis results to generate the necessary new content.

[0923] Input: New text data and latest sentiment state data

[0924] Output: Newly generated visual, audio, and video content

[0925] Step 14:

[0926] The device receives newly generated content and delivers it to the user using its display and speaker. For example, it might display a graph showing battery life on the display and provide an explanation with voice guidance.

[0927] Input: Newly generated content data

[0928] Output: Display of visual content and playback of audio content

[0929] Step 15:

[0930] The user gives a voice command to end the interaction. For example, they might say, "I'm done."

[0931] Input: Voice command to end

[0932] Output: Captured audio data

[0933] Step 16:

[0934] The terminal captures the voice command to terminate the process, converts it into text data using speech recognition software, and then puts the system into standby mode.

[0935] Input: Voice command to end

[0936] Output: System transition to standby mode

[0937] (Application Example 2)

[0938] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0939] Conventional product information systems in physical stores have problems in providing quick and detailed responses to customer questions, and furthermore, in providing interactive responses that match the customer's emotions. In addition, there is a lack of means for customers to obtain specific product information visually and aurally, making it difficult to increase their purchasing intent. Therefore, the present invention aims to solve these problems and realize more effective and customized information provision to customers.

[0940] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0941] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for analyzing the user's facial expressions and recognizing their emotional state, and means for customizing the generated content according to the user's emotional state and presenting it visually and aurally. As a result, customers can obtain specific, visual, and auditory product information in real time within a physical store, and since that information is customized according to the customer's emotional state, an interactive experience that enhances their desire to purchase becomes possible.

[0942] "Means of inputting voice" refers to devices or software that receive voice signals from a user and process them as digital signals.

[0943] "Methods for converting speech to text data" refer to algorithms or software that analyze an input speech signal and convert it into corresponding text data.

[0944] "Means for analyzing text data and generating related visual, auditory, and video content" refers to a system that uses natural language processing and generative AI models based on text data to generate content in various formats.

[0945] "Means for displaying and playing generated content" refers to a system that outputs generated visual and auditory content to a display device or speaker so that the user can see and hear it.

[0946] "Methods for analyzing a user's facial expressions to recognize their emotional state" refer to software or algorithms that capture a user's facial expressions through cameras or sensors, analyze them, and estimate the user's emotional state.

[0947] "Means for customizing generated content according to the user's emotional state and presenting it visually and aurally" refers to a system that adjusts the format and content of the content considering the user's emotional state and provides information to the user in an appropriate manner.

[0948] One embodiment of the present invention relates to a system for effectively providing product information to customers in a physical store. This system uses voice input, speech recognition, natural language processing, sentiment recognition, and generative AI models to provide customized information in response to customer inquiries.

[0949] Configuration of the main components

[0950] Hardware configuration

[0951] Terminal: An interactive device equipped with voice input, a display, speakers, and a camera. It captures the user's voice and facial expressions and displays and plays the generated content.

[0952] Server: A central system located in the cloud or locally, which performs data analysis and content generation.

[0953] Software Configuration

[0954] The following software and algorithms operate between the terminal and the server:

[0955] Speech recognition software: Converts user speech into text data. Example: SpeechRecognition library.

[0956] Natural Language Processing Engine: Analyzes text data and generates the best possible answer to a question. Example: Transformers in Hugging Face.

[0957] Generative AI models: Generate relevant visual and auditory content. Example: A generative model for Hugging Face.

[0958] Emotion recognition engine: Analyzes the user's facial expressions to recognize their emotional state. Example: EmotionRecognizer.

[0959] Communication protocol: Enables efficient transmission and reception of data between terminals and servers.

[0960] Content generation process

[0961] 1. Voice Input: Users can ask questions about the product using voice. Example: "Please tell me about the features of this smartphone."

[0962] 2. Speech Recognition: The device's microphone captures speech and converts it to text using SpeechRecognition software.

[0963] 3. Text Analysis: The converted text data is sent to the server and analyzed by a natural language processing engine.

[0964] 4. Emotion Recognition: The device's camera captures the user's facial expressions, and the EmotionRecognizer recognizes their emotional state.

[0965] 5. Content Generation: Based on the analysis results and emotional state, the AI ​​generation model generates relevant content such as visuals, audio, and video.

[0966] 6. Content display and playback: The generated content is sent to the device, displayed on the screen, and the audio is played from the speaker.

[0967] Specific example

[0968] For example, if a user asks, "What are the features of this smartphone?", the system will operate as follows:

[0969] Speech recognition software converts the user's voice into text data.

[0970] A natural language processing engine analyzes text data to identify information about the smartphone's features (e.g., camera functions, battery life).

[0971] The generative AI model generates visual content such as high-resolution camera images and graphs showing battery life.

[0972] The emotion recognition engine analyzes the user's facial expressions, and if it determines that the user is in an excited state, it provides additional information such as a voice message saying, "You can take very clear photos."

[0973] Example of a prompt

[0974] "Please tell me about the features of this smartphone."

[0975] "How long does the battery last?"

[0976] This allows customers to receive specific and detailed information in real time, and since that information is customized according to the customer's emotional state, it enables an interactive experience that increases their desire to buy.

[0977] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0978] Step 1:

[0979] Users ask questions about the product using voice.

[0980] Input: User's voice.

[0981] Operation: The device's microphone captures the user's voice.

[0982] Output: Captured audio data.

[0983] Step 2:

[0984] The speech recognition software converts the captured audio data into text data.

[0985] Input: Captured audio data.

[0986] Operation: The device's speech recognition software (e.g., SpeechRecognition library) analyzes the speech data and converts it into corresponding text.

[0987] Output: Converted text data.

[0988] Step 3:

[0989] The terminal sends text data to the server.

[0990] Input: Converted text data.

[0991] Operation: Text data is sent to the server via a communication protocol.

[0992] Output: Text data received by the server.

[0993] Step 4:

[0994] The server's natural language processing engine analyzes the text data and generates the best possible answer to the question.

[0995] Input: Received text data.

[0996] Operation: The server's natural language processing engine (e.g., Hugging Face's Transformers) analyzes the text data, understands the user's intent in the question, and generates an appropriate answer.

[0997] Output: Generated answer text.

[0998] Step 5:

[0999] The device's camera captures the user's facial expressions, and the emotion recognition engine recognizes the user's emotional state.

[1000] Input: User's facial expression image.

[1001] Operation: The device's camera captures the user's facial expressions, and an emotion recognition engine (e.g., EmotionRecognizer) analyzes the facial data to estimate the emotional state.

[1002] Output: Estimated emotional state data.

[1003] Step 6:

[1004] The server uses a generation AI model to generate visual and auditory content based on response text and sentiment state data.

[1005] Input: Generated response text, estimated sentiment state data.

[1006] Operation: The server's generation AI model considers the response text and emotional state to generate relevant visual content (e.g., photos, graphs) and auditory content (e.g., voice guidance).

[1007] Output: Generated visual and auditory content.

[1008] Step 7:

[1009] The generated content is sent to the device, displayed on the screen, and played through the speaker.

[1010] Input: Generated visual and auditory content.

[1011] Operation: Content is sent from the server to the terminal, visual content is displayed on the screen, and audio content is played from the speaker.

[1012] Output: Users visually and aurally perceive the content.

[1013] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1014] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1015] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1016] [Third Embodiment]

[1017] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1018] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1019] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1020] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1021] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1022] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1023] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1024] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1025] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1026] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1027] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1028] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1029] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice instructions via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those instructions.

[1030] System Configuration

[1031] Hardware configuration

[1032] This system consists of the following main components:

[1033] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[1034] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[1035] Software Configuration

[1036] The following software and algorithms operate between the terminal and the server:

[1037] Speech recognition software: Converts voice input from the user into text data.

[1038] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[1039] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[1040] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[1041] System operation

[1042] User voice input

[1043] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[1044] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[1045] 3. The terminal sends text data to the server.

[1046] Data analysis and content generation

[1047] 4. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[1048] 5. The server uses the generated AI model to produce relevant visual, audio, and video content.

[1049] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[1050] Content display and response

[1051] 6. The device receives the generated content.

[1052] 7. Display visual content on the device's screen and play audio content using the speaker.

[1053] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[1054] 8. If the user requests further information (e.g., "How long does the battery last?"), they input voice again, and the process is repeated.

[1055] This allows users to visually and audibly experience the features and benefits of a product through a robot, thereby increasing their desire to purchase. The present invention is a system that, through this interactive process, solves the problems of conventional, limited demonstration methods and enables the effective provision of product information.

[1056] The following describes the processing flow.

[1057] Step 1:

[1058] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[1059] Step 2:

[1060] The terminal (robot) uses its built-in microphone to capture the user's voice.

[1061] Step 3:

[1062] The device uses speech recognition software to convert the captured audio into text data.

[1063] Step 4:

[1064] The terminal sends the converted text data to the server.

[1065] Step 5:

[1066] The server runs a natural language processing (NLP) engine to analyze the text data it receives.

[1067] Step 6:

[1068] The server uses an NLP engine to understand the user's request and identify the necessary information (smartphone features).

[1069] Step 7:

[1070] The server uses a generated AI model to create content such as visuals, audio, and video based on identified information.

[1071] Step 8:

[1072] The server sends the generated content to the terminal.

[1073] Step 9:

[1074] The device displays the received content on its screen and plays the audio through its speaker.

[1075] Step 10:

[1076] The user visually and audibly confirms the content that has been displayed and played.

[1077] Step 11:

[1078] If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[1079] Step 12:

[1080] The device captures the user's voice again, converts it into text data, and sends it to the server.

[1081] Step 13:

[1082] The server analyzes the new text data, generates the necessary new content, and sends it to the terminal.

[1083] Step 14:

[1084] The device displays new content on its screen and plays audio through its speaker.

[1085] Step 15:

[1086] When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[1087] (Example 1)

[1088] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1089] Current product demonstration systems lack the interactivity necessary for users to effectively understand product information. Traditional systems only provide pre-prepared information unilaterally, lacking the flexibility to respond to specific user questions and requests. Furthermore, the lack of integration of visual and auditory content limits the user experience. This can lead to users not fully understanding the product's features and benefits, potentially reducing their purchase intent.

[1090] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1091] In this invention, the server includes means for analyzing voice data and understanding the user's intent, means for generating relevant content such as visuals, audio, and video using a generative AI model, and means for transmitting the generated content to a terminal. This makes it possible to generate and provide appropriate and interactive content in response to the user's specific questions and requests.

[1092] A "user" refers to a person who operates a system and obtains information.

[1093] "Means of inputting voice" refers to devices such as microphones that capture the user's voice and input it into the system.

[1094] "Means of converting to text data" refers to the process or equipment used to convert captured audio into text information.

[1095] "Methods for analyzing and understanding user intent" refers to the process of analyzing text data converted from speech to understand the information and questions the user is seeking.

[1096] A "generative AI model" refers to an artificial intelligence algorithm that learns from large amounts of data and generates new content.

[1097] "Means for generating visual, audio, and video content" refers to processes and devices that generate visual and auditory content for users based on analysis results.

[1098] "Means for sending generated content to a terminal" refers to communication means for transmitting content generated on a server to a terminal.

[1099] "Means of visual and auditory presentation" refers to devices and methods that provide generated content to users visually and aurally, using displays, speakers, etc.

[1100] A "prompt statement" refers to a sentence used as a specific input to a generative AI model.

[1101] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those commands.

[1102] Hardware configuration

[1103] This system consists of the following main components:

[1104] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[1105] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[1106] Software Configuration

[1107] The following software and algorithms operate between the terminal and the server:

[1108] Speech recognition software: Converts voice input from the user into text data.

[1109] Specific software to use: For example, Google Cloud Speech-to-Text API.

[1110] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[1111] Specific software to be used: for example, the Python-based NLTK library or SpaCy.

[1112] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[1113] Specific models to use: For example, OpenAI GPT-3.

[1114] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[1115] Specific protocols to use: HTTP, MQTT, etc.

[1116] System operation

[1117] User voice input

[1118] The user asks a question about product information, and the device's microphone captures the user's voice. Speech recognition software converts this into text data, and the device sends the text data to the server.

[1119] Data analysis and content generation

[1120] The server receives text data, and the NLP engine analyzes it. Based on the analysis, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.). Next, the server uses a generative AI model to generate relevant visual, audio, and video content.

[1121] Examples of generated content: photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[1122] Content display and response

[1123] The device receives the generated content, displays the visual content on the screen, and plays the audio content using the speaker. If the user requests further information, they can input voice again, and the process repeats.

[1124] Specific example

[1125] When a user asks, "What are the features of this smartphone?", the device's microphone captures the audio, and speech recognition software converts it into text data. Next, the text data is sent to a server and analyzed by an NLP engine. Based on the analysis, a generative AI model generates content with the prompt, "Please describe the smartphone's camera features." Then, a photo taken with the high-resolution camera is displayed on the device's screen, and explanatory audio is played through the speaker.

[1126] In this way, users can visually and audibly experience the features and benefits of a product through the robot, thereby increasing their desire to purchase. This system solves the problems of conventional, limited demonstration methods and enables the effective delivery of product information.

[1127] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1128] Step 1:

[1129] A user asks a question about product information. The user's voice is input to the microphone. Example of a user question: "What are the features of this smartphone?" The input is voice data. Specifically, the microphone captures the voice signal and temporarily stores it in internal memory.

[1130] Step 2:

[1131] The device's speech recognition software captures audio and converts it into text data. The audio data is then sent to the Google Cloud Speech-to-Text API to retrieve the text data. The input is audio data, and the output is text data. Specifically, the device sends audio data to an external speech recognition API and receives text data as a response from the API.

[1132] Step 3:

[1133] The device sends the generated text data to the server. The text data is sent to the server's API endpoint using the HTTP protocol. The input is text data, and the output is an HTTP request. Specifically, the device sends the text data to the server as the payload of the HTTP request.

[1134] Step 4:

[1135] The server receives text data and performs analysis using an NLP engine. It uses Python's NLTK and SpaCy libraries to analyze the text data and extract the user's intent. The input is text data, and the output is the analysis result (user intent). Specifically, the server inputs text data into the NLP library and retrieves the analysis result.

[1136] Step 5:

[1137] The server generates relevant visual, audio, and video content using a generative AI model based on the analysis results. A prompt is input to the generative AI model (e.g., OpenAI GPT-3) to generate the necessary content. The input is the analysis results and the prompt, and the output is the generated content. Specifically, a prompt is created and input to the generative AI model to obtain content. Example: Prompt: "Please explain the camera functions of a smartphone."

[1138] Step 6:

[1139] The server sends the generated content to the terminal. The generated visual, audio, and video files are sent to the terminal as an HTTP response. The input is the generated content, and the output is the HTTP response. Specifically, the server sends the generated content to the terminal as an HTTP response.

[1140] Step 7:

[1141] The device receives generated content, displays visual content on the screen, and plays audio content through the speaker. The input is the HTTP response (generated content), and the output is the display of visual and auditory content. Specifically, the display module displays images and videos, and the audio module plays audio. Example: The display shows a photo taken with a high-resolution camera and plays the audio, "This smartphone has a high-resolution camera and can take very clear photos."

[1142] Step 8:

[1143] If the user requests further information, they input voice again, and the process is repeated. The input is new voice data, and the output is updated text data. Specifically, the device captures the user's new question again, and the aforementioned processing steps are executed again. Example: "How long does the battery last?"

[1144] (Application Example 1)

[1145] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1146] In traditional brick-and-mortar stores, customers often relied on sales staff to obtain detailed product information. This method suffers from inconsistencies in information delivery due to variations in staff knowledge and explanation skills. Furthermore, during busy periods, adequate service may be difficult to provide. Additionally, insufficient visual and auditory information makes it difficult for customers to fully understand product features and benefits.

[1147] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1148] In this invention, the server includes means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for presenting the generated content to the user visually and aurally, means for providing relevant content based on questions from customers via a robot installed in a physical store, and means for processing the analyzed and generated content through a cloud or local server. This enables users to receive consistent, high-quality information and to better understand the features and benefits of products.

[1149] "Means of inputting voice" refers to devices or systems that effectively capture voice from users.

[1150] "Methods for converting audio to text data" refer to software or algorithms that extract text information from captured audio and convert it into text data.

[1151] "Means for analyzing text data and generating related visual, auditory, and video content" refers to technologies that understand user intent and questions based on transcribed data and create corresponding visual, auditory, and video content.

[1152] "Means for displaying and playing generated content" refers to devices and mechanisms for showing or letting users hear generated visual, auditory, and video content.

[1153] "Means for presenting the generated content to the user visually and audibly" refers to a system that transmits the generated content to the user through a screen, speakers, or the like.

[1154] "A means of providing relevant content based on customer questions via robots installed in physical stores" refers to a system in which robots placed in actual stores generate and provide various content in response to user questions.

[1155] "Means for processing the analyzed and generated content via a cloud or local server" refers to a technology for efficiently processing the generated and analyzed content on a cloud server or a server on an internal network.

[1156] This invention provides a system that allows customers to receive product information visually and aurally using a robot installed in a physical store. The system comprises means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating relevant visual, auditory, and video content, means for displaying and playing the generated content, means for presenting it visually and aurally, means for providing relevant content based on customer questions via a robot installed in the physical store, and means for processing the generated content through a cloud or local server.

[1157] Hardware configuration

[1158] Terminal (robot body):

[1159] Microphone: A high-sensitivity microphone is used to capture the voices of customers.

[1160] Display: Use a high-resolution display to show the generated visual content.

[1161] Speakers: High-quality speakers are used to play the generated audio content.

[1162] server:

[1163] Cloud servers or local servers: Used for data analysis and content generation.

[1164] Software Configuration

[1165] Software and algorithms that operate between terminals and servers:

[1166] Speech recognition software: Uses the Google Speech-to-Text API to convert user speech into text data.

[1167] Natural Language Processing (NLP) Engine: Uses SpaCy or Google Natural Language API to analyze text data and understand user intent.

[1168] Generative AI Model: Generates relevant visual, auditory, and video content using OpenAI GPT-4, DALL-E, and a TTS engine.

[1169] Specific examples of actions

[1170] 1. A customer asks, "Could you tell me about the features of this TV?"

[1171] The voices of customers are captured using a microphone.

[1172] The captured audio is converted into text data using the Google Speech-to-Text API.

[1173] Using an NLP engine, text data is analyzed to understand the customer's intentions.

[1174] Relevant visual, auditory, and video content is generated using generative AI models (GPT-4, DALL-E, TTS engine).

[1175] 2. The generated content is displayed on the robot's screen and played back as audio through its speaker.

[1176] Example: A high-resolution television image is displayed on the screen, and a voice message plays from the speaker saying, "This television is 4K compatible, so you can enjoy high-resolution images."

[1177] A short demo video is played to visually demonstrate the video quality.

[1178] Example of a prompt:

[1179] Input: "Please tell me about the features of this television."

[1180] Prompt message:

[1181] Please provide the following information, including the features of your television:

[1182] 1. Resolution

[1183] 2. Special features (4K, HDR, smart features, etc.)

[1184] 3. Size

[1185] 4. Price range

[1186] 5. Buyer Reviews

[1187] Proposed output:

[1188] This TV is 4K compatible, allowing you to enjoy high-resolution images. It also features HDR functionality, providing vibrant colors. It's a 55-inch model and priced at approximately 100,000 yen. It has received high praise from many buyers.

[1189] In this way, customers can receive detailed product information through the robot. The entire system is managed via the cloud or a local server, enabling efficient and effective information delivery.

[1190] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1191] Step 1:

[1192] Process name: Voice input capture

[1193] Subject: User

[1194] Specific actions:

[1195] The user asks a question to the robot. An example message might be, "Please tell me about the features of this television." A high-sensitivity microphone built into the robot captures this audio.

[1196] Input: User's voice

[1197] Output: Audio data

[1198] Step 2:

[1199] Process name: Text conversion of audio data

[1200] Subject: terminal

[1201] Specific actions:

[1202] The captured audio data is converted into text data in real time by the device's speech recognition software (Google Speech-to-Text API).

[1203] Input: Audio data

[1204] Output: Text data

[1205] Step 3:

[1206] Process name: Sending text data

[1207] Subject: terminal

[1208] Specific actions:

[1209] Text data is sent from the terminal to the server. A secure communication protocol is used for this transmission.

[1210] Input: Text data

[1211] Output: Text data sent to the server

[1212] Step 4:

[1213] Process name: Text data analysis

[1214] Subject: Server

[1215] Specific actions:

[1216] The server analyzes the received text data using a natural language processing (NLP) engine (SpaCy or Google Natural Language API) to understand the user's intent and requests. Specifically, if the question is related to a product, it identifies information such as the product's features, benefits, price, and ratings.

[1217] Input: Text data sent to the server

[1218] Output: Analyzed text data (information such as product features and benefits)

[1219] Step 5:

[1220] Process name: Content generation

[1221] Subject: Server

[1222] Specific actions:

[1223] The server uses a generation AI model (GPT-4, DALL-E, TTS engine) to generate visual, auditory, and video content based on the analyzed text data. Examples include high-resolution images of the product, graphs and explanatory audio demonstrating product features, and demo videos.

[1224] Input: Analyzed text data

[1225] Output: Generated visual, auditory, and video content

[1226] Step 6:

[1227] Process name: Send content

[1228] Subject: Server

[1229] Specific actions:

[1230] The generated visual, auditory, and video content is transmitted from the server to the terminal. Protocols are used to maintain data integrity and security during this process.

[1231] Input: Generated visual, auditory, and video content

[1232] Output: Content sent to the device

[1233] Step 7:

[1234] Process name: Display and play content

[1235] Subject: Terminal (robot body)

[1236] Specific actions:

[1237] The device displays and plays received content. Images and graphs are shown on a high-resolution display, and explanatory audio is played using high-quality speakers. Demo videos are also played on the display.

[1238] For example, a voice explanation such as, "This TV is 4K compatible, allowing you to enjoy high-resolution images," is played while product images and demo videos are displayed.

[1239] Input: Content sent to the device

[1240] Output: Visual and auditory content displayed and played for the user.

[1241] This series of steps allows customers to interactively obtain product information.

[1242] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1243] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands through a robot, which then generates and presents various forms of data (images, audio, video, etc.) based on those commands. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it enables interactive responses that correspond to the user's emotional state.

[1244] System Configuration

[1245] Hardware configuration

[1246] This system consists of the following main components:

[1247] Terminal (robot body): Equipped with voice input means, display, speaker, and camera, it serves as the main component for interactive demonstrations.

[1248] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[1249] Software Configuration

[1250] The following software and algorithms operate between the terminal and the server:

[1251] Speech recognition software: Converts voice input from the user into text data.

[1252] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[1253] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[1254] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state.

[1255] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[1256] System operation

[1257] User voice input and emotion recognition

[1258] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[1259] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[1260] 3. The device uses its camera to capture the user's facial expressions, and the emotion engine analyzes this to recognize the emotional state (e.g., excitement, interest, questioning, etc.).

[1261] 4. The device sends text data and sentiment analysis results to the server.

[1262] Data analysis and content generation

[1263] 5. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[1264] 6. The server takes the sentiment analysis results into account and uses a generative AI model to generate relevant visual, audio, and video content.

[1265] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[1266] Content display and response

[1267] 7. The device receives the generated content.

[1268] 8. Display visual content on the device's screen and play audio content using the speaker.

[1269] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[1270] 9. The user visually and audibly confirms the content that has been displayed and played.

[1271] 10. The emotion engine continuously monitors user responses and updates the emotion state to the server as needed.

[1272] Providing additional information and continuing the interaction

[1273] 11. If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[1274] 12. The device captures the user's voice again, converts it to text data, and sends it to the server.

[1275] 13. The server analyzes the new text data and the latest sentiment analysis results, generates the necessary new content, and sends it to the terminal.

[1276] End of interaction

[1277] 14. When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[1278] In this way, the system provides detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state. This is expected to further pique the user's interest and increase their desire to purchase.

[1279] The following describes the processing flow.

[1280] Step 1:

[1281] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[1282] Step 2:

[1283] The device's microphone captures the user's voice.

[1284] Step 3:

[1285] The device uses speech recognition software to convert the captured audio into text data.

[1286] Step 4:

[1287] The device's camera captures the user's facial expressions.

[1288] Step 5:

[1289] The device uses an emotion engine to analyze captured facial data and recognize emotional states (e.g., excitement, interest, questioning).

[1290] Step 6:

[1291] The terminal sends the converted text data and sentiment analysis results to the server.

[1292] Step 7:

[1293] The server receives text data and performs analysis using an NLP engine. Based on the analysis results, it identifies the information the user is requesting (e.g., smartphone camera functions, battery life, etc.).

[1294] Step 8:

[1295] The server uses a generative AI model, taking sentiment analysis results into account, to generate relevant visual, audio, and video content.

[1296] Step 9:

[1297] The server sends the generated content to the terminal.

[1298] Step 10:

[1299] The device receives the generated content.

[1300] Step 11:

[1301] Visual content is displayed on the device's screen, and audio content is played through the speaker. For example, a photo taken with a high-resolution camera is displayed on the screen, and an audio message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[1302] Step 12:

[1303] Users visually and aurally confirm the displayed content and played audio.

[1304] Step 13:

[1305] The emotion engine continuously monitors the user's reactions and determines their emotional state.

[1306] Step 14:

[1307] If the user requests more detailed information (for example, "How long does the battery last?"), additional voice instructions will be provided.

[1308] Step 15:

[1309] The device captures the user's voice again, converts it into text data using speech recognition software, and sends it to the server.

[1310] Step 16:

[1311] The server receives new text data and the latest sentiment analysis results, and analyzes them again using the NLP engine. Based on the analysis results, it generates the necessary new content and sends it to the terminal.

[1312] Step 17:

[1313] The device receives new content, displays it on the screen, and plays it through the speaker. For example, a graph showing battery life might be displayed, and a voice message might say, "This smartphone's battery lasts 24 hours with normal use."

[1314] Step 18:

[1315] If the user is satisfied, they instruct the robot to end the interaction.

[1316] Step 19:

[1317] The terminal captures the voice command to terminate, and the system enters standby mode.

[1318] (Example 2)

[1319] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1320] Conventional systems, when providing information based on user voice input, lacked the ability to respond in a way that took into account the user's emotional state, resulting in a uniform user experience and low satisfaction. Furthermore, they lacked the ability to integrate and deliver diverse content formats, limiting the effectiveness of information transmission.

[1321] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1322] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for capturing the user's facial expressions and analyzing their emotional state, means for transmitting the text data and emotional state, means for analyzing the text data and generating related content such as visuals, audio, and video, means for displaying and playing the generated content, and means for presenting the generated content to the user visually and aurally. This enables the provision of interactive content that reflects the user's emotional state, thereby improving the quality of the user experience.

[1323] A "user" is an entity that gives instructions or asks questions to a system using voice.

[1324] A "terminal" is a device that performs voice input, voice recognition, facial expression capture, result display, and playback.

[1325] A "server" is a device that receives text data and sentiment analysis results, and performs data analysis and content generation.

[1326] "Voice input means" refers to devices such as microphones that capture the user's voice.

[1327] "Speech recognition software" is a program that converts captured audio into text data.

[1328] "Text data" refers to character information converted by speech recognition software.

[1329] "Facial expression capture means" refers to devices such as cameras that capture the user's facial expressions.

[1330] An "emotion engine" is software that analyzes a user's emotional state from captured facial expressions.

[1331] A "communication protocol" is a set of communication rules for efficiently sending and receiving data between a terminal and a server.

[1332] A "natural language processing (NLP) engine" is a program that analyzes text data to understand the user's intent.

[1333] A "generative AI model" is an algorithm that generates relevant visual, audio, and video content based on data analysis results.

[1334] "Visual content" refers to visual information such as images and graphics.

[1335] "Audio content" refers to auditory information such as spoken language.

[1336] "Video content" refers to information that combines dynamic visual and audio information.

[1337] "Content presentation means" refers to a function that provides generated visual, audio, and video content to the user visually and aurally.

[1338] The system of the present invention enhances the user experience by dynamically generating and presenting relevant visual, audio, and video content based on the user's voice input and emotion analysis. The embodiments for carrying out the present invention will be described in detail below.

[1339] Hardware configuration

[1340] This system consists of the following main hardware components:

[1341] Terminal: Equipped with voice input, display, speaker, and camera, it serves as the main component for interactive demonstrations. Specifically, it includes a high-sensitivity microphone, high-resolution display, speaker, and high-resolution camera.

[1342] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation. Servers with high-performance computing resources are required.

[1343] Software Configuration

[1344] The following software and algorithms operate between the terminal and the server:

[1345] Speech recognition software: Converts voice input from a user into text data. For example, technologies such as Google Cloud Speech-to-Text are used for speech recognition.

[1346] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent. OpenAI GPT-3 is an example of this.

[1347] Generative AI models: These models generate relevant content such as visuals, audio, and video based on analysis results. For example, OpenAI DALL-E and GPT-3 are used.

[1348] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state. For example, the Microsoft Azure Emotion API is used.

[1349] Communication protocol: A protocol that efficiently sends and receives data between a terminal and a server. For example, HTTPS is used.

[1350] System operation

[1351] User voice input and emotion recognition

[1352] The user asks questions about product information using voice.

[1353] The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[1354] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state.

[1355] The device sends text data and sentiment analysis results to the server.

[1356] Data analysis and content generation

[1357] The server analyzes the received text data using a natural language processing engine to identify the user's intent.

[1358] The server uses a generation AI model that takes sentiment analysis results into account to generate content in various formats.

[1359] For example, if a user asks, "What are the features of this smartwatch?", the system will generate a demo video of the heart rate monitoring function and a graph showing battery life.

[1360] Content display and response

[1361] The device receives the generated content, displays the visual content on its screen, and plays the audio content using its speaker.

[1362] For example, a demo video of the heart rate measurement function is displayed on the screen, and a voice guide plays saying, "This smartwatch is capable of accurate heart rate measurement."

[1363] Check the content that users have viewed and played.

[1364] The emotion engine monitors the user's reactions and updates the emotion state to the server as needed.

[1365] Examples of specific cases and prompt statements

[1366] As a concrete example, consider a scenario where a user asks, "Tell me about the features of this smartwatch." In this case, the user's voice is captured by the device's microphone and converted into text data by speech recognition software. The server then analyzes this text data with a natural language processing engine, and a generative AI model generates a demo video of the heart rate measurement function and a graph of battery life. If the emotion engine analyzes the user's facial expressions and recognizes an excited emotional state, it continues with a more detailed explanation of the functions.

[1367] An example of a prompt might be: "Consider a scenario where a user asks, 'What are the features of this smartwatch?' and come up with prompts that would allow the generative AI model to create appropriate visual, audio, and video content."

[1368] As described above, the system of the present invention can provide detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state, thereby significantly improving the quality of the user experience.

[1369] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1370] Step 1:

[1371] The user asks a question about product information using voice. For example, they might say, "Tell me about the features of this smartwatch."

[1372] Input: User voice input

[1373] Output: Captured audio data

[1374] Step 2:

[1375] The device's microphone captures the user's voice, and speech recognition software converts this into text data. For example, Google Cloud Speech-to-Text can be used.

[1376] Input: Captured audio data

[1377] Output: Converted text data (e.g., "Tell me the features of this smartwatch")

[1378] Step 3:

[1379] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state. For example, the Microsoft Azure Emotion API can be used.

[1380] Input: Captured facial expression data

[1381] Output: Analyzed emotional state data (e.g., excitement, interest)

[1382] Step 4:

[1383] The device sends text data and sentiment analysis results to the server. For example, HTTPS is used to efficiently transmit the data.

[1384] Input: Text data and sentiment state data

[1385] Output: Sending data to the server

[1386] Step 5:

[1387] The server analyzes the received text data using a natural language processing (NLP) engine. For example, OpenAI GPT-3 can be used.

[1388] Input: Text data

[1389] Output: Analysis results of user intent (e.g., smartwatch heart rate monitoring function, battery life)

[1390] Step 6:

[1391] The server considers the sentiment analysis results and uses a generative AI model to generate relevant visual, audio, and video content. For example, it might use OpenAI DALL-E or GPT-3.

[1392] Input: Analysis results of user intent and emotional state data

[1393] Output: Generated visual, audio, and video content (e.g., a demo video of the heart rate measurement function, a graph showing battery life)

[1394] Step 7:

[1395] The device receives the generated content from the server.

[1396] Input: Content data sent from the server

[1397] Output: Content data stored on the device

[1398] Step 8:

[1399] The device displays visual content on its screen and plays audio content through its speaker. For example, it might show a demo video of the heart rate measurement function on the screen and play an audio guide through the speaker saying, "This smartwatch is capable of accurate heart rate measurement."

[1400] Input: Content data stored on the device

[1401] Output: Display of visual content and playback of audio content

[1402] Step 9:

[1403] Review the content displayed and played by the user. For example, review the display image and speaker description.

[1404] Input: Visual and audio content

[1405] Output: User understanding and response

[1406] Step 10:

[1407] The emotion engine continuously monitors the user's reactions and updates the emotional state to the server as needed. For example, if the user's facial expression changes to one of surprise, the emotion engine analyzes this and sends the information to the server.

[1408] Input: User's facial expression data

[1409] Output: Updated sentiment state data

[1410] Step 11:

[1411] The user asks additional questions by voice. For example, they might say, "How long does the battery last?"

[1412] Input: Additional voice questions

[1413] Output: Captured audio data

[1414] Step 12:

[1415] The device then captures the user's voice again and converts it into text data using speech recognition software.

[1416] Input: Captured audio data

[1417] Output: Converted text data

[1418] Step 13:

[1419] The server analyzes the new text data and the latest sentiment analysis results to generate the necessary new content.

[1420] Input: New text data and latest sentiment state data

[1421] Output: Newly generated visual, audio, and video content

[1422] Step 14:

[1423] The device receives newly generated content and delivers it to the user using its display and speaker. For example, it might display a graph showing battery life on the display and provide an explanation with voice guidance.

[1424] Input: Newly generated content data

[1425] Output: Display of visual content and playback of audio content

[1426] Step 15:

[1427] The user gives a voice command to end the interaction. For example, they might say, "I'm done."

[1428] Input: Voice command to end

[1429] Output: Captured audio data

[1430] Step 16:

[1431] The terminal captures the voice command to terminate the process, converts it into text data using speech recognition software, and then puts the system into standby mode.

[1432] Input: Voice command to end

[1433] Output: System transition to standby mode

[1434] (Application Example 2)

[1435] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1436] Conventional product information systems in physical stores have problems in providing quick and detailed responses to customer questions, and furthermore, in providing interactive responses that match the customer's emotions. In addition, there is a lack of means for customers to obtain specific product information visually and aurally, making it difficult to increase their purchasing intent. Therefore, the present invention aims to solve these problems and realize more effective and customized information provision to customers.

[1437] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1438] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for analyzing the user's facial expressions and recognizing their emotional state, and means for customizing the generated content according to the user's emotional state and presenting it visually and aurally. As a result, customers can obtain specific, visual, and auditory product information in real time within a physical store, and since that information is customized according to the customer's emotional state, an interactive experience that enhances their desire to purchase becomes possible.

[1439] "Means of inputting voice" refers to devices or software that receive voice signals from a user and process them as digital signals.

[1440] "Methods for converting speech to text data" refer to algorithms or software that analyze an input speech signal and convert it into corresponding text data.

[1441] "Means for analyzing text data and generating related visual, auditory, and video content" refers to a system that uses natural language processing and generative AI models based on text data to generate content in various formats.

[1442] "Means for displaying and playing generated content" refers to a system that outputs generated visual and auditory content to a display device or speaker so that the user can see and hear it.

[1443] "Methods for analyzing a user's facial expressions to recognize their emotional state" refer to software or algorithms that capture a user's facial expressions through cameras or sensors, analyze them, and estimate the user's emotional state.

[1444] "Means for customizing generated content according to the user's emotional state and presenting it visually and aurally" refers to a system that adjusts the format and content of the content considering the user's emotional state and provides information to the user in an appropriate manner.

[1445] One embodiment of the present invention relates to a system for effectively providing product information to customers in a physical store. This system uses voice input, speech recognition, natural language processing, sentiment recognition, and generative AI models to provide customized information in response to customer inquiries.

[1446] Configuration of the main components

[1447] Hardware configuration

[1448] Terminal: An interactive device equipped with voice input, a display, speakers, and a camera. It captures the user's voice and facial expressions and displays and plays the generated content.

[1449] Server: A central system located in the cloud or locally, which performs data analysis and content generation.

[1450] Software Configuration

[1451] The following software and algorithms operate between the terminal and the server:

[1452] Speech recognition software: Converts user speech into text data. Example: SpeechRecognition library.

[1453] Natural Language Processing Engine: Analyzes text data and generates the best possible answer to a question. Example: Transformers in Hugging Face.

[1454] Generative AI models: Generate relevant visual and auditory content. Example: A generative model for Hugging Face.

[1455] Emotion recognition engine: Analyzes the user's facial expressions to recognize their emotional state. Example: EmotionRecognizer.

[1456] Communication protocol: Enables efficient transmission and reception of data between terminals and servers.

[1457] Content generation process

[1458] 1. Voice Input: Users can ask questions about the product using voice. Example: "Please tell me about the features of this smartphone."

[1459] 2. Speech Recognition: The device's microphone captures speech and converts it to text using SpeechRecognition software.

[1460] 3. Text Analysis: The converted text data is sent to the server and analyzed by a natural language processing engine.

[1461] 4. Emotion Recognition: The device's camera captures the user's facial expressions, and the EmotionRecognizer recognizes their emotional state.

[1462] 5. Content Generation: Based on the analysis results and emotional state, the AI ​​generation model generates relevant content such as visuals, audio, and video.

[1463] 6. Content display and playback: The generated content is sent to the device, displayed on the screen, and the audio is played from the speaker.

[1464] Specific example

[1465] For example, if a user asks, "What are the features of this smartphone?", the system will operate as follows:

[1466] Speech recognition software converts the user's voice into text data.

[1467] A natural language processing engine analyzes text data to identify information about the smartphone's features (e.g., camera functions, battery life).

[1468] The generative AI model generates visual content such as high-resolution camera images and graphs showing battery life.

[1469] The emotion recognition engine analyzes the user's facial expressions, and if it determines that the user is in an excited state, it provides additional information such as a voice message saying, "You can take very clear photos."

[1470] Example of a prompt

[1471] "Please tell me about the features of this smartphone."

[1472] "How long does the battery last?"

[1473] This allows customers to receive specific and detailed information in real time, and since that information is customized according to the customer's emotional state, it enables an interactive experience that increases their desire to buy.

[1474] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1475] Step 1:

[1476] Users ask questions about the product using voice.

[1477] Input: User's voice.

[1478] Operation: The device's microphone captures the user's voice.

[1479] Output: Captured audio data.

[1480] Step 2:

[1481] The speech recognition software converts the captured audio data into text data.

[1482] Input: Captured audio data.

[1483] Operation: The device's speech recognition software (e.g., SpeechRecognition library) analyzes the speech data and converts it into corresponding text.

[1484] Output: Converted text data.

[1485] Step 3:

[1486] The terminal sends text data to the server.

[1487] Input: Converted text data.

[1488] Operation: Text data is sent to the server via a communication protocol.

[1489] Output: Text data received by the server.

[1490] Step 4:

[1491] The server's natural language processing engine analyzes the text data and generates the best possible answer to the question.

[1492] Input: Received text data.

[1493] Operation: The server's natural language processing engine (e.g., Hugging Face's Transformers) analyzes the text data, understands the user's intent in the question, and generates an appropriate answer.

[1494] Output: Generated answer text.

[1495] Step 5:

[1496] The device's camera captures the user's facial expressions, and the emotion recognition engine recognizes the user's emotional state.

[1497] Input: User's facial expression image.

[1498] Operation: The device's camera captures the user's facial expressions, and an emotion recognition engine (e.g., EmotionRecognizer) analyzes the facial data to estimate the emotional state.

[1499] Output: Estimated emotional state data.

[1500] Step 6:

[1501] The server uses a generation AI model to generate visual and auditory content based on response text and sentiment state data.

[1502] Input: Generated response text, estimated sentiment state data.

[1503] Operation: The server's generation AI model considers the response text and emotional state to generate relevant visual content (e.g., photos, graphs) and auditory content (e.g., voice guidance).

[1504] Output: Generated visual and auditory content.

[1505] Step 7:

[1506] The generated content is sent to the device, displayed on the screen, and played through the speaker.

[1507] Input: Generated visual and auditory content.

[1508] Operation: Content is sent from the server to the terminal, visual content is displayed on the screen, and audio content is played from the speaker.

[1509] Output: Users visually and aurally perceive the content.

[1510] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1511] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1512] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1513] [Fourth Embodiment]

[1514] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1515] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1516] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1517] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1518] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1519] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1520] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1521] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1522] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1523] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1524] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1525] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1526] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1527] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice instructions via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those instructions.

[1528] System Configuration

[1529] Hardware configuration

[1530] This system consists of the following main components:

[1531] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[1532] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[1533] Software Configuration

[1534] The following software and algorithms operate between the terminal and the server:

[1535] Speech recognition software: Converts voice input from the user into text data.

[1536] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[1537] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[1538] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[1539] System operation

[1540] User voice input

[1541] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[1542] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[1543] 3. The terminal sends text data to the server.

[1544] Data analysis and content generation

[1545] 4. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[1546] 5. The server uses the generated AI model to produce relevant visual, audio, and video content.

[1547] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[1548] Content display and response

[1549] 6. The device receives the generated content.

[1550] 7. Display visual content on the device's screen and play audio content using the speaker.

[1551] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[1552] 8. If the user requests further information (e.g., "How long does the battery last?"), they input voice again, and the process is repeated.

[1553] This allows users to visually and audibly experience the features and benefits of a product through a robot, thereby increasing their desire to purchase. The present invention is a system that, through this interactive process, solves the problems of conventional, limited demonstration methods and enables the effective provision of product information.

[1554] The following describes the processing flow.

[1555] Step 1:

[1556] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[1557] Step 2:

[1558] The terminal (robot) uses its built-in microphone to capture the user's voice.

[1559] Step 3:

[1560] The device uses speech recognition software to convert the captured audio into text data.

[1561] Step 4:

[1562] The terminal sends the converted text data to the server.

[1563] Step 5:

[1564] The server runs a natural language processing (NLP) engine to analyze the text data it receives.

[1565] Step 6:

[1566] The server uses an NLP engine to understand the user's request and identify the necessary information (smartphone features).

[1567] Step 7:

[1568] The server uses a generated AI model to create content such as visuals, audio, and video based on identified information.

[1569] Step 8:

[1570] The server sends the generated content to the terminal.

[1571] Step 9:

[1572] The device displays the received content on its screen and plays the audio through its speaker.

[1573] Step 10:

[1574] The user visually and audibly confirms the content that has been displayed and played.

[1575] Step 11:

[1576] If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[1577] Step 12:

[1578] The device captures the user's voice again, converts it into text data, and sends it to the server.

[1579] Step 13:

[1580] The server analyzes the new text data, generates the necessary new content, and sends it to the terminal.

[1581] Step 14:

[1582] The device displays new content on its screen and plays audio through its speaker.

[1583] Step 15:

[1584] When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[1585] (Example 1)

[1586] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1587] Current product demonstration systems lack the interactivity necessary for users to effectively understand product information. Traditional systems only provide pre-prepared information unilaterally, lacking the flexibility to respond to specific user questions and requests. Furthermore, the lack of integration of visual and auditory content limits the user experience. This can lead to users not fully understanding the product's features and benefits, potentially reducing their purchase intent.

[1588] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1589] In this invention, the server includes means for analyzing voice data and understanding the user's intent, means for generating relevant content such as visuals, audio, and video using a generative AI model, and means for transmitting the generated content to a terminal. This makes it possible to generate and provide appropriate and interactive content in response to the user's specific questions and requests.

[1590] A "user" refers to a person who operates a system and obtains information.

[1591] "Means of inputting voice" refers to devices such as microphones that capture the user's voice and input it into the system.

[1592] "Means of converting to text data" refers to the process or equipment used to convert captured audio into text information.

[1593] "Methods for analyzing and understanding user intent" refers to the process of analyzing text data converted from speech to understand the information and questions the user is seeking.

[1594] A "generative AI model" refers to an artificial intelligence algorithm that learns from large amounts of data and generates new content.

[1595] "Means for generating visual, audio, and video content" refers to processes and devices that generate visual and auditory content for users based on analysis results.

[1596] "Means for sending generated content to a terminal" refers to communication means for transmitting content generated on a server to a terminal.

[1597] "Means of visual and auditory presentation" refers to devices and methods that provide generated content to users visually and aurally, using displays, speakers, etc.

[1598] A "prompt statement" refers to a sentence used as a specific input to a generative AI model.

[1599] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands via a robot, and for the robot to generate and present various forms of data (images, audio, video, etc.) based on those commands.

[1600] Hardware configuration

[1601] This system consists of the following main components:

[1602] Terminal (robot body): Equipped with voice input means, display, and speaker, it serves as the main component for interactive demonstrations.

[1603] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[1604] Software Configuration

[1605] The following software and algorithms operate between the terminal and the server:

[1606] Speech recognition software: Converts voice input from the user into text data.

[1607] Specific software to use: For example, Google Cloud Speech-to-Text API.

[1608] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[1609] Specific software to be used: for example, the Python-based NLTK library or SpaCy.

[1610] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[1611] Specific models to use: For example, OpenAI GPT-3.

[1612] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[1613] Specific protocols to use: HTTP, MQTT, etc.

[1614] System operation

[1615] User voice input

[1616] The user asks a question about product information, and the device's microphone captures the user's voice. Speech recognition software converts this into text data, and the device sends the text data to the server.

[1617] Data analysis and content generation

[1618] The server receives text data, and the NLP engine analyzes it. Based on the analysis, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.). Next, the server uses a generative AI model to generate relevant visual, audio, and video content.

[1619] Examples of generated content: photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[1620] Content display and response

[1621] The device receives the generated content, displays the visual content on the screen, and plays the audio content using the speaker. If the user requests further information, they can input voice again, and the process repeats.

[1622] Specific example

[1623] When a user asks, "What are the features of this smartphone?", the device's microphone captures the audio, and speech recognition software converts it into text data. Next, the text data is sent to a server and analyzed by an NLP engine. Based on the analysis, a generative AI model generates content with the prompt, "Please describe the smartphone's camera features." Then, a photo taken with the high-resolution camera is displayed on the device's screen, and explanatory audio is played through the speaker.

[1624] In this way, users can visually and audibly experience the features and benefits of a product through the robot, thereby increasing their desire to purchase. This system solves the problems of conventional, limited demonstration methods and enables the effective delivery of product information.

[1625] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1626] Step 1:

[1627] A user asks a question about product information. The user's voice is input to the microphone. Example of a user question: "What are the features of this smartphone?" The input is voice data. Specifically, the microphone captures the voice signal and temporarily stores it in internal memory.

[1628] Step 2:

[1629] The device's speech recognition software captures audio and converts it into text data. The audio data is then sent to the Google Cloud Speech-to-Text API to retrieve the text data. The input is audio data, and the output is text data. Specifically, the device sends audio data to an external speech recognition API and receives text data as a response from the API.

[1630] Step 3:

[1631] The device sends the generated text data to the server. The text data is sent to the server's API endpoint using the HTTP protocol. The input is text data, and the output is an HTTP request. Specifically, the device sends the text data to the server as the payload of the HTTP request.

[1632] Step 4:

[1633] The server receives text data and performs analysis using an NLP engine. It uses Python's NLTK and SpaCy libraries to analyze the text data and extract the user's intent. The input is text data, and the output is the analysis result (user intent). Specifically, the server inputs text data into the NLP library and retrieves the analysis result.

[1634] Step 5:

[1635] The server generates relevant visual, audio, and video content using a generative AI model based on the analysis results. A prompt is input to the generative AI model (e.g., OpenAI GPT-3) to generate the necessary content. The input is the analysis results and the prompt, and the output is the generated content. Specifically, a prompt is created and input to the generative AI model to obtain content. Example: Prompt: "Please explain the camera functions of a smartphone."

[1636] Step 6:

[1637] The server sends the generated content to the terminal. The generated visual, audio, and video files are sent to the terminal as an HTTP response. The input is the generated content, and the output is the HTTP response. Specifically, the server sends the generated content to the terminal as an HTTP response.

[1638] Step 7:

[1639] The device receives generated content, displays visual content on the screen, and plays audio content through the speaker. The input is the HTTP response (generated content), and the output is the display of visual and auditory content. Specifically, the display module displays images and videos, and the audio module plays audio. Example: The display shows a photo taken with a high-resolution camera and plays the audio, "This smartphone has a high-resolution camera and can take very clear photos."

[1640] Step 8:

[1641] If the user requests further information, they input voice again, and the process is repeated. The input is new voice data, and the output is updated text data. Specifically, the device captures the user's new question again, and the aforementioned processing steps are executed again. Example: "How long does the battery last?"

[1642] (Application Example 1)

[1643] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1644] In traditional brick-and-mortar stores, customers often relied on sales staff to obtain detailed product information. This method suffers from inconsistencies in information delivery due to variations in staff knowledge and explanation skills. Furthermore, during busy periods, adequate service may be difficult to provide. Additionally, insufficient visual and auditory information makes it difficult for customers to fully understand product features and benefits.

[1645] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1646] In this invention, the server includes means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for presenting the generated content to the user visually and aurally, means for providing relevant content based on questions from customers via a robot installed in a physical store, and means for processing the analyzed and generated content through a cloud or local server. This enables users to receive consistent, high-quality information and to better understand the features and benefits of products.

[1647] "Means of inputting voice" refers to devices or systems that effectively capture voice from users.

[1648] "Methods for converting audio to text data" refer to software or algorithms that extract text information from captured audio and convert it into text data.

[1649] "Means for analyzing text data and generating related visual, auditory, and video content" refers to technologies that understand user intent and questions based on transcribed data and create corresponding visual, auditory, and video content.

[1650] "Means for displaying and playing generated content" refers to devices and mechanisms for showing or letting users hear generated visual, auditory, and video content.

[1651] "Means for presenting the generated content to the user visually and audibly" refers to a system that transmits the generated content to the user through a screen, speakers, or the like.

[1652] "A means of providing relevant content based on customer questions via robots installed in physical stores" refers to a system in which robots placed in actual stores generate and provide various content in response to user questions.

[1653] "Means for processing the analyzed and generated content via a cloud or local server" refers to a technology for efficiently processing the generated and analyzed content on a cloud server or a server on an internal network.

[1654] This invention provides a system that allows customers to receive product information visually and aurally using a robot installed in a physical store. The system comprises means for inputting voice, means for converting voice into text data, means for analyzing the text data and generating relevant visual, auditory, and video content, means for displaying and playing the generated content, means for presenting it visually and aurally, means for providing relevant content based on customer questions via a robot installed in the physical store, and means for processing the generated content through a cloud or local server.

[1655] Hardware configuration

[1656] Terminal (robot body):

[1657] Microphone: A high-sensitivity microphone is used to capture the voices of customers.

[1658] Display: Use a high-resolution display to show the generated visual content.

[1659] Speakers: High-quality speakers are used to play the generated audio content.

[1660] server:

[1661] Cloud servers or local servers: Used for data analysis and content generation.

[1662] Software Configuration

[1663] Software and algorithms that operate between terminals and servers:

[1664] Speech recognition software: Uses the Google Speech-to-Text API to convert user speech into text data.

[1665] Natural Language Processing (NLP) Engine: Uses SpaCy or Google Natural Language API to analyze text data and understand user intent.

[1666] Generative AI Model: Generates relevant visual, auditory, and video content using OpenAI GPT-4, DALL-E, and a TTS engine.

[1667] Specific examples of actions

[1668] 1. A customer asks, "Could you tell me about the features of this TV?"

[1669] The voices of customers are captured using a microphone.

[1670] The captured audio is converted into text data using the Google Speech-to-Text API.

[1671] Using an NLP engine, text data is analyzed to understand the customer's intentions.

[1672] Relevant visual, auditory, and video content is generated using generative AI models (GPT-4, DALL-E, TTS engine).

[1673] 2. The generated content is displayed on the robot's screen and played back as audio through its speaker.

[1674] Example: A high-resolution television image is displayed on the screen, and a voice message plays from the speaker saying, "This television is 4K compatible, so you can enjoy high-resolution images."

[1675] A short demo video is played to visually demonstrate the video quality.

[1676] Example of a prompt:

[1677] Input: "Please tell me about the features of this television."

[1678] Prompt message:

[1679] Please provide the following information, including the features of your television:

[1680] 1. Resolution

[1681] 2. Special features (4K, HDR, smart features, etc.)

[1682] 3. Size

[1683] 4. Price range

[1684] 5. Buyer Reviews

[1685] Proposed output:

[1686] This TV is 4K compatible, allowing you to enjoy high-resolution images. It also features HDR functionality, providing vibrant colors. It's a 55-inch model and priced at approximately 100,000 yen. It has received high praise from many buyers.

[1687] In this way, customers can receive detailed product information through the robot. The entire system is managed via the cloud or a local server, enabling efficient and effective information delivery.

[1688] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1689] Step 1:

[1690] Process name: Voice input capture

[1691] Subject: User

[1692] Specific actions:

[1693] The user asks a question to the robot. An example message might be, "Please tell me about the features of this television." A high-sensitivity microphone built into the robot captures this audio.

[1694] Input: User's voice

[1695] Output: Audio data

[1696] Step 2:

[1697] Process name: Text conversion of audio data

[1698] Subject: terminal

[1699] Specific actions:

[1700] The captured audio data is converted into text data in real time by the device's speech recognition software (Google Speech-to-Text API).

[1701] Input: Audio data

[1702] Output: Text data

[1703] Step 3:

[1704] Process name: Sending text data

[1705] Subject: terminal

[1706] Specific actions:

[1707] Text data is sent from the terminal to the server. A secure communication protocol is used for this transmission.

[1708] Input: Text data

[1709] Output: Text data sent to the server

[1710] Step 4:

[1711] Process name: Text data analysis

[1712] Subject: Server

[1713] Specific actions:

[1714] The server analyzes the received text data using a natural language processing (NLP) engine (SpaCy or Google Natural Language API) to understand the user's intent and requests. Specifically, if the question is related to a product, it identifies information such as the product's features, benefits, price, and ratings.

[1715] Input: Text data sent to the server

[1716] Output: Analyzed text data (information such as product features and benefits)

[1717] Step 5:

[1718] Process name: Content generation

[1719] Subject: Server

[1720] Specific actions:

[1721] The server uses a generation AI model (GPT-4, DALL-E, TTS engine) to generate visual, auditory, and video content based on the analyzed text data. Examples include high-resolution images of the product, graphs and explanatory audio demonstrating product features, and demo videos.

[1722] Input: Analyzed text data

[1723] Output: Generated visual, auditory, and video content

[1724] Step 6:

[1725] Process name: Send content

[1726] Subject: Server

[1727] Specific actions:

[1728] The generated visual, auditory, and video content is transmitted from the server to the terminal. Protocols are used to maintain data integrity and security during this process.

[1729] Input: Generated visual, auditory, and video content

[1730] Output: Content sent to the device

[1731] Step 7:

[1732] Process name: Display and play content

[1733] Subject: Terminal (robot body)

[1734] Specific actions:

[1735] The device displays and plays received content. Images and graphs are shown on a high-resolution display, and explanatory audio is played using high-quality speakers. Demo videos are also played on the display.

[1736] For example, a voice explanation such as, "This TV is 4K compatible, allowing you to enjoy high-resolution images," is played while product images and demo videos are displayed.

[1737] Input: Content sent to the device

[1738] Output: Visual and auditory content displayed and played for the user.

[1739] This series of steps allows customers to interactively obtain product information.

[1740] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1741] This invention provides a system that effectively communicates the features and benefits of a product by allowing a user to input voice commands through a robot, which then generates and presents various forms of data (images, audio, video, etc.) based on those commands. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it enables interactive responses that correspond to the user's emotional state.

[1742] System Configuration

[1743] Hardware configuration

[1744] This system consists of the following main components:

[1745] Terminal (robot body): Equipped with voice input means, display, speaker, and camera, it serves as the main component for interactive demonstrations.

[1746] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation.

[1747] Software Configuration

[1748] The following software and algorithms operate between the terminal and the server:

[1749] Speech recognition software: Converts voice input from the user into text data.

[1750] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent.

[1751] Generative AI model: Generates relevant content such as visuals, audio, and video based on analysis results.

[1752] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state.

[1753] Communication protocol: Efficiently sends and receives data between a terminal and a server.

[1754] System operation

[1755] User voice input and emotion recognition

[1756] 1. The user asks a question about product information (e.g., "What are the features of this smartphone?").

[1757] 2. The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[1758] 3. The device uses its camera to capture the user's facial expressions, and the emotion engine analyzes this to recognize the emotional state (e.g., excitement, interest, questioning, etc.).

[1759] 4. The device sends text data and sentiment analysis results to the server.

[1760] Data analysis and content generation

[1761] 5. The server receives the text data, and the NLP engine performs analysis. Based on the analysis results, it identifies the information the user is looking for (e.g., smartphone camera functions, battery life, etc.).

[1762] 6. The server takes the sentiment analysis results into account and uses a generative AI model to generate relevant visual, audio, and video content.

[1763] Examples: Photos taken with a smartphone's high-resolution camera, a graph showing battery life, and a demo video of the facial recognition function.

[1764] Content display and response

[1765] 7. The device receives the generated content.

[1766] 8. Display visual content on the device's screen and play audio content using the speaker.

[1767] Example: An image taken with a high-resolution camera is displayed on the screen, and a voice message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[1768] 9. The user visually and audibly confirms the content that has been displayed and played.

[1769] 10. The emotion engine continuously monitors user responses and updates the emotion state to the server as needed.

[1770] Providing additional information and continuing the interaction

[1771] 11. If the user requests more detailed information, ask additional questions (e.g., "How long does the battery last?").

[1772] 12. The device captures the user's voice again, converts it to text data, and sends it to the server.

[1773] 13. The server analyzes the new text data and the latest sentiment analysis results, generates the necessary new content, and sends it to the terminal.

[1774] End of interaction

[1775] 14. When the user instructs the robot to end the interaction, the terminal captures the voice command to end the interaction, and the system enters standby mode.

[1776] In this way, the system provides detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state. This is expected to further pique the user's interest and increase their desire to purchase.

[1777] The following describes the processing flow.

[1778] Step 1:

[1779] The user speaks to the robot and asks, "Please tell me about the features of this smartphone."

[1780] Step 2:

[1781] The device's microphone captures the user's voice.

[1782] Step 3:

[1783] The device uses speech recognition software to convert the captured audio into text data.

[1784] Step 4:

[1785] The device's camera captures the user's facial expressions.

[1786] Step 5:

[1787] The device uses an emotion engine to analyze captured facial data and recognize emotional states (e.g., excitement, interest, questioning).

[1788] Step 6:

[1789] The terminal sends the converted text data and sentiment analysis results to the server.

[1790] Step 7:

[1791] The server receives text data and performs analysis using an NLP engine. Based on the analysis results, it identifies the information the user is requesting (e.g., smartphone camera functions, battery life, etc.).

[1792] Step 8:

[1793] The server uses a generative AI model, taking sentiment analysis results into account, to generate relevant visual, audio, and video content.

[1794] Step 9:

[1795] The server sends the generated content to the terminal.

[1796] Step 10:

[1797] The device receives the generated content.

[1798] Step 11:

[1799] Visual content is displayed on the device's screen, and audio content is played through the speaker. For example, a photo taken with a high-resolution camera is displayed on the screen, and an audio message is played saying, "This smartphone is equipped with a high-resolution camera and can take very clear photos."

[1800] Step 12:

[1801] Users visually and aurally confirm the displayed content and played audio.

[1802] Step 13:

[1803] The emotion engine continuously monitors the user's reactions and determines their emotional state.

[1804] Step 14:

[1805] If the user requests more detailed information (for example, "How long does the battery last?"), additional voice instructions will be provided.

[1806] Step 15:

[1807] The device captures the user's voice again, converts it into text data using speech recognition software, and sends it to the server.

[1808] Step 16:

[1809] The server receives new text data and the latest sentiment analysis results, and analyzes them again using the NLP engine. Based on the analysis results, it generates the necessary new content and sends it to the terminal.

[1810] Step 17:

[1811] The device receives new content, displays it on the screen, and plays it through the speaker. For example, a graph showing battery life might be displayed, and a voice message might say, "This smartphone's battery lasts 24 hours with normal use."

[1812] Step 18:

[1813] If the user is satisfied, they instruct the robot to end the interaction.

[1814] Step 19:

[1815] The terminal captures the voice command to terminate, and the system enters standby mode.

[1816] (Example 2)

[1817] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1818] Conventional systems, when providing information based on user voice input, lacked the ability to respond in a way that took into account the user's emotional state, resulting in a uniform user experience and low satisfaction. Furthermore, they lacked the ability to integrate and deliver diverse content formats, limiting the effectiveness of information transmission.

[1819] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1820] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for capturing the user's facial expressions and analyzing their emotional state, means for transmitting the text data and emotional state, means for analyzing the text data and generating related content such as visuals, audio, and video, means for displaying and playing the generated content, and means for presenting the generated content to the user visually and aurally. This enables the provision of interactive content that reflects the user's emotional state, thereby improving the quality of the user experience.

[1821] A "user" is an entity that gives instructions or asks questions to a system using voice.

[1822] A "terminal" is a device that performs voice input, voice recognition, facial expression capture, result display, and playback.

[1823] A "server" is a device that receives text data and sentiment analysis results, and performs data analysis and content generation.

[1824] "Voice input means" refers to devices such as microphones that capture the user's voice.

[1825] "Speech recognition software" is a program that converts captured audio into text data.

[1826] "Text data" refers to character information converted by speech recognition software.

[1827] "Facial expression capture means" refers to devices such as cameras that capture the user's facial expressions.

[1828] An "emotion engine" is software that analyzes a user's emotional state from captured facial expressions.

[1829] A "communication protocol" is a set of communication rules for efficiently sending and receiving data between a terminal and a server.

[1830] A "natural language processing (NLP) engine" is a program that analyzes text data to understand the user's intent.

[1831] A "generative AI model" is an algorithm that generates relevant visual, audio, and video content based on data analysis results.

[1832] "Visual content" refers to visual information such as images and graphics.

[1833] "Audio content" refers to auditory information such as spoken language.

[1834] "Video content" refers to information that combines dynamic visual and audio information.

[1835] "Content presentation means" refers to a function that provides generated visual, audio, and video content to the user visually and aurally.

[1836] The system of the present invention enhances the user experience by dynamically generating and presenting relevant visual, audio, and video content based on the user's voice input and emotion analysis. The embodiments for carrying out the present invention will be described in detail below.

[1837] Hardware configuration

[1838] This system consists of the following main hardware components:

[1839] Terminal: Equipped with voice input, display, speaker, and camera, it serves as the main component for interactive demonstrations. Specifically, it includes a high-sensitivity microphone, high-resolution display, speaker, and high-resolution camera.

[1840] Servers: Located in the cloud or locally, they are responsible for data analysis and content generation. Servers with high-performance computing resources are required.

[1841] Software Configuration

[1842] The following software and algorithms operate between the terminal and the server:

[1843] Speech recognition software: Converts voice input from a user into text data. For example, technologies such as Google Cloud Speech-to-Text are used for speech recognition.

[1844] Natural Language Processing (NLP) engine: Analyzes text data to understand user intent. OpenAI GPT-3 is an example of this.

[1845] Generative AI models: These models generate relevant content such as visuals, audio, and video based on analysis results. For example, OpenAI DALL-E and GPT-3 are used.

[1846] Emotion Engine: Analyzes the user's voice and facial expressions to recognize their emotional state. For example, the Microsoft Azure Emotion API is used.

[1847] Communication protocol: A protocol that efficiently sends and receives data between a terminal and a server. For example, HTTPS is used.

[1848] System operation

[1849] User voice input and emotion recognition

[1850] The user asks questions about product information using voice.

[1851] The device's microphone captures the user's voice, and speech recognition software converts this into text data.

[1852] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state.

[1853] The device sends text data and sentiment analysis results to the server.

[1854] Data analysis and content generation

[1855] The server analyzes the received text data using a natural language processing engine to identify the user's intent.

[1856] The server uses a generation AI model that takes sentiment analysis results into account to generate content in various formats.

[1857] For example, if a user asks, "What are the features of this smartwatch?", the system will generate a demo video of the heart rate monitoring function and a graph showing battery life.

[1858] Content display and response

[1859] The device receives the generated content, displays the visual content on its screen, and plays the audio content using its speaker.

[1860] For example, a demo video of the heart rate measurement function is displayed on the screen, and a voice guide plays saying, "This smartwatch is capable of accurate heart rate measurement."

[1861] Check the content that users have viewed and played.

[1862] The emotion engine monitors the user's reactions and updates the emotion state to the server as needed.

[1863] Examples of specific cases and prompt statements

[1864] As a concrete example, consider a scenario where a user asks, "Tell me about the features of this smartwatch." In this case, the user's voice is captured by the device's microphone and converted into text data by speech recognition software. The server then analyzes this text data with a natural language processing engine, and a generative AI model generates a demo video of the heart rate measurement function and a graph of battery life. If the emotion engine analyzes the user's facial expressions and recognizes an excited emotional state, it continues with a more detailed explanation of the functions.

[1865] An example of a prompt might be: "Consider a scenario where a user asks, 'What are the features of this smartwatch?' and come up with prompts that would allow the generative AI model to create appropriate visual, audio, and video content."

[1866] As described above, the system of the present invention can provide detailed product information in a step-by-step and interactive manner based on the user's voice input and emotional state, thereby significantly improving the quality of the user experience.

[1867] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1868] Step 1:

[1869] The user asks a question about product information using voice. For example, they might say, "Tell me about the features of this smartwatch."

[1870] Input: User voice input

[1871] Output: Captured audio data

[1872] Step 2:

[1873] The device's microphone captures the user's voice, and speech recognition software converts this into text data. For example, Google Cloud Speech-to-Text can be used.

[1874] Input: Captured audio data

[1875] Output: Converted text data (e.g., "Tell me the features of this smartwatch")

[1876] Step 3:

[1877] The device's camera captures the user's facial expressions, and an emotion engine analyzes this to recognize their emotional state. For example, the Microsoft Azure Emotion API can be used.

[1878] Input: Captured facial expression data

[1879] Output: Analyzed emotional state data (e.g., excitement, interest)

[1880] Step 4:

[1881] The device sends text data and sentiment analysis results to the server. For example, HTTPS is used to efficiently transmit the data.

[1882] Input: Text data and sentiment state data

[1883] Output: Sending data to the server

[1884] Step 5:

[1885] The server analyzes the received text data using a natural language processing (NLP) engine. For example, OpenAI GPT-3 can be used.

[1886] Input: Text data

[1887] Output: Analysis results of user intent (e.g., smartwatch heart rate monitoring function, battery life)

[1888] Step 6:

[1889] The server considers the sentiment analysis results and uses a generative AI model to generate relevant visual, audio, and video content. For example, it might use OpenAI DALL-E or GPT-3.

[1890] Input: Analysis results of user intent and emotional state data

[1891] Output: Generated visual, audio, and video content (e.g., a demo video of the heart rate measurement function, a graph showing battery life)

[1892] Step 7:

[1893] The device receives the generated content from the server.

[1894] Input: Content data sent from the server

[1895] Output: Content data stored on the device

[1896] Step 8:

[1897] The device displays visual content on its screen and plays audio content through its speaker. For example, it might show a demo video of the heart rate measurement function on the screen and play an audio guide through the speaker saying, "This smartwatch is capable of accurate heart rate measurement."

[1898] Input: Content data stored on the device

[1899] Output: Display of visual content and playback of audio content

[1900] Step 9:

[1901] Review the content displayed and played by the user. For example, review the display image and speaker description.

[1902] Input: Visual and audio content

[1903] Output: User understanding and response

[1904] Step 10:

[1905] The emotion engine continuously monitors the user's reactions and updates the emotional state to the server as needed. For example, if the user's facial expression changes to one of surprise, the emotion engine analyzes this and sends the information to the server.

[1906] Input: User's facial expression data

[1907] Output: Updated sentiment state data

[1908] Step 11:

[1909] The user asks additional questions by voice. For example, they might say, "How long does the battery last?"

[1910] Input: Additional voice questions

[1911] Output: Captured audio data

[1912] Step 12:

[1913] The device then captures the user's voice again and converts it into text data using speech recognition software.

[1914] Input: Captured audio data

[1915] Output: Converted text data

[1916] Step 13:

[1917] The server analyzes the new text data and the latest sentiment analysis results to generate the necessary new content.

[1918] Input: New text data and latest sentiment state data

[1919] Output: Newly generated visual, audio, and video content

[1920] Step 14:

[1921] The device receives newly generated content and delivers it to the user using its display and speaker. For example, it might display a graph showing battery life on the display and provide an explanation with voice guidance.

[1922] Input: Newly generated content data

[1923] Output: Display of visual content and playback of audio content

[1924] Step 15:

[1925] The user gives a voice command to end the interaction. For example, they might say, "I'm done."

[1926] Input: Voice command to end

[1927] Output: Captured audio data

[1928] Step 16:

[1929] The terminal captures the voice command to terminate the process, converts it into text data using speech recognition software, and then puts the system into standby mode.

[1930] Input: Voice command to end

[1931] Output: System transition to standby mode

[1932] (Application Example 2)

[1933] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1934] Conventional product information systems in physical stores have problems in providing quick and detailed responses to customer questions, and furthermore, in providing interactive responses that match the customer's emotions. In addition, there is a lack of means for customers to obtain specific product information visually and aurally, making it difficult to increase their purchasing intent. Therefore, the present invention aims to solve these problems and realize more effective and customized information provision to customers.

[1935] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1936] In this invention, the server includes means for inputting the user's voice, means for converting the voice into text data, means for analyzing the text data and generating related visual, auditory, and video content, means for displaying and playing the generated content, means for analyzing the user's facial expressions and recognizing their emotional state, and means for customizing the generated content according to the user's emotional state and presenting it visually and aurally. As a result, customers can obtain specific, visual, and auditory product information in real time within a physical store, and since that information is customized according to the customer's emotional state, an interactive experience that enhances their desire to purchase becomes possible.

[1937] "Means of inputting voice" refers to devices or software that receive voice signals from a user and process them as digital signals.

[1938] "Methods for converting speech to text data" refer to algorithms or software that analyze an input speech signal and convert it into corresponding text data.

[1939] "Means for analyzing text data and generating related visual, auditory, and video content" refers to a system that uses natural language processing and generative AI models based on text data to generate content in various formats.

[1940] "Means for displaying and playing generated content" refers to a system that outputs generated visual and auditory content to a display device or speaker so that the user can see and hear it.

[1941] "Methods for analyzing a user's facial expressions to recognize their emotional state" refer to software or algorithms that capture a user's facial expressions through cameras or sensors, analyze them, and estimate the user's emotional state.

[1942] "Means for customizing generated content according to the user's emotional state and presenting it visually and aurally" refers to a system that adjusts the format and content of the content considering the user's emotional state and provides information to the user in an appropriate manner.

[1943] One embodiment of the present invention relates to a system for effectively providing product information to customers in a physical store. This system uses voice input, speech recognition, natural language processing, sentiment recognition, and generative AI models to provide customized information in response to customer inquiries.

[1944] Configuration of the main components

[1945] Hardware configuration

[1946] Terminal: An interactive device equipped with voice input, a display, speakers, and a camera. It captures the user's voice and facial expressions and displays and plays the generated content.

[1947] Server: A central system located in the cloud or locally, which performs data analysis and content generation.

[1948] Software Configuration

[1949] The following software and algorithms operate between the terminal and the server:

[1950] Speech recognition software: Converts user speech into text data. Example: SpeechRecognition library.

[1951] Natural Language Processing Engine: Analyzes text data and generates the best possible answer to a question. Example: Transformers in Hugging Face.

[1952] Generative AI models: Generate relevant visual and auditory content. Example: A generative model for Hugging Face.

[1953] Emotion recognition engine: Analyzes the user's facial expressions to recognize their emotional state. Example: EmotionRecognizer.

[1954] Communication protocol: Enables efficient transmission and reception of data between terminals and servers.

[1955] Content generation process

[1956] 1. Voice Input: Users can ask questions about the product using voice. Example: "Please tell me about the features of this smartphone."

[1957] 2. Speech Recognition: The device's microphone captures speech and converts it to text using SpeechRecognition software.

[1958] 3. Text Analysis: The converted text data is sent to the server and analyzed by a natural language processing engine.

[1959] 4. Emotion Recognition: The device's camera captures the user's facial expressions, and the EmotionRecognizer recognizes their emotional state.

[1960] 5. Content Generation: Based on the analysis results and emotional state, the AI ​​generation model generates relevant content such as visuals, audio, and video.

[1961] 6. Content display and playback: The generated content is sent to the device, displayed on the screen, and the audio is played from the speaker.

[1962] Specific example

[1963] For example, if a user asks, "What are the features of this smartphone?", the system will operate as follows:

[1964] Speech recognition software converts the user's voice into text data.

[1965] A natural language processing engine analyzes text data to identify information about the smartphone's features (e.g., camera functions, battery life).

[1966] The generative AI model generates visual content such as high-resolution camera images and graphs showing battery life.

[1967] The emotion recognition engine analyzes the user's facial expressions, and if it determines that the user is in an excited state, it provides additional information such as a voice message saying, "You can take very clear photos."

[1968] Example of a prompt

[1969] "Please tell me about the features of this smartphone."

[1970] "How long does the battery last?"

[1971] This allows customers to receive specific and detailed information in real time, and since that information is customized according to the customer's emotional state, it enables an interactive experience that increases their desire to buy.

[1972] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1973] Step 1:

[1974] Users ask questions about the product using voice.

[1975] Input: User's voice.

[1976] Operation: The device's microphone captures the user's voice.

[1977] Output: Captured audio data.

[1978] Step 2:

[1979] The speech recognition software converts the captured audio data into text data.

[1980] Input: Captured audio data.

[1981] Operation: The device's speech recognition software (e.g., SpeechRecognition library) analyzes the speech data and converts it into corresponding text.

[1982] Output: Converted text data.

[1983] Step 3:

[1984] The terminal sends text data to the server.

[1985] Input: Converted text data.

[1986] Operation: Text data is sent to the server via a communication protocol.

[1987] Output: Text data received by the server.

[1988] Step 4:

[1989] The server's natural language processing engine analyzes the text data and generates the best possible answer to the question.

[1990] Input: Received text data.

[1991] Operation: The server's natural language processing engine (e.g., Hugging Face's Transformers) analyzes the text data, understands the user's intent in the question, and generates an appropriate answer.

[1992] Output: Generated answer text.

[1993] Step 5:

[1994] The device's camera captures the user's facial expressions, and the emotion recognition engine recognizes the user's emotional state.

[1995] Input: User's facial expression image.

[1996] Operation: The device's camera captures the user's facial expressions, and an emotion recognition engine (e.g., EmotionRecognizer) analyzes the facial data to estimate the emotional state.

[1997] Output: Estimated emotional state data.

[1998] Step 6:

[1999] The server uses a generation AI model to generate visual and auditory content based on response text and sentiment state data.

[2000] Input: Generated response text, estimated sentiment state data.

[2001] Operation: The server's generation AI model considers the response text and emotional state to generate relevant visual content (e.g., photos, graphs) and auditory content (e.g., voice guidance).

[2002] Output: Generated visual and auditory content.

[2003] Step 7:

[2004] The generated content is sent to the device, displayed on the screen, and played through the speaker.

[2005] Input: Generated visual and auditory content.

[2006] Operation: Content is sent from the server to the terminal, visual content is displayed on the screen, and audio content is played from the speaker.

[2007] Output: Users visually and aurally perceive the content.

[2008] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2009] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2010] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[2011] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2012] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[2013] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[2014] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[2015] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[2016] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[2017] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[2018] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[2019] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[2020] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[2021] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2022] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[2023] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[2024] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[2025] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[2026] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[2027] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[2028] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[2029] The following is further disclosed regarding the embodiments described above.

[2030] (Claim 1)

[2031] A means of inputting the user's voice,

[2032] A means for converting the audio into text data,

[2033] A means for analyzing the text data and generating related visual, audio, video, and other content,

[2034] Means for displaying and playing the generated content,

[2035] A system including means for presenting the generated content to a user visually and aurally.

[2036] (Claim 2)

[2037] In the system described in claim 1,

[2038] A system in which the voice input means includes means for recognizing a user's question or request.

[2039] (Claim 3)

[2040] In the system described in claim 1,

[2041] A system in which the content generation means includes means for integrating multiple data formats using artificial intelligence.

[2042] "Example 1"

[2043] (Claim 1)

[2044] A means of inputting the user's voice,

[2045] A means for converting the audio into text data,

[2046] A means for analyzing the text data and understanding the user's intent,

[2047] A means of generating relevant visual, audio, video, and other content using a generative AI model,

[2048] A means of sending the generated content to the terminal,

[2049] Means for presenting the generated content to the user visually and aurally,

[2050] A system including means for constructing a prompt statement corresponding to the generated content.

[2051] (Claim 2)

[2052] The system according to claim 1, wherein the voice input means includes means for recognizing a user's question or request.

[2053] (Claim 3)

[2054] The system according to claim 1, wherein the content generation means includes means for integrating multiple data formats using artificial intelligence.

[2055] "Application Example 1"

[2056] (Claim 1)

[2057] A means of inputting voice,

[2058] A means for converting the audio into text data,

[2059] A means for analyzing the text data and generating related visual, auditory, and video content,

[2060] Means for displaying and playing the generated content,

[2061] Means for presenting the generated content to the user visually and aurally,

[2062] A means of providing relevant content based on customer questions via a robot installed in a physical store,

[2063] A system including means for processing the aforementioned analysis and generated content via a cloud or local server.

[2064] (Claim 2)

[2065] The system according to claim 1, which recognizes a user's question or request.

[2066] (Claim 3)

[2067] The system according to claim 1, which uses a generative AI model to integrate multiple data formats.

[2068] "Example 2 of combining an emotion engine"

[2069] (Claim 1)

[2070] A means of inputting the user's voice,

[2071] A means for converting the audio into text data,

[2072] A means of capturing the user's facial expressions and analyzing their emotional state,

[2073] Means for transmitting text data and emotional states,

[2074] A means for analyzing the text data and generating related visual, audio, video, and other content,

[2075] Means for displaying and playing the generated content,

[2076] A system including means for presenting the generated content to a user visually and aurally.

[2077] (Claim 2)

[2078] The system according to claim 1 that generates content taking into account the aforementioned emotional state.

[2079] (Claim 3)

[2080] The system according to claim 1, which provides the generated content to the user visually and aurally.

[2081] "Application example 2 when combining with an emotional engine"

[2082] (Claim 1)

[2083] A means of inputting the user's voice,

[2084] A means for converting the audio into text data,

[2085] A means for analyzing the text data and generating related visual, auditory, video, and other content,

[2086] Means for displaying and playing the generated content,

[2087] A means of analyzing the user's facial expressions to recognize their emotional state,

[2088] A system including means for customizing the generated content according to the user's emotional state and presenting it visually and audibly.

[2089] (Claim 2)

[2090] The system according to claim 1, comprising means for recognizing a user's question or request.

[2091] (Claim 3)

[2092] The system according to claim 1, comprising means for integrating multiple data formats using artificial intelligence. [Explanation of Symbols]

[2093] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of inputting the user's voice, A means for converting the audio into text data, A means for analyzing the text data and generating related visual, audio, video, and other content, Means for displaying and playing the generated content, A system including means for presenting the generated content to a user visually and aurally.

2. In the system described in claim 1, A system in which the voice input means includes means for recognizing a user's question or request.

3. In the system described in claim 1, A system in which the content generation means includes means for integrating multiple data formats using artificial intelligence.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A