System
The system addresses the complexity of AI-based systems by using a stuffed toy-type terminal with voice recognition and synthesis, offering integrated services like information provision, healthcare support, and purchasing advice, enhancing user experience and accessibility.
Patent Information
- Application Number
- JP2024119124
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Existing AI-based systems require users to operate multiple applications and possess a certain level of knowledge, making them cumbersome and less user-friendly, limiting widespread adoption.
A system comprising a stuffed toy-type terminal with voice recognition and synthesis capabilities, integrated with a server managing user profiles, location information, and external services, providing a voice interface for information provision, healthcare support, and purchasing advice.
Enables intuitive operation and multifunctional support through a user-friendly interface, enhancing familiarity and comfort, suitable for a wide range of users.
Smart Images

Figure 2026018063000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, AI-based functions such as information provision, healthcare support, and purchasing advice are important for enriching users' daily lives. However, these functions are typically provided through two-dimensional interfaces such as smartphones and computers. This requires users to operate multiple applications, which not only makes operation cumbersome but also lacks familiarity and comfort. Furthermore, a certain level of knowledge is required to master advanced AI technology, which is a barrier to widespread adoption. Therefore, there is a need to provide AI interfaces that are simpler, more user-friendly, and yet multifunctional. [Means for solving the problem]
[0005] The present invention provides a system including: a server means for managing user profile information; a stuffed toy-type terminal means equipped with a microphone, a voice recognition device, and a voice synthesizer and providing a voice interface with the user; a communication means for communicating with the server means and transmitting the user's input profile information to the server means; a location information device for acquiring the user's location information; a weather information acquisition means for acquiring weather information from an external weather information provider based on the acquired location information; and a voice output means for synthesizing the acquired weather information into voice and providing it to the user. This allows the user to intuitively operate the device through the voice interface and to use functions such as information provision, healthcare support, and purchasing advice in an integrated manner. Furthermore, the use of a stuffed toy-type terminal means provides elements of familiarity and comfort, and is expected to be popular among a wide range of users.
[0006] The "server means" is a computer device that manages profile information about users and communicates with other devices and external services.
[0007] The "terminal means" is a device for providing a voice interface with the user, and in this invention is a device that particularly employs a stuffed toy design.
[0008] A "microphone" is a device that picks up a user's voice and converts it into an audio signal.
[0009] A "voice recognition device" is a device that has the function of analyzing picked-up voice signals and converting them into text data.
[0010] A "speech synthesizer" is a device that converts text data into a voice signal and outputs the voice through a speaker.
[0011] "Communication means" refers to the functions and technologies that enable data communication between terminal means and server means or other devices.
[0012] A "location information device" is a device for identifying a user's current location and acquiring that information.
[0013] The "weather information providing means" is a means for accessing an external weather information service and acquiring weather data.
[0014] A "healthcare device" is a device for acquiring health information such as data on the amount of exercise of a user.
[0015] "External weather information sources" refers to external providers or APIs that provide weather information services.
[0016] The "online store communication means" is a communication means for issuing an order to the online store based on the user's purchase instructions. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention is a system that manages profile information about users and provides a variety of information and support, and is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, an audio output means, and a healthcare device.
[0039] User profile registration
[0040] (overview)
[0041] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[0042] (Processing flow)
[0043] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[0044] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[0045] 3. The device uses voice recognition to convert the input information into text data.
[0046] 4. The device sends the converted text data to the server.
[0047] 5. The server stores the received profile information in a database.
[0048] Weather information provided
[0049] (overview)
[0050] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[0051] (Processing flow)
[0052] 1. The user asks the device, "What's the weather like today?"
[0053] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[0054] 3. The server obtains the user's location information and sends a request to the weather information API.
[0055] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0056] 5. The server returns the generated text data to the terminal.
[0057] 6. The device synthesizes the received text data into voice and tells the user, for example, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0058] Providing healthcare advice
[0059] (overview)
[0060] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[0061] (Processing flow)
[0062] 1. The user instructs the device, "Tell me how much exercise I did today."
[0063] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[0064] 3. The server obtains data from the healthcare device through the API.
[0065] 4. The server analyzes the acquired exercise data and generates advice for the user.
[0066] 5. The server returns the generated advice text to the terminal.
[0067] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[0068] Product proposals and orders
[0069] (overview)
[0070] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[0071] (Processing flow)
[0072] 1. The user asks the terminal, "How much shampoo is in stock?"
[0073] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[0074] 3. The server checks the user's purchasing history in the database and searches for inventory information.
[0075] 4. The server generates text data suggesting replenishment if stock is low or uncertain.
[0076] 5. The server returns the generated text data to the terminal.
[0077] 6. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[0078] 7. The user instructs the device to "purchase."
[0079] 8. The device uses voice recognition to convert the instructions into text data and send it to the server.
[0080] 9. The server receives the user's instructions and issues an order to the affiliated online store via API.
[0081] 10. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0082] 11. The terminal will report to the user by voice, "Your shampoo order has been completed."
[0083] In this way, the system of the present invention enriches the user's life through a series of operations and provides multifunctional support with a user-friendly interface.
[0084] The processing flow will be explained below.
[0085] User profile registration
[0086] Step 1:
[0087] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[0088] Step 2:
[0089] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[0090] Step 3:
[0091] The device uses voice recognition to convert the input information into text data.
[0092] Step 4:
[0093] The terminal transmits the converted text data to the server.
[0094] Step 5:
[0095] The server stores the received profile information in a database.
[0096] Weather information provided
[0097] Step 1:
[0098] The user asks the device, "What's the weather like today?"
[0099] Step 2:
[0100] The device uses voice recognition to convert the question into text data.
[0101] Step 3:
[0102] The terminal transmits the text data to the server.
[0103] Step 4:
[0104] The server obtains the user's location information and sends a request to the weather information API.
[0105] Step 5:
[0106] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0107] Step 6:
[0108] The server returns the generated text data to the terminal.
[0109] Step 7:
[0110] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0111] Providing healthcare advice
[0112] Step 1:
[0113] The user instructs the terminal, "Tell me how much exercise I did today."
[0114] Step 2:
[0115] The device uses voice recognition to convert instructions into text data.
[0116] Step 3:
[0117] The terminal transmits the text data to the server.
[0118] Step 4:
[0119] The server obtains data from healthcare devices through an API.
[0120] Step 5:
[0121] The server analyzes the acquired exercise data and generates advice for the user.
[0122] Step 6:
[0123] The server returns the generated advice text to the terminal.
[0124] Step 7:
[0125] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[0126] Product proposals and orders
[0127] Step 1:
[0128] The user asks the terminal, "How much shampoo is in stock?"
[0129] Step 2:
[0130] The device uses voice recognition to convert questions into text data.
[0131] Step 3:
[0132] The terminal transmits the text data to the server.
[0133] Step 4:
[0134] The server checks the user's purchasing history in a database and searches for inventory information.
[0135] Step 5:
[0136] The server generates text data suggesting replenishment when stock is low or uncertain.
[0137] Step 6:
[0138] The server returns the generated text data to the terminal.
[0139] Step 7:
[0140] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[0141] Step 8:
[0142] The user instructs the terminal to "purchase."
[0143] Step 9:
[0144] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[0145] Step 10:
[0146] The server receives the user's instructions and issues an order to the affiliated online store via API.
[0147] Step 11:
[0148] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0149] Step 12:
[0150] The terminal will report to the user via voice, "Your shampoo order has been completed."
[0151] Example 1
[0152] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0153] Conventional voice interface systems have difficulty effectively managing user profile information, location information, physical activity data, and purchase history, and providing a variety of information and support. Furthermore, the complexity of the system when integrating multiple functions and improving the user experience have also been issues.
[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0155] In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with users; communication means for communicating with the server means and transmitting the profile information input by the users to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; and voice output means for synthesizing the acquired weather information and providing it to the user. This makes it possible to integrate and manage the user's profile information, location information, exercise data, and purchase history, and to provide information and support efficiently and effectively.
[0156] The "server means" is a device or system for managing profile information about users and processing and storing data in cooperation with various databases.
[0157] The "terminal means" is a device or system that includes a microphone, a voice recognition device, and a voice synthesis device, provides a voice interface with the user, and collects input data from the user.
[0158] "Communication means" refers to a device or system for transmitting and receiving data between terminal means and server means.
[0159] A "location information device" is a device for acquiring a user's current location, and mainly uses location information technology such as GPS.
[0160] The "weather information acquisition means" is a device or system for acquiring the latest weather information from an external weather information providing means based on the acquired user's location information.
[0161] The "audio output means" is a device or system that converts acquired weather information and various data into audio and provides it to the user.
[0162] A "healthcare device" is a device for acquiring a user's exercise data, and primarily includes fitness trackers and smartwatches.
[0163] The "online store communication means" is a device or system for issuing an order to the online store based on the user's purchase instructions.
[0164] MODE FOR CARRYING OUT THE INVENTION
[0165] The present invention is a system that manages a user's profile information, location information, exercise data, and purchase history, and provides a variety of information and support based on this data. The system is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and a healthcare device.
[0166] User profile registration
[0167] overview
[0168] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[0169] Detailed Description
[0170] The device enters initial setup mode and starts a voice prompt saying, "Nice to meet you! Please register your profile." The user uses the microphone to input information such as their name, age, gender, and lifestyle patterns by voice. The device uses a voice recognition engine (for example, Google Voice Recognition API) to convert the voice into text data. The converted text data is sent to the server using a communication method. The server stores the received data in a database.
[0171] Specific examples
[0172] For example, if a user speaks "My name is Yamada Taro and I'm 30 years old," the device converts it into text and the server stores it in a database. An example of a prompt is "Write a program that inputs a user's name and age and stores that information in a database."
[0173] Weather information provided
[0174] overview
[0175] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[0176] Detailed Description
[0177] When a user asks a device, "What's the weather like today?", the device uses its voice recognition function to convert the question into text data and sends it to the server. The server uses a location information device to obtain the user's current location and sends a request to a weather information providing API (for example, OpenWeatherMap API). The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand. The generated text data is sent back to the device and conveyed to the user using the voice synthesis function.
[0178] Specific examples
[0179] When the user asks, "What's the weather like today?", the device responds, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees." An example of a prompt is, "Write a program that obtains weather information based on the user's location and relays it to them in voice."
[0180] Providing healthcare advice
[0181] overview
[0182] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[0183] Detailed Description
[0184] When a user instructs the device to "tell me how much exercise I did today," the device uses voice recognition to convert the instruction into text data and send it to the server. The server acquires the exercise data via the healthcare device, analyzes it, and generates advice. The advice is then sent back to the device, converted into voice, and conveyed to the user.
[0185] Specific examples
[0186] When the user asks, "How much exercise did you do today?", the device responds, "You've done 5,000 steps today. You're halfway to your goal of 10,000 steps." An example of a prompt is, "Write a program that obtains a user's exercise data and provides advice based on that data."
[0187] Product proposals and orders
[0188] overview
[0189] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[0190] Detailed Description
[0191] When a user asks the terminal, "What shampoo is in stock?", the terminal uses voice recognition to convert the question into text data and sends it to the server. The server checks the purchase history database to find inventory information. If inventory is low or unclear, the server generates a replenishment suggestion and sends it back to the terminal. The terminal notifies the user of this by voice, receives the user's purchase instructions, and issues an order.
[0192] Specific examples
[0193] For example, if a user asks, "How much shampoo is in stock?" and the terminal responds, "Your shampoo is low in stock. Would you like to buy more?", and the user responds by saying, "Purchase," the terminal reports, "Your shampoo order is complete." An example of a prompt sentence is, "Write a program that provides inventory information based on the user's purchasing history and places an order for the product if necessary."
[0194] In this way, the system of the present invention provides multifunctional services through detailed and specific processes to support the user's life.
[0195] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0196] User profile registration
[0197] Step 1:
[0198] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[0199] Specific behavior: The device plays a message aloud, prompting the user to prepare to start typing.
[0200] Input: Command to start the initial setting mode on the terminal side
[0201] Output: Start of voice prompt
[0202] Step 2:
[0203] The user follows the instructions on the device and enters their information by voice, for example, saying, "My name is Taro Yamada and I'm 30 years old."
[0204] Specific operation: The user verbally transmits their profile information through a microphone.
[0205] Input: Voice input (user profile information)
[0206] Output: Audio data
[0207] Step 3:
[0208] The device uses a voice recognition engine to convert the input voice data into text data, using the Google Voice Recognition API.
[0209] Specific operation: Converts voice data into text in real time.
[0210] Input: Audio data
[0211] Output: Text data
[0212] Step 4:
[0213] The device sends the converted text data to the server via a communication method, using an AWS S3 bucket.
[0214] Specific operation: Text data is transferred to the server using a communication protocol.
[0215] Input: Text data
[0216] Output: Send data to the server
[0217] Step 5:
[0218] The server stores the received profile information in a database, using a MySQL database.
[0219] Specific behavior: Saves and validates data.
[0220] Input: Text data
[0221] Output: Save information to a database
[0222] Weather information provided
[0223] Step 1:
[0224] The user asks the device, "What's the weather like today?" and inputs the request by voice.
[0225] Specific operation: The user verbally instructs the device that he or she wants to know weather information.
[0226] Input: Voice input (weather information request)
[0227] Output: Audio data
[0228] Step 2:
[0229] The device uses its voice recognition function to convert the question into text data and send it to the server, using the Google Voice Recognition API.
[0230] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[0231] Input: Audio data
[0232] Output: Text data (weather information request)
[0233] Step 3:
[0234] The server uses the location device to obtain the user's current location and sends a request to the OpenWeatherMap API.
[0235] Specific operation: Obtains location information and sends a request to an external API.
[0236] Input: Text data, location information
[0237] Output: Weather information request
[0238] Step 4:
[0239] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0240] What it does: Parses weather information and formats it for the user.
[0241] Input: Weather information data
[0242] Output: Text data (analyzed weather information)
[0243] Step 5:
[0244] The server returns the generated text data to the terminal, which converts it into voice and conveys it to the user.
[0245] Specific operation: Text data is synthesized into speech and information is provided to the user.
[0246] Input: Text data (analyzed weather information)
[0247] Output: Audio output (weather information)
[0248] Providing healthcare advice
[0249] Step 1:
[0250] The user instructs the device, "Tell me how much exercise I did today." The request is input by voice.
[0251] Specific operation: The user gives voice instructions to the terminal that he / she wants to know the exercise amount data.
[0252] Input: Voice input (request for exercise data)
[0253] Output: Audio data
[0254] Step 2:
[0255] The device uses voice recognition to convert instructions into text data and send it to the server.
[0256] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[0257] Input: Audio data
[0258] Output: Text data (momentum data request)
[0259] Step 3:
[0260] The server obtains exercise data via the healthcare device through the FitBit API.
[0261] Specific behavior: Uses an external API to obtain momentum data.
[0262] Input: Request data
[0263] Output: Momentum data
[0264] Step 4:
[0265] The server analyzes the acquired exercise data and generates advice for the user.
[0266] Specific operation: Analyze the momentum data and generate advice as text data.
[0267] Input: Momentum data
[0268] Output: Text data (advice)
[0269] Step 5:
[0270] The server returns the generated text data of the advice to the terminal, which converts it into voice and conveys it to the user.
[0271] Specific operation: Text data is synthesized into speech and information is provided to the user.
[0272] Input: Text data (advice)
[0273] Output: Audio output (advice)
[0274] Product proposals and orders
[0275] Step 1:
[0276] The user asks the terminal, "What shampoo is in stock?" and inputs the request by voice.
[0277] Specific operation: The user verbally instructs the terminal that he / she wants to know inventory information.
[0278] Input: Voice input (request for stock information)
[0279] Output: Audio data
[0280] Step 2:
[0281] The device uses voice recognition to convert the question into text data and send it to the server.
[0282] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[0283] Input: Audio data
[0284] Output: Text data (request for inventory information)
[0285] Step 3:
[0286] The server checks the user's purchasing history in a database and searches for inventory information.
[0287] Specific operation: Retrieve purchase history from the database and check inventory information.
[0288] Input: Text data (request for inventory information)
[0289] Output: Inventory information
[0290] Step 4:
[0291] The server generates replenishment suggestions when inventory is low or uncertain.
[0292] Specific operation: Generate replenishment suggestions as text data based on inventory information.
[0293] Input: Inventory information
[0294] Output: Text data (supplement proposal)
[0295] Step 5:
[0296] The server returns the generated text data of the proposal to the terminal, which converts it into voice and conveys it to the user.
[0297] Specific operation: Text data is synthesized into speech and information is provided to the user.
[0298] Input: Text data (supplement proposal)
[0299] Output: Audio output (supplement suggestion)
[0300] Step 6:
[0301] The user instructs the terminal to "purchase," and the terminal converts the instruction into text using voice recognition and sends it to the server.
[0302] Specific operation: Converts voice input into text data and sends it to the server.
[0303] Input: Voice input (purchase instructions)
[0304] Output: Text data (purchase instructions)
[0305] Step 7:
[0306] The server receives the user's instructions and places an order with the affiliated online store.
[0307] What it does: Places an order using the online store's API.
[0308] Input: Text data (purchase instructions)
[0309] Output: Order data
[0310] Step 8:
[0311] The server receives confirmation that the order has been completed and sends a notification to the terminal, which then tells the user, "Your shampoo order has been completed."
[0312] Specific operation: The notification data is synthesized into voice and the information is provided to the user.
[0313] Input: Order completion notification
[0314] Output: Audio output (order completion notification)
[0315] In this way, by performing specific input, data processing, and output at each step, multifunctional services are realized for users.
[0316] (Application example 1)
[0317] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0318] In modern society, users lead busy lives and require timely recommendations for optimal products and services based on weather information and health status to facilitate smooth purchasing behavior. It is also necessary to improve the quality of life by managing this information in an integrated manner and providing appropriate advice to users. However, conventional systems have had difficulty integrating various data, such as user profiles, location information, weather information, purchase history, and health data, to provide consistent services to users.
[0319] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0320] In this invention, the server includes server means for managing profile information about users, terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting the profile information entered by the user to the server means, a location information device for acquiring user location information, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information and providing it to the user, proposal generation means for proposing optimal products and services based on the acquired weather information and user preference information, and health support means for proposing optimal health foods and lifestyle-related products based on the user's purchase history and health data. This allows users to obtain information necessary for their daily lives in a centralized manner and improve their quality of life.
[0321] The "server means" is a central computer device that manages user profile information, purchase history, health data, etc., and integrates various data to provide optimal services to users.
[0322] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and provides a voice interface with the user.
[0323] The "communication means" is an interface for the terminal means to send and receive data to and from the server means.
[0324] A "location information device" is a device for obtaining the current location of a user.
[0325] The "weather information acquisition means" is a means for acquiring the latest weather information from an external weather information providing service based on the user's location information.
[0326] The "audio output means" is a device that synthesizes acquired information into voice and provides it to the user in the form of voice.
[0327] The "proposal generating means" is a means for proposing optimal products and services based on the acquired weather information and user preference information.
[0328] "Health support tools" are tools for suggesting optimal health foods and lifestyle-related products based on a user's purchasing history and health data.
[0329] This invention is a system that provides various information and support based on user profile information. The system is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a suggestion generation means, and a health support means.
[0330] Hardware and software used
[0331] 1. Server means: A central computer device (e.g., a cloud server) that manages user profile information, purchase history, and health data.
[0332] 2. Terminal means: A device equipped with a microphone, a voice recognition device, and a voice synthesis device (examples include smartphones, smart glasses, and head-mounted displays).
[0333] 3. Communication means: An interface that transmits and receives data between terminal means and server means (example: Internet communication).
[0334] 4. Location information device: A device that obtains the user's current location (example: GPS).
[0335] 5. Weather information acquisition method: A method for acquiring the latest weather information from an external weather information service (e.g., weather API).
[0336] 6. Voice output means: A device that provides information to the user by voice synthesis (example: speaker).
[0337] 7. Proposal generation means: A means for proposing optimal products and services based on weather information and user preference information (specific example: proposal generation algorithm).
[0338] 8. Health support tools: Tools that suggest optimal health foods and lifestyle-related products based on a user's purchasing history and health data (example: health advice algorithm).
[0339] Data processing and calculation
[0340] Registration of profile information: The user enters his / her profile information by voice through the terminal means, and the voice recognition device converts the information into text data and transmits it to the server means.
[0341] Obtaining weather information: Based on the user's location information obtained by the location information device, the weather information providing means obtains the latest weather information, and the audio output means provides the information by audio.
[0342] Proposal generation: The proposal generator integrates weather information and user preference information to suggest the most suitable products and services to the user.
[0343] Health support: Health support tools analyze users' purchasing history and health data to recommend optimal health foods and lifestyle-related products.
[0344] Examples of concrete examples and prompts
[0345] Example 1: "When a user asks their device, 'What's the weather like today?' the system obtains their location, retrieves weather information from a weather API, and provides, via voice synthesis, the answer, 'It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees.'"
[0346] Example prompt: "Suggest restaurants that deliver to a given user based on their location and weather."
[0347] Example 2: "When a user asks the device, 'What do you recommend for lunch today?' the system will suggest the best restaurant and dish based on weather information and the user's preferences."
[0348] Example prompt: "Based on the user's healthcare data, please recommend a healthy meal for them."
[0349] In this way, the system integrates multiple data sets, provides users with the information and services they need in a unified manner, and improves their quality of life.
[0350] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0351] System processing flow
[0352] Step 1:
[0353] The user inputs voice into the terminal means. The user asks questions such as "What's the weather like today?" or "What do you recommend for lunch today?" This voice input is captured by the microphone in the terminal means.
[0354] Step 2:
[0355] The terminal means converts the user's voice input into text data using a voice recognition device, and the converted text data is sent to the server means for further processing, where the voice data is processed as text data.
[0356] Step 3:
[0357] The server analyzes the received text data. For example, it analyzes the text "Tell me the weather today" and recognizes that the user's intention is to obtain weather information. Based on the analysis result, it obtains the user's current location from the location information device.
[0358] Step 4:
[0359] The server means uses the acquired location information to send a request to the weather information API, which acquires the latest weather information from an external weather information service based on the location information.
[0360] Step 5:
[0361] The server means receives and analyzes the weather information returned from the weather information providing API. Based on the analysis results, the server means generates text data for a voice response. For example, the generated text data is "It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees."
[0362] Step 6:
[0363] The generated text data is transmitted from the server means to the terminal means. Specifically, the server means transfers the text data for speech synthesis to the terminal means.
[0364] Step 7:
[0365] The terminal means converts the transmitted text data into voice using a voice synthesizer. This provides the user with a specific weather forecast. The terminal means outputs a voice saying, "It's sunny today. The maximum temperature will be 25 degrees, and the minimum temperature will be 15 degrees."
[0366] Step 8:
[0367] If the user wants to ask more detailed information or another question, they can input the question into the terminal again by voice. For example, they can ask, "What do you recommend for lunch today?" In this case, the process starts again from step 1.
[0368] Detailed System Operation
[0369] 1. Capturing voice input: The user makes a voice input, which is captured by the microphone of the terminal means.
[0370] Input: User's voice
[0371] Output: Captured audio data
[0372] 2. Speech recognition: The captured voice data is analyzed by a voice recognizer and converted into text data.
[0373] Input: Audio data
[0374] Output: Text data
[0375] 3. Analysis of text data: The server receives the text data, analyzes it, and recognizes the user's intention.
[0376] Input: Text data
[0377] Output: User request information
[0378] 4. Obtaining location information: The server means obtains the user's current location from the location information device.
[0379] Input: User information
[0380] Output: Current location data
[0381] 5. Obtaining weather information: The server sends a request to the weather information API to obtain the latest weather information.
[0382] Input: Current location data
[0383] Output: Weather information
[0384] 6. Generation of voice response: The server means generates text data for a voice response based on the weather information.
[0385] Input: Weather information
[0386] Output: Text data
[0387] 7. Sending text data: The server means sends the generated text data to the terminal means.
[0388] Input: Text data
[0389] Output: Text data transferred to the terminal means
[0390] 8. Speech synthesis: The terminal means synthesizes text data into speech and outputs it as speech.
[0391] Input: Text data
[0392] Output: Audio output
[0393] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0394] The present invention is a system that manages user profile information and provides a variety of information and support, and is particularly capable of interacting with the user while taking their emotions into consideration. The system is primarily composed of a server, a stuffed toy-type terminal, communication means, a location information device, weather information acquisition means, audio output means, a healthcare device, and an emotion engine.
[0395] User profile registration
[0396] (overview)
[0397] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[0398] (Processing flow)
[0399] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[0400] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[0401] 3. The device uses voice recognition to convert the input information into text data.
[0402] 4. The device sends the converted text data to the server.
[0403] 5. The server stores the received profile information in a database.
[0404] Weather information provided
[0405] (overview)
[0406] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[0407] (Processing flow)
[0408] 1. The user asks the device, "What's the weather like today?"
[0409] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[0410] 3. The server obtains the user's location information and sends a request to the weather information API.
[0411] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0412] 5. The server returns the generated text data to the terminal.
[0413] 6. The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0414] Providing healthcare advice
[0415] (overview)
[0416] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[0417] (Processing flow)
[0418] 1. The user instructs the device, "Tell me how much exercise I did today."
[0419] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[0420] 3. The server obtains data from the healthcare device through the API.
[0421] 4. The server analyzes the acquired exercise data and generates advice for the user.
[0422] 5. The server returns the generated advice text to the terminal.
[0423] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[0424] Product proposals and orders
[0425] (overview)
[0426] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[0427] (Processing flow)
[0428] 1. The user asks the terminal, "How much shampoo is in stock?"
[0429] 2. The device uses voice recognition to convert the question into text data.
[0430] 3. The device sends the text data to the server.
[0431] 4. The server checks the user's purchasing history in the database and searches for inventory information.
[0432] 5. The server generates text data suggesting replenishment if stock is low or uncertain.
[0433] 6. The server returns the generated text data to the terminal.
[0434] 7. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[0435] 8. The user instructs the device to "purchase."
[0436] 9. The device uses voice recognition to convert the instructions into text data and send it to the server.
[0437] 10. The server receives the user's instructions and issues an order to the affiliated online store via API.
[0438] 11. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0439] 12. The terminal will report to the user by voice, "Your shampoo order has been completed."
[0440] Implementing the Emotion Engine
[0441] (overview)
[0442] The system uses an emotion engine to recognize the user's emotions and interact with them based on those emotions, allowing for more personalized responses.
[0443] (Processing flow)
[0444] 1. The device acquires emotional data from voice and facial expressions while interacting with the user.
[0445] 2. The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[0446] 3. The device sends the generated emotional state text data to the server.
[0447] 4. The server generates appropriate feedback and advice for the user based on the received emotional data.
[0448] 5. The server sends the generated feedback and advice text data back to the device.
[0449] 6. The device synthesizes feedback and advice into voice and conveys it to the user.
[0450] Example: If the user is feeling stressed, the device might suggest, "Would you like me to play some music to help you calm down?"
[0451] In this way, by incorporating an emotion engine, it becomes possible to respond according to the user's emotional state, resulting in a system that provides more friendly and personalized services.
[0452] The processing flow will be explained below.
[0453] User profile registration
[0454] Step 1:
[0455] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[0456] Step 2:
[0457] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[0458] Step 3:
[0459] The device uses voice recognition to convert the input information into text data.
[0460] Step 4:
[0461] The terminal transmits the converted text data to the server.
[0462] Step 5:
[0463] The server stores the received profile information in a database.
[0464] Weather information provided
[0465] Step 1:
[0466] The user asks the device, "What's the weather like today?"
[0467] Step 2:
[0468] The device uses voice recognition to convert the question into text data.
[0469] Step 3:
[0470] The terminal transmits the text data to the server.
[0471] Step 4:
[0472] The server obtains the user's location information and sends a request to the weather information API.
[0473] Step 5:
[0474] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0475] Step 6:
[0476] The server returns the generated text data to the terminal.
[0477] Step 7:
[0478] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0479] Providing healthcare advice
[0480] Step 1:
[0481] The user instructs the terminal, "Tell me how much exercise I did today."
[0482] Step 2:
[0483] The device uses voice recognition to convert instructions into text data.
[0484] Step 3:
[0485] The terminal transmits the text data to the server.
[0486] Step 4:
[0487] The server obtains data from healthcare devices through an API.
[0488] Step 5:
[0489] The server analyzes the acquired exercise data and generates advice for the user.
[0490] Step 6:
[0491] The server returns the generated advice text data to the terminal.
[0492] Step 7:
[0493] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[0494] Product proposals and orders
[0495] Step 1:
[0496] The user asks the terminal, "How much shampoo is in stock?"
[0497] Step 2:
[0498] The device uses voice recognition to convert questions into text data.
[0499] Step 3:
[0500] The terminal transmits the text data to the server.
[0501] Step 4:
[0502] The server checks the user's purchasing history in a database and searches for inventory information.
[0503] Step 5:
[0504] The server generates text data suggesting replenishment when stock is low or uncertain.
[0505] Step 6:
[0506] The server returns the generated text data to the terminal.
[0507] Step 7:
[0508] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[0509] Step 8:
[0510] The user instructs the terminal to "purchase."
[0511] Step 9:
[0512] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[0513] Step 10:
[0514] The server receives the user's instructions and issues an order to the affiliated online store via API.
[0515] Step 11:
[0516] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0517] Step 12:
[0518] The terminal will report to the user via voice, "Your shampoo order has been completed."
[0519] Implementing the Emotion Engine
[0520] Step 1:
[0521] During a conversation with the user, the device uses a microphone and camera to acquire emotional data from voice and facial expressions.
[0522] Step 2:
[0523] The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[0524] Step 3:
[0525] The terminal transmits the generated text data of the emotional state to the server.
[0526] Step 4:
[0527] The server generates appropriate feedback and advice for the user based on the received emotional data.
[0528] Step 5:
[0529] The server returns the generated feedback and advice text data to the terminal.
[0530] Step 6:
[0531] The device synthesizes feedback and advice into voice and conveys it to the user.
[0532] Example: If the user is feeling stressed, the device might suggest, "Would you like me to play some music to help you calm down?"
[0533] In this way, by incorporating an emotion engine, the system is able to respond according to the user's emotional state, providing a more friendly and personalized service.
[0534] Example 2
[0535] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0536] In order to meet today's diverse and changing user needs, systems that can properly manage user profile information and provide personalized services based on that information are necessary. However, conventional systems have had difficulty comprehensively managing users' emotions, health status, and product purchase history, and providing feedback and advice in real time. To solve these problems, systems with more advanced recognition and response capabilities are required.
[0537] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0538] In this invention, the server includes means for managing profile information about users, stuffed toy-type terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting profile information entered by the user to the server means, a location information device for acquiring location information of the user, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information into voice and providing it to the user, emotion recognition means having an emotion recognition device and acquiring emotion data from the user's voice and facial expressions and transmitting it to the server, and server means for analyzing the emotion data and generating appropriate feedback and advice. This enables personalized responses according to the user's emotional state and health condition, as well as real-time feedback and advice.
[0539] The "server means" is a computer processing device that manages user profile information, purchase history, emotional data, etc., and has data analysis and feedback generation functions.
[0540] The "stuffed toy type terminal means" is a stuffed toy type device that incorporates a microphone, a voice recognition device, and a voice synthesis device and provides a voice interface with the user.
[0541] "Communication means" refers to a means for transmitting and receiving data between a server and a terminal means, and includes wireless communication technologies such as Wi-Fi, 4G, and 5G.
[0542] A "location information device" is a device that obtains a user's current location and uses GPS or other location identification technology.
[0543] The "weather information acquisition means" is a means for using weather information acquired from an external weather information providing service.
[0544] The "audio output means" is a means for converting acquired information into audio and providing it to the user, and utilizes voice synthesis technology.
[0545] The "emotion recognition means" is a means for acquiring and analyzing emotion data from the user's voice and facial expressions.
[0546] A "healthcare device" is a device for acquiring data on the amount of exercise performed by a user and has a measurement function for health management.
[0547] "Online Store Communication Means" means a means for communicating with the e-commerce platform and placing an order.
[0548] The present invention is a system that manages user profile information and provides a variety of information and support. In particular, the present invention provides a system that enables interaction that takes into account the user's emotions. The system mainly consists of a server means, a stuffed toy-type terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a health care device, and an emotion recognition device.
[0549] The overall operation of the system proceeds as follows: First, the user registers profile information using the stuffed toy-type terminal means. The terminal means uses a built-in microphone and voice recognition device to convert the user's voice input into text data and transmits that data to the server means. The server means stores the received data in a local database. Through this procedure, the user's profile information is managed.
[0550] Next, when the user asks, "What's the weather like today?", the terminal means uses a voice recognition device to convert the question into text data and sends it to the server means. The server means acquires the user's location information and sends a request to a weather information service (e.g., OpenWeatherMap API) to acquire the latest weather information. The acquired weather information is analyzed, converted into a format that is easy for the user to understand, and sent to the terminal means. Finally, the terminal means uses a voice synthesis device to convey the text data to the user, thereby providing the weather information.
[0551] The system also works with healthcare devices (e.g., Fitbit or Apple Watch) to acquire data on the user's exercise volume. When the user asks, "Tell me how much exercise I did today," the terminal device sends the instruction to the server device, which then acquires the exercise volume data from the healthcare device. Based on the data, the system generates health advice and communicates it to the user via the terminal device.
[0552] Furthermore, it also has a function to provide product inventory information based on the user's purchase history. When a user asks, "How much shampoo is in stock?", the server means checks the purchase history in the database and searches for inventory information. If the stock is low, it suggests replenishing the product, and when the user gives the instruction to "purchase," it places an order with the affiliated e-commerce platform and notifies the terminal means of confirmation of the order completion.
[0553] Finally, an emotion recognition device is used to analyze the user's emotional state and provide appropriate feedback in real time. If the user feels stressed, the system will suggest, "Would you like me to play some music to help you relax?" This allows for more personalized responses.
[0554] (Example of a prompt)
[0555] "Please explain the process for registering profile information. The information to enter is name, age, and gender."
[0556] "Please explain the process for providing weather information based on your current location."
[0557] "Please explain the process for capturing exercise data and providing health advice to users."
[0558] "Please explain the process of suggesting products based on a user's purchase history and placing an order."
[0559] "Please explain your process for recognizing user emotions and providing appropriate feedback."
[0560] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0561] User profile registration
[0562] Step 1:
[0563] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[0564] Input: Start the system in initial setting mode
[0565] Output: Voice message "Nice to meet you! Please register your profile."
[0566] Specific operation: The device uses a voice synthesizer to call out to the user in a familiar voice.
[0567] Step 2:
[0568] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[0569] Input: Voice data such as name, age, gender, lifestyle patterns, etc.
[0570] Output: User's voice input
[0571] Specific operation: The user speaks, "My name is Yamada Taro. I am 30 years old. I am male."
[0572] Step 3:
[0573] The device uses Google Cloud Speech-to-Text to convert input voice information into text data.
[0574] Input: User voice input
[0575] Output: User profile information converted to text data
[0576] Specific operation: The device analyzes the voice data, converts it into text, and temporarily stores it in its internal memory.
[0577] Step 4:
[0578] The terminal transmits the converted text data to the server.
[0579] Input: User profile information converted to text data
[0580] Output: Send data to the server
[0581] Specific operation: The device uses Wi-Fi or 4G / 5G communication to send text data to the server.
[0582] Step 5:
[0583] The server stores the received profile information in a MySQL database.
[0584] Input: User profile information converted to text data
[0585] Output: Message that data has been saved to the server
[0586] Specific operation: The server analyzes the text data sent and stores it in the corresponding field in the database.
[0587] Weather information provided
[0588] Step 1:
[0589] The user asks the device, "What's the weather like today?"
[0590] Input: Say "What's the weather like today?"
[0591] Output: User's voice input
[0592] Specific operation: The user asks a question into the device, and the device's microphone picks up the voice.
[0593] Step 2:
[0594] The device uses voice recognition to convert the question into text data and send it to the server.
[0595] Input: Say "What's the weather like today?"
[0596] Output: Question content converted into text data
[0597] Specific operation: The device converts the voice into text data and sends it to the server using Wi-Fi or 4G / 5G communication.
[0598] Step 3:
[0599] The server obtains the user's location information and sends a request to a weather information API (e.g., OpenWeatherMap).
[0600] Input: Question content converted into text data
[0601] Output: Request sent to the weather API
[0602] Specific operation: The server retrieves the user's location information from a database or GPS information, and uses that location information to send a request to the weather information API.
[0603] Step 4:
[0604] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0605] Input: Weather information returned from the API (JSON format)
[0606] Output: Weather information text formatted for user
[0607] Specific operation: The server analyzes weather information and generates text data such as "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0608] Step 5:
[0609] The server returns the generated text data to the terminal.
[0610] Input: Weather information text formatted for the user
[0611] Output: Send text data to the terminal
[0612] Specific operation: The server sends the generated text data to the terminal.
[0613] Step 6:
[0614] The terminal synthesizes the received text data into voice and conveys it to the user.
[0615] Input: Weather information converted into text data
[0616] Output: Voiced weather information
[0617] Specific operation: The device uses Google Text-to-Speech to convert the text data into speech and tells the user, "Today is sunny. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0618] Providing healthcare advice
[0619] Step 1:
[0620] The user instructs the terminal, "Tell me how much exercise I did today."
[0621] Input: Say "How much exercise did I do today?"
[0622] Output: User's voice input
[0623] Specific operation: The user speaks to the device, "Tell me how much exercise I did today."
[0624] Step 2:
[0625] The device uses voice recognition to convert instructions into text data and send it to the server.
[0626] Input: Say "How much exercise did I do today?"
[0627] Output: User instructions converted to text data
[0628] Specific operation: The device converts the voice into text data and sends it to the server via Wi-Fi or 4G / 5G.
[0629] Step 3:
[0630] The server obtains data from healthcare devices through an API.
[0631] Input: User instructions converted into text data
[0632] Output: Sending a request to the API and retrieving exercise data
[0633] Specific operation: The server obtains the user's exercise data using the API of a healthcare device (e.g., Fitbit).
[0634] Step 4:
[0635] The server analyzes the acquired exercise data and generates advice for the user.
[0636] Input: Exercise data obtained from healthcare devices
[0637] Output: Text data generation of advice
[0638] Specific operation: The server analyzes the exercise data and generates advice text such as, "Today's exercise amount is 5,000 steps. You are halfway to your goal of 10,000 steps."
[0639] Step 5:
[0640] The server returns the generated advice text to the terminal.
[0641] Input: Text data of advice
[0642] Output: Send advice to terminal
[0643] Specific operation: The server sends the generated text data to the terminal.
[0644] Step 6:
[0645] The terminal converts the text data into speech and conveys it to the user.
[0646] Input: Text data of advice
[0647] Output: Spoken advice
[0648] Specific operation: The device uses Google Text-to-Speech to convert the text into audio and tells the user, "You've taken 5,000 steps today. You're halfway to your goal of 10,000 steps."
[0649] Product proposals and orders
[0650] Step 1:
[0651] The user asks the terminal, "How much shampoo is in stock?"
[0652] Input: Speak "What shampoo is in stock?"
[0653] Output: User's voice input
[0654] Specific operation: The user speaks to the terminal, "How much shampoo is in stock?"
[0655] Step 2:
[0656] The device uses voice recognition to convert the question into text data.
[0657] Input: Speak "What shampoo is in stock?"
[0658] Output: Question content converted into text data
[0659] Specific operation: The device converts the voice into text data.
[0660] Step 3:
[0661] The terminal transmits the text data to the server.
[0662] Input: Question content converted into text data
[0663] Output: Send data to the server
[0664] Specific operation: The device sends text data to the server using Wi-Fi or 4G / 5G.
[0665] Step 4:
[0666] The server checks the user's purchasing history in a database and searches for inventory information.
[0667] Input: Question content converted into text data
[0668] Output: Stock information search results
[0669] Specific operation: The server retrieves the user's purchase history from the database and searches for inventory information.
[0670] Step 5:
[0671] The server generates text data suggesting replenishment when stock is low or uncertain.
[0672] Input: Search results for inventory information
[0673] Output: Text data of inventory replenishment proposal
[0674] Specific operation: The server generates text data such as "Stock is low. Would you like to purchase?"
[0675] Step 6:
[0676] The server returns the generated text data to the terminal.
[0677] Input: Text data of inventory replenishment proposal
[0678] Output: Sending data to the terminal
[0679] Specific operation: The server sends the generated text data to the terminal.
[0680] Step 7:
[0681] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[0682] Input: Text data of inventory replenishment proposal
[0683] Output: Vocalized refill suggestions
[0684] Specific operation: The device uses Google Text-to-Speech to communicate with the user aloud.
[0685] Step 8:
[0686] The user instructs the terminal to "purchase."
[0687] Input: Say "Buy"
[0688] Output: User's voice input
[0689] Specific action: The user says "purchase."
[0690] Step 9:
[0691] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[0692] Input: Say "Buy"
[0693] Output: User instructions converted to text data
[0694] Specific operation: The device converts the user's voice into text data and sends it to the server.
[0695] Step 10:
[0696] The server receives the user's instructions and issues an order through the API of the affiliated online store.
[0697] Input: User instructions converted into text data
[0698] Output: Place order to online store
[0699] Specific operation: The server calls the online store's API and places an order.
[0700] Step 11:
[0701] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0702] Input: Order completion notification from online store
[0703] Output: Order completion notification to terminal
[0704] Specific operation: The server receives confirmation of the order completion and notifies the terminal.
[0705] Step 12:
[0706] The terminal will report to the user via voice, "Your shampoo order has been completed."
[0707] Input: Notification of order completion
[0708] Output: Vocalized report
[0709] Specific operation: The device uses Google Text-to-Speech to report the completion of the order via voice.
[0710] Implementing the Emotion Engine
[0711] Step 1:
[0712] The terminal acquires emotional data from voice and facial expressions during a conversation with the user.
[0713] Input: User's voice and facial expression data
[0714] Output: Emotion data
[0715] Specific operation: The device uses an emotion recognition device to collect the user's voice and facial expression data.
[0716] Step 2:
[0717] The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[0718] Input: User's voice and facial expression data
[0719] Output: Emotional state converted into text data
[0720] Specific operation: The device analyzes the emotional data and generates text data that indicates the emotional state, such as "The user is feeling stressed."
[0721] Step 3:
[0722] The terminal transmits the generated text data of the emotional state to the server.
[0723] Input: Emotional state expressed as text data
[0724] Output: Send data to the server
[0725] Specific operation: The emotion data generated by the device is sent to the server.
[0726] Step 4:
[0727] The server generates appropriate feedback and advice for the user based on the received emotional data.
[0728] Input: Emotional state expressed as text data
[0729] Output: Text data of feedback and advice
[0730] Specific behavior: The server uses the emotional data to generate feedback and advice such as "Shall I play some music to help you relax?"
[0731] Step 5:
[0732] The server returns the generated feedback and advice text data to the terminal.
[0733] Input: Text data for feedback and advice
[0734] Output: Sending data to the terminal
[0735] Specific operation: The server sends the generated feedback to the device.
[0736] Step 6:
[0737] The device synthesizes feedback and advice into voice and conveys it to the user.
[0738] Input: Text data for feedback and advice
[0739] Output: Spoken feedback and advice
[0740] What it does: Your device will use Google Text-to-Speech to suggest aloud, "Would you like me to play some music to help you relax?"
[0741] The above is a specific description of the processing of the system.
[0742] (Application example 2)
[0743] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0744] Conventional systems provide limited information and services to users, and lack personalized support that takes emotion recognition into account. Furthermore, providing real-time, emotion-aware support is difficult for customer service in brick-and-mortar stores. The present invention aims to solve these problems and realize a system that provides interactive services tailored to the user's emotional state.
[0745] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user; communication means for communicating with the server means and transmitting profile information entered by the user to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; voice output means for synthesizing the acquired weather information and providing it to the user; emotion recognition means for acquiring and analyzing user emotion data and generating appropriate feedback and advice; and voice output means for synthesizing the feedback and advice generated by the emotion recognition means and providing it to the user. This enables the provision of highly personalized information and service support based on the user's emotional state.
[0746] The "server means" is a central system that manages user profile information and emotional data, collects and analyzes external information, and generates appropriate feedback.
[0747] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and that provides a voice interface with the user.
[0748] The "communication means" is a means for transmitting and receiving data between the server means and the terminal means.
[0749] A "location information device" is a device for acquiring the current location of a user.
[0750] The "weather information acquisition means" is a means for acquiring weather information from an external weather information providing service based on the acquired location information.
[0751] The "voice output means" is a means for synthesizing acquired or generated information into voice and conveying it to the user.
[0752] The "emotion recognition means" is a means for acquiring and analyzing the user's emotional data to generate appropriate feedback and advice.
[0753] "Feedback" refers to responses or advice to the user that are generated based on the user's emotional data and other information.
[0754] "Personalization" refers to providing information and services that are customized according to a user's individual profile and emotional state.
[0755] The system for implementing this invention provides personalized information based on a user's profile information and emotion recognition. Specifically, it is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and an emotion recognition means.
[0756] First, the server manages user profile information and collects and analyzes external information. The server includes a database system for storing profile information and emotional data. Furthermore, the server analyzes data using Natural Language Processing (NLP) APIs and speech recognition APIs from Azure and Google Cloud Services. Based on the results of this analysis, the server generates appropriate feedback and advice for the user.
[0757] The terminal is a device that provides a voice interface with the user and is equipped with a microphone, a voice recognition device, and a voice synthesis device. When the user inputs a question or instruction by voice, it is converted into text data and sent to the server. The converted text data is analyzed by the server and appropriate feedback is generated. The generated feedback is sent back to the terminal and conveyed to the user via voice output means.
[0758] The communication means is used to transmit and receive data between the server and the terminal. This communication is performed in real time, and the information required by the user is provided promptly.
[0759] The location information device is used to acquire the user's current location, and based on the acquired location information, the weather information acquisition means acquires weather information from an external weather information service, allowing the user to know the weather information for their current location in real time.
[0760] The emotion recognition means acquires and analyzes the user's emotion data to generate appropriate feedback and advice. This emotion recognition is used to understand the user's emotional state from voice input and text data and provide the optimal response to the user. For example, if the user is feeling stressed, it is possible to suggest relaxing music.
[0761] The following are specific examples of embodiments:
[0762] When a user speaks into their smartphone, "What's today's promotion?", the speech is converted into text data and sent to the server. The server analyzes the text data and generates promotion information, taking into account the user's emotional state. The device then synthesizes the generated promotion information into voice and provides it to the user.
[0763] Examples of such prompts include "Determine whether the user is interested in sale days and provide appropriate promotional information" and "Determine the emotion of the following text:- Text: {user's voice input}- Emotion: happy, sad, angry, neutral."
[0764] This allows users to get the information they need in real time within a physical store and receive personalized support based on their emotional state.
[0765] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0766] Step 1:
[0767] The user speaks a question or command into the device. For example, they might say, "What are today's promotions?" The voice input is picked up by the device's microphone.
[0768] Step 2:
[0769] The terminal analyzes the voice input with a voice recognition device and converts it into text data. After this text data is generated, it is sent to the server. The input is voice data and the output is text data.
[0770] Step 3:
[0771] The server analyzes the received text data using a Natural Language Processing (NLP) API. During the analysis, it also extracts the question content and emotional part. This analysis reveals the user's question content and emotional state. The input is text data, and the output is the analysis result.
[0772] Step 4:
[0773] The server generates an appropriate response based on the analysis results and profile information. For example, if the question is about a promotion, it generates promotion information taking into account inventory information and cross-selling information. The response also changes depending on the emotional state. The input is the analysis results and profile data, and the output is the response data.
[0774] Step 5:
[0775] The server then personalizes the generated response data according to the user's emotional state and sends the response in text format to the terminal. The input is the response data, and the output is the personalized text data.
[0776] Step 6:
[0777] The terminal synthesizes the received personalized text data using a voice output means and conveys the speech to the user. For example, it might say, "Today is sale day! Your favorite product is 20% off!" The input is personalized text data, and the output is voice data.
[0778] Step 7:
[0779] After receiving a response from the device, the user acts based on the content. For example, a user receives information about a sale day and decides which specific product to purchase. This step does not provide any feedback to the system, but it may affect future interactions. The input is voice data, and the output is the user's actions.
[0780] This allows users to receive personalized information and responses in real time that are tailored to their emotions.
[0781] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0782] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0783] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0784] [Second embodiment]
[0785] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0786] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0787] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0788] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0789] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0790] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0791] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0792] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0793] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0794] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0795] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0796] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0797] The present invention is a system that manages profile information about users and provides a variety of information and support, and is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, an audio output means, and a healthcare device.
[0798] User profile registration
[0799] (overview)
[0800] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[0801] (Processing flow)
[0802] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[0803] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[0804] 3. The device uses voice recognition to convert the input information into text data.
[0805] 4. The device sends the converted text data to the server.
[0806] 5. The server stores the received profile information in a database.
[0807] Weather information provided
[0808] (overview)
[0809] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[0810] (Processing flow)
[0811] 1. The user asks the device, "What's the weather like today?"
[0812] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[0813] 3. The server obtains the user's location information and sends a request to the weather information API.
[0814] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0815] 5. The server returns the generated text data to the terminal.
[0816] 6. The device synthesizes the received text data into voice and tells the user, for example, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0817] Providing healthcare advice
[0818] (overview)
[0819] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[0820] (Processing flow)
[0821] 1. The user instructs the device, "Tell me how much exercise I did today."
[0822] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[0823] 3. The server obtains data from the healthcare device through the API.
[0824] 4. The server analyzes the acquired exercise data and generates advice for the user.
[0825] 5. The server returns the generated advice text to the terminal.
[0826] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[0827] Product proposals and orders
[0828] (overview)
[0829] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[0830] (Processing flow)
[0831] 1. The user asks the terminal, "How much shampoo is in stock?"
[0832] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[0833] 3. The server checks the user's purchasing history in the database and searches for inventory information.
[0834] 4. The server generates text data suggesting replenishment if stock is low or uncertain.
[0835] 5. The server returns the generated text data to the terminal.
[0836] 6. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[0837] 7. The user instructs the device to "purchase."
[0838] 8. The device uses voice recognition to convert the instructions into text data and send it to the server.
[0839] 9. The server receives the user's instructions and issues an order to the affiliated online store via API.
[0840] 10. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0841] 11. The terminal will report to the user by voice, "Your shampoo order has been completed."
[0842] In this way, the system of the present invention enriches the user's life through a series of operations and provides multifunctional support with a user-friendly interface.
[0843] The processing flow will be explained below.
[0844] User profile registration
[0845] Step 1:
[0846] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[0847] Step 2:
[0848] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[0849] Step 3:
[0850] The device uses voice recognition to convert the input information into text data.
[0851] Step 4:
[0852] The terminal transmits the converted text data to the server.
[0853] Step 5:
[0854] The server stores the received profile information in a database.
[0855] Weather information provided
[0856] Step 1:
[0857] The user asks the device, "What's the weather like today?"
[0858] Step 2:
[0859] The device uses voice recognition to convert the question into text data.
[0860] Step 3:
[0861] The terminal transmits the text data to the server.
[0862] Step 4:
[0863] The server obtains the user's location information and sends a request to the weather information API.
[0864] Step 5:
[0865] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0866] Step 6:
[0867] The server returns the generated text data to the terminal.
[0868] Step 7:
[0869] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[0870] Providing healthcare advice
[0871] Step 1:
[0872] The user instructs the terminal, "Tell me how much exercise I did today."
[0873] Step 2:
[0874] The device uses voice recognition to convert instructions into text data.
[0875] Step 3:
[0876] The terminal transmits the text data to the server.
[0877] Step 4:
[0878] The server obtains data from healthcare devices through an API.
[0879] Step 5:
[0880] The server analyzes the acquired exercise data and generates advice for the user.
[0881] Step 6:
[0882] The server returns the generated advice text to the terminal.
[0883] Step 7:
[0884] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[0885] Product proposals and orders
[0886] Step 1:
[0887] The user asks the terminal, "How much shampoo is in stock?"
[0888] Step 2:
[0889] The device uses voice recognition to convert questions into text data.
[0890] Step 3:
[0891] The terminal transmits the text data to the server.
[0892] Step 4:
[0893] The server checks the user's purchasing history in a database and searches for inventory information.
[0894] Step 5:
[0895] The server generates text data suggesting replenishment when stock is low or uncertain.
[0896] Step 6:
[0897] The server returns the generated text data to the terminal.
[0898] Step 7:
[0899] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[0900] Step 8:
[0901] The user instructs the terminal to "purchase."
[0902] Step 9:
[0903] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[0904] Step 10:
[0905] The server receives the user's instructions and issues an order to the affiliated online store via API.
[0906] Step 11:
[0907] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[0908] Step 12:
[0909] The terminal will report to the user via voice, "Your shampoo order has been completed."
[0910] Example 1
[0911] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0912] Conventional voice interface systems have difficulty effectively managing user profile information, location information, physical activity data, and purchase history, and providing a variety of information and support. Furthermore, the complexity of the system when integrating multiple functions and improving the user experience have also been issues.
[0913] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0914] In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with users; communication means for communicating with the server means and transmitting the profile information input by the users to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; and voice output means for synthesizing the acquired weather information and providing it to the user. This makes it possible to integrate and manage the user's profile information, location information, exercise data, and purchase history, and to provide information and support efficiently and effectively.
[0915] The "server means" is a device or system for managing profile information about users and processing and storing data in cooperation with various databases.
[0916] The "terminal means" is a device or system that includes a microphone, a voice recognition device, and a voice synthesis device, provides a voice interface with the user, and collects input data from the user.
[0917] "Communication means" refers to a device or system for transmitting and receiving data between terminal means and server means.
[0918] A "location information device" is a device for acquiring a user's current location, and mainly uses location information technology such as GPS.
[0919] The "weather information acquisition means" is a device or system for acquiring the latest weather information from an external weather information providing means based on the acquired user's location information.
[0920] The "audio output means" is a device or system that converts acquired weather information and various data into audio and provides it to the user.
[0921] A "healthcare device" is a device for acquiring a user's exercise data, and primarily includes fitness trackers and smartwatches.
[0922] The "online store communication means" is a device or system for issuing an order to the online store based on the user's purchase instructions.
[0923] MODE FOR CARRYING OUT THE INVENTION
[0924] The present invention is a system that manages a user's profile information, location information, exercise data, and purchase history, and provides a variety of information and support based on this data. The system is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and a healthcare device.
[0925] User profile registration
[0926] overview
[0927] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[0928] Detailed Description
[0929] The device enters initial setup mode and starts a voice prompt saying, "Nice to meet you! Please register your profile." The user uses the microphone to input information such as their name, age, gender, and lifestyle patterns by voice. The device uses a voice recognition engine (for example, Google Voice Recognition API) to convert the voice into text data. The converted text data is sent to the server using a communication method. The server stores the received data in a database.
[0930] Specific examples
[0931] For example, if a user speaks "My name is Yamada Taro and I'm 30 years old," the device converts it into text and the server stores it in a database. An example of a prompt is "Write a program that inputs a user's name and age and stores that information in a database."
[0932] Weather information provided
[0933] overview
[0934] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[0935] Detailed Description
[0936] When a user asks a device, "What's the weather like today?", the device uses its voice recognition function to convert the question into text data and sends it to the server. The server uses a location information device to obtain the user's current location and sends a request to a weather information providing API (for example, OpenWeatherMap API). The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand. The generated text data is sent back to the device and conveyed to the user using the voice synthesis function.
[0937] Specific examples
[0938] When the user asks, "What's the weather like today?", the device responds, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees." An example of a prompt is, "Write a program that obtains weather information based on the user's location and relays it to them in voice."
[0939] Providing healthcare advice
[0940] overview
[0941] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[0942] Detailed Description
[0943] When a user instructs the device to "tell me how much exercise I did today," the device uses voice recognition to convert the instruction into text data and send it to the server. The server acquires the exercise data via the healthcare device, analyzes it, and generates advice. The advice is then sent back to the device, converted into voice, and conveyed to the user.
[0944] Specific examples
[0945] When the user asks, "How much exercise did you do today?", the device responds, "You've done 5,000 steps today. You're halfway to your goal of 10,000 steps." An example of a prompt is, "Write a program that obtains a user's exercise data and provides advice based on that data."
[0946] Product proposals and orders
[0947] overview
[0948] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[0949] Detailed Description
[0950] When a user asks the terminal, "What shampoo is in stock?", the terminal uses voice recognition to convert the question into text data and sends it to the server. The server checks the purchase history database to find inventory information. If inventory is low or unclear, the server generates a replenishment suggestion and sends it back to the terminal. The terminal notifies the user of this by voice, receives the user's purchase instructions, and issues an order.
[0951] Specific examples
[0952] For example, if a user asks, "How much shampoo is in stock?" and the terminal responds, "Your shampoo is low in stock. Would you like to buy more?", and the user responds by saying, "Purchase," the terminal reports, "Your shampoo order is complete." An example of a prompt sentence is, "Write a program that provides inventory information based on the user's purchasing history and places an order for the product if necessary."
[0953] In this way, the system of the present invention provides multifunctional services through detailed and specific processes to support the user's life.
[0954] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0955] User profile registration
[0956] Step 1:
[0957] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[0958] Specific behavior: The device plays a message aloud, prompting the user to prepare to start typing.
[0959] Input: Command to start the initial setting mode on the terminal side
[0960] Output: Start of voice prompt
[0961] Step 2:
[0962] The user follows the instructions on the device and enters their information by voice, for example, saying, "My name is Taro Yamada and I'm 30 years old."
[0963] Specific operation: The user verbally transmits their profile information through a microphone.
[0964] Input: Voice input (user profile information)
[0965] Output: Audio data
[0966] Step 3:
[0967] The device uses a voice recognition engine to convert the input voice data into text data, using the Google Voice Recognition API.
[0968] Specific operation: Converts voice data into text in real time.
[0969] Input: Audio data
[0970] Output: Text data
[0971] Step 4:
[0972] The device sends the converted text data to the server via a communication method, using an AWS S3 bucket.
[0973] Specific operation: Text data is transferred to the server using a communication protocol.
[0974] Input: Text data
[0975] Output: Send data to the server
[0976] Step 5:
[0977] The server stores the received profile information in a database, using a MySQL database.
[0978] Specific behavior: Saves and validates data.
[0979] Input: Text data
[0980] Output: Save information to a database
[0981] Weather information provided
[0982] Step 1:
[0983] The user asks the device, "What's the weather like today?" and inputs the request by voice.
[0984] Specific operation: The user verbally instructs the device that he or she wants to know weather information.
[0985] Input: Voice input (weather information request)
[0986] Output: Audio data
[0987] Step 2:
[0988] The device uses its voice recognition function to convert the question into text data and send it to the server, using the Google Voice Recognition API.
[0989] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[0990] Input: Audio data
[0991] Output: Text data (weather information request)
[0992] Step 3:
[0993] The server uses the location device to obtain the user's current location and sends a request to the OpenWeatherMap API.
[0994] Specific operation: Obtains location information and sends a request to an external API.
[0995] Input: Text data, location information
[0996] Output: Weather information request
[0997] Step 4:
[0998] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[0999] What it does: Parses weather information and formats it for the user.
[1000] Input: Weather information data
[1001] Output: Text data (analyzed weather information)
[1002] Step 5:
[1003] The server returns the generated text data to the terminal, which converts it into voice and conveys it to the user.
[1004] Specific operation: Text data is synthesized into speech and information is provided to the user.
[1005] Input: Text data (analyzed weather information)
[1006] Output: Audio output (weather information)
[1007] Providing healthcare advice
[1008] Step 1:
[1009] The user instructs the device, "Tell me how much exercise I did today." The request is input by voice.
[1010] Specific operation: The user gives voice instructions to the terminal that he / she wants to know the exercise amount data.
[1011] Input: Voice input (request for exercise data)
[1012] Output: Audio data
[1013] Step 2:
[1014] The device uses voice recognition to convert instructions into text data and send it to the server.
[1015] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[1016] Input: Audio data
[1017] Output: Text data (momentum data request)
[1018] Step 3:
[1019] The server obtains exercise data via the healthcare device through the FitBit API.
[1020] Specific behavior: Uses an external API to obtain momentum data.
[1021] Input: Request data
[1022] Output: Momentum data
[1023] Step 4:
[1024] The server analyzes the acquired exercise data and generates advice for the user.
[1025] Specific operation: Analyze the momentum data and generate advice as text data.
[1026] Input: Momentum data
[1027] Output: Text data (advice)
[1028] Step 5:
[1029] The server returns the generated text data of the advice to the terminal, which converts it into voice and conveys it to the user.
[1030] Specific operation: Text data is synthesized into speech and information is provided to the user.
[1031] Input: Text data (advice)
[1032] Output: Audio output (advice)
[1033] Product proposals and orders
[1034] Step 1:
[1035] The user asks the terminal, "What shampoo is in stock?" and inputs the request by voice.
[1036] Specific operation: The user verbally instructs the terminal that he / she wants to know inventory information.
[1037] Input: Voice input (request for stock information)
[1038] Output: Audio data
[1039] Step 2:
[1040] The device uses voice recognition to convert the question into text data and send it to the server.
[1041] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[1042] Input: Audio data
[1043] Output: Text data (request for inventory information)
[1044] Step 3:
[1045] The server checks the user's purchasing history in a database and searches for inventory information.
[1046] Specific operation: Retrieve purchase history from the database and check inventory information.
[1047] Input: Text data (request for inventory information)
[1048] Output: Inventory information
[1049] Step 4:
[1050] The server generates replenishment suggestions when inventory is low or uncertain.
[1051] Specific operation: Generate replenishment suggestions as text data based on inventory information.
[1052] Input: Inventory information
[1053] Output: Text data (supplement proposal)
[1054] Step 5:
[1055] The server returns the generated text data of the proposal to the terminal, which converts it into voice and conveys it to the user.
[1056] Specific operation: Text data is synthesized into speech and information is provided to the user.
[1057] Input: Text data (supplement proposal)
[1058] Output: Audio output (supplement suggestion)
[1059] Step 6:
[1060] The user instructs the terminal to "purchase," and the terminal converts the instruction into text using voice recognition and sends it to the server.
[1061] Specific operation: Converts voice input into text data and sends it to the server.
[1062] Input: Voice input (purchase instructions)
[1063] Output: Text data (purchase instructions)
[1064] Step 7:
[1065] The server receives the user's instructions and places an order with the affiliated online store.
[1066] What it does: Places an order using the online store's API.
[1067] Input: Text data (purchase instructions)
[1068] Output: Order data
[1069] Step 8:
[1070] The server receives confirmation that the order has been completed and sends a notification to the terminal, which then tells the user, "Your shampoo order has been completed."
[1071] Specific operation: The notification data is synthesized into voice and the information is provided to the user.
[1072] Input: Order completion notification
[1073] Output: Audio output (order completion notification)
[1074] In this way, by performing specific input, data processing, and output at each step, multifunctional services are realized for users.
[1075] (Application example 1)
[1076] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1077] In modern society, users lead busy lives and require timely recommendations for optimal products and services based on weather information and health status to facilitate smooth purchasing behavior. It is also necessary to improve the quality of life by managing this information in an integrated manner and providing appropriate advice to users. However, conventional systems have had difficulty integrating various data, such as user profiles, location information, weather information, purchase history, and health data, to provide consistent services to users.
[1078] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1079] In this invention, the server includes server means for managing profile information about users, terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting the profile information entered by the user to the server means, a location information device for acquiring user location information, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information and providing it to the user, proposal generation means for proposing optimal products and services based on the acquired weather information and user preference information, and health support means for proposing optimal health foods and lifestyle-related products based on the user's purchase history and health data. This allows users to obtain information necessary for their daily lives in a centralized manner and improve their quality of life.
[1080] The "server means" is a central computer device that manages user profile information, purchase history, health data, etc., and integrates various data to provide optimal services to users.
[1081] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and provides a voice interface with the user.
[1082] The "communication means" is an interface for the terminal means to send and receive data to and from the server means.
[1083] A "location information device" is a device for obtaining the current location of a user.
[1084] The "weather information acquisition means" is a means for acquiring the latest weather information from an external weather information providing service based on the user's location information.
[1085] The "audio output means" is a device that synthesizes acquired information into voice and provides it to the user in the form of voice.
[1086] The "proposal generating means" is a means for proposing optimal products and services based on the acquired weather information and user preference information.
[1087] "Health support tools" are tools for suggesting optimal health foods and lifestyle-related products based on a user's purchasing history and health data.
[1088] This invention is a system that provides various information and support based on user profile information. The system is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a suggestion generation means, and a health support means.
[1089] Hardware and software used
[1090] 1. Server means: A central computer device (e.g., a cloud server) that manages user profile information, purchase history, and health data.
[1091] 2. Terminal means: A device equipped with a microphone, a voice recognition device, and a voice synthesis device (examples include smartphones, smart glasses, and head-mounted displays).
[1092] 3. Communication means: An interface that transmits and receives data between terminal means and server means (example: Internet communication).
[1093] 4. Location information device: A device that obtains the user's current location (example: GPS).
[1094] 5. Weather information acquisition method: A method for acquiring the latest weather information from an external weather information service (e.g., weather API).
[1095] 6. Voice output means: A device that provides information to the user by voice synthesis (example: speaker).
[1096] 7. Proposal generation means: A means for proposing optimal products and services based on weather information and user preference information (specific example: proposal generation algorithm).
[1097] 8. Health support tools: Tools that suggest optimal health foods and lifestyle-related products based on a user's purchasing history and health data (example: health advice algorithm).
[1098] Data processing and calculation
[1099] Registration of profile information: The user enters his / her profile information by voice through the terminal means, and the voice recognition device converts the information into text data and transmits it to the server means.
[1100] Obtaining weather information: Based on the user's location information obtained by the location information device, the weather information providing means obtains the latest weather information, and the audio output means provides the information by audio.
[1101] Proposal generation: The proposal generator integrates weather information and user preference information to suggest the most suitable products and services to the user.
[1102] Health support: Health support tools analyze users' purchasing history and health data to recommend optimal health foods and lifestyle-related products.
[1103] Examples of concrete examples and prompts
[1104] Example 1: "When a user asks their device, 'What's the weather like today?' the system obtains their location, retrieves weather information from a weather API, and provides, via voice synthesis, the answer, 'It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees.'"
[1105] Example prompt: "Suggest restaurants that deliver to a given user based on their location and weather."
[1106] Example 2: "When a user asks the device, 'What do you recommend for lunch today?' the system will suggest the best restaurant and dish based on weather information and the user's preferences."
[1107] Example prompt: "Based on the user's healthcare data, please recommend a healthy meal for them."
[1108] In this way, the system integrates multiple data sets, provides users with the information and services they need in a unified manner, and improves their quality of life.
[1109] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1110] System processing flow
[1111] Step 1:
[1112] The user inputs voice into the terminal means. The user asks questions such as "What's the weather like today?" or "What do you recommend for lunch today?" This voice input is captured by the microphone in the terminal means.
[1113] Step 2:
[1114] The terminal means converts the user's voice input into text data using a voice recognition device, and the converted text data is sent to the server means for further processing, where the voice data is processed as text data.
[1115] Step 3:
[1116] The server analyzes the received text data. For example, it analyzes the text "Tell me the weather today" and recognizes that the user's intention is to obtain weather information. Based on the analysis result, it obtains the user's current location from the location information device.
[1117] Step 4:
[1118] The server means uses the acquired location information to send a request to the weather information API, which acquires the latest weather information from an external weather information service based on the location information.
[1119] Step 5:
[1120] The server means receives and analyzes the weather information returned from the weather information providing API. Based on the analysis results, the server means generates text data for a voice response. For example, the generated text data is "It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees."
[1121] Step 6:
[1122] The generated text data is transmitted from the server means to the terminal means. Specifically, the server means transfers the text data for speech synthesis to the terminal means.
[1123] Step 7:
[1124] The terminal means converts the transmitted text data into voice using a voice synthesizer. This provides the user with a specific weather forecast. The terminal means outputs a voice saying, "It's sunny today. The maximum temperature will be 25 degrees, and the minimum temperature will be 15 degrees."
[1125] Step 8:
[1126] If the user wants to ask more detailed information or another question, they can input the question into the terminal again by voice. For example, they can ask, "What do you recommend for lunch today?" In this case, the process starts again from step 1.
[1127] Detailed System Operation
[1128] 1. Capturing voice input: The user makes a voice input, which is captured by the microphone of the terminal means.
[1129] Input: User's voice
[1130] Output: Captured audio data
[1131] 2. Speech recognition: The captured voice data is analyzed by a voice recognizer and converted into text data.
[1132] Input: Audio data
[1133] Output: Text data
[1134] 3. Analysis of text data: The server receives the text data, analyzes it, and recognizes the user's intention.
[1135] Input: Text data
[1136] Output: User request information
[1137] 4. Obtaining location information: The server means obtains the user's current location from the location information device.
[1138] Input: User information
[1139] Output: Current location data
[1140] 5. Obtaining weather information: The server sends a request to the weather information API to obtain the latest weather information.
[1141] Input: Current location data
[1142] Output: Weather information
[1143] 6. Generation of voice response: The server means generates text data for a voice response based on the weather information.
[1144] Input: Weather information
[1145] Output: Text data
[1146] 7. Sending text data: The server means sends the generated text data to the terminal means.
[1147] Input: Text data
[1148] Output: Text data transferred to the terminal means
[1149] 8. Speech synthesis: The terminal means synthesizes text data into speech and outputs it as speech.
[1150] Input: Text data
[1151] Output: Audio output
[1152] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1153] The present invention is a system that manages user profile information and provides a variety of information and support, and is particularly capable of interacting with the user while taking their emotions into consideration. The system is primarily composed of a server, a stuffed toy-type terminal, communication means, a location information device, weather information acquisition means, audio output means, a healthcare device, and an emotion engine.
[1154] User profile registration
[1155] (overview)
[1156] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[1157] (Processing flow)
[1158] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[1159] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[1160] 3. The device uses voice recognition to convert the input information into text data.
[1161] 4. The device sends the converted text data to the server.
[1162] 5. The server stores the received profile information in a database.
[1163] Weather information provided
[1164] (overview)
[1165] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[1166] (Processing flow)
[1167] 1. The user asks the device, "What's the weather like today?"
[1168] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[1169] 3. The server obtains the user's location information and sends a request to the weather information API.
[1170] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1171] 5. The server returns the generated text data to the terminal.
[1172] 6. The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1173] Providing healthcare advice
[1174] (overview)
[1175] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[1176] (Processing flow)
[1177] 1. The user instructs the device, "Tell me how much exercise I did today."
[1178] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[1179] 3. The server obtains data from the healthcare device through the API.
[1180] 4. The server analyzes the acquired exercise data and generates advice for the user.
[1181] 5. The server returns the generated advice text to the terminal.
[1182] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[1183] Product proposals and orders
[1184] (overview)
[1185] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[1186] (Processing flow)
[1187] 1. The user asks the terminal, "How much shampoo is in stock?"
[1188] 2. The device uses voice recognition to convert the question into text data.
[1189] 3. The device sends the text data to the server.
[1190] 4. The server checks the user's purchasing history in the database and searches for inventory information.
[1191] 5. The server generates text data suggesting replenishment if stock is low or uncertain.
[1192] 6. The server returns the generated text data to the terminal.
[1193] 7. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[1194] 8. The user instructs the device to "purchase."
[1195] 9. The device uses voice recognition to convert the instructions into text data and send it to the server.
[1196] 10. The server receives the user's instructions and issues an order to the affiliated online store via API.
[1197] 11. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[1198] 12. The terminal will report to the user by voice, "Your shampoo order has been completed."
[1199] Implementing the Emotion Engine
[1200] (overview)
[1201] The system uses an emotion engine to recognize the user's emotions and interact with them based on those emotions, allowing for more personalized responses.
[1202] (Processing flow)
[1203] 1. The device acquires emotional data from voice and facial expressions while interacting with the user.
[1204] 2. The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[1205] 3. The device sends the generated emotional state text data to the server.
[1206] 4. The server generates appropriate feedback and advice for the user based on the received emotional data.
[1207] 5. The server sends the generated feedback and advice text data back to the device.
[1208] 6. The device synthesizes feedback and advice into voice and conveys it to the user.
[1209] Example: If the user is feeling stressed, the device might suggest, "Would you like me to play some music to help you calm down?"
[1210] In this way, by incorporating an emotion engine, it becomes possible to respond according to the user's emotional state, resulting in a system that provides more friendly and personalized services.
[1211] The processing flow will be explained below.
[1212] User profile registration
[1213] Step 1:
[1214] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[1215] Step 2:
[1216] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[1217] Step 3:
[1218] The device uses voice recognition to convert the input information into text data.
[1219] Step 4:
[1220] The terminal transmits the converted text data to the server.
[1221] Step 5:
[1222] The server stores the received profile information in a database.
[1223] Weather information provided
[1224] Step 1:
[1225] The user asks the device, "What's the weather like today?"
[1226] Step 2:
[1227] The device uses voice recognition to convert the question into text data.
[1228] Step 3:
[1229] The terminal transmits the text data to the server.
[1230] Step 4:
[1231] The server obtains the user's location information and sends a request to the weather information API.
[1232] Step 5:
[1233] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1234] Step 6:
[1235] The server returns the generated text data to the terminal.
[1236] Step 7:
[1237] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1238] Providing healthcare advice
[1239] Step 1:
[1240] The user instructs the terminal, "Tell me how much exercise I did today."
[1241] Step 2:
[1242] The device uses voice recognition to convert instructions into text data.
[1243] Step 3:
[1244] The terminal transmits the text data to the server.
[1245] Step 4:
[1246] The server obtains data from healthcare devices through an API.
[1247] Step 5:
[1248] The server analyzes the acquired exercise data and generates advice for the user.
[1249] Step 6:
[1250] The server returns the generated advice text data to the terminal.
[1251] Step 7:
[1252] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[1253] Product proposals and orders
[1254] Step 1:
[1255] The user asks the terminal, "How much shampoo is in stock?"
[1256] Step 2:
[1257] The device uses voice recognition to convert questions into text data.
[1258] Step 3:
[1259] The terminal transmits the text data to the server.
[1260] Step 4:
[1261] The server checks the user's purchasing history in a database and searches for inventory information.
[1262] Step 5:
[1263] The server generates text data suggesting replenishment when stock is low or uncertain.
[1264] Step 6:
[1265] The server returns the generated text data to the terminal.
[1266] Step 7:
[1267] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[1268] Step 8:
[1269] The user instructs the terminal to "purchase."
[1270] Step 9:
[1271] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[1272] Step 10:
[1273] The server receives the user's instructions and issues an order to the affiliated online store via API.
[1274] Step 11:
[1275] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[1276] Step 12:
[1277] The terminal will report to the user via voice, "Your shampoo order has been completed."
[1278] Implementing the Emotion Engine
[1279] Step 1:
[1280] During a conversation with the user, the device uses a microphone and camera to acquire emotional data from voice and facial expressions.
[1281] Step 2:
[1282] The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[1283] Step 3:
[1284] The terminal transmits the generated text data of the emotional state to the server.
[1285] Step 4:
[1286] The server generates appropriate feedback and advice for the user based on the received emotional data.
[1287] Step 5:
[1288] The server returns the generated feedback and advice text data to the terminal.
[1289] Step 6:
[1290] The device synthesizes feedback and advice into voice and conveys it to the user.
[1291] Example: If the user is feeling stressed, the device might suggest, "Would you like me to play some music to help you calm down?"
[1292] In this way, by incorporating an emotion engine, the system is able to respond according to the user's emotional state, providing a more friendly and personalized service.
[1293] Example 2
[1294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1295] In order to meet today's diverse and changing user needs, systems that can properly manage user profile information and provide personalized services based on that information are necessary. However, conventional systems have had difficulty comprehensively managing users' emotions, health status, and product purchase history, and providing feedback and advice in real time. To solve these problems, systems with more advanced recognition and response capabilities are required.
[1296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1297] In this invention, the server includes means for managing profile information about users, stuffed toy-type terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting profile information entered by the user to the server means, a location information device for acquiring location information of the user, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information into voice and providing it to the user, emotion recognition means having an emotion recognition device and acquiring emotion data from the user's voice and facial expressions and transmitting it to the server, and server means for analyzing the emotion data and generating appropriate feedback and advice. This enables personalized responses according to the user's emotional state and health condition, as well as real-time feedback and advice.
[1298] The "server means" is a computer processing device that manages user profile information, purchase history, emotional data, etc., and has data analysis and feedback generation functions.
[1299] The "stuffed toy type terminal means" is a stuffed toy type device that incorporates a microphone, a voice recognition device, and a voice synthesis device and provides a voice interface with the user.
[1300] "Communication means" refers to a means for transmitting and receiving data between a server and a terminal means, and includes wireless communication technologies such as Wi-Fi, 4G, and 5G.
[1301] A "location information device" is a device that obtains a user's current location and uses GPS or other location identification technology.
[1302] The "weather information acquisition means" is a means for using weather information acquired from an external weather information providing service.
[1303] The "audio output means" is a means for converting acquired information into audio and providing it to the user, and utilizes voice synthesis technology.
[1304] The "emotion recognition means" is a means for acquiring and analyzing emotion data from the user's voice and facial expressions.
[1305] A "healthcare device" is a device for acquiring data on the amount of exercise performed by a user and has a measurement function for health management.
[1306] "Online Store Communication Means" means a means for communicating with the e-commerce platform and placing an order.
[1307] The present invention is a system that manages user profile information and provides a variety of information and support. In particular, the present invention provides a system that enables interaction that takes into account the user's emotions. The system mainly consists of a server means, a stuffed toy-type terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a health care device, and an emotion recognition device.
[1308] The overall operation of the system proceeds as follows: First, the user registers profile information using the stuffed toy-type terminal means. The terminal means uses a built-in microphone and voice recognition device to convert the user's voice input into text data and transmits that data to the server means. The server means stores the received data in a local database. Through this procedure, the user's profile information is managed.
[1309] Next, when the user asks, "What's the weather like today?", the terminal means uses a voice recognition device to convert the question into text data and sends it to the server means. The server means acquires the user's location information and sends a request to a weather information service (e.g., OpenWeatherMap API) to acquire the latest weather information. The acquired weather information is analyzed, converted into a format that is easy for the user to understand, and sent to the terminal means. Finally, the terminal means uses a voice synthesis device to convey the text data to the user, thereby providing the weather information.
[1310] The system also works with healthcare devices (e.g., Fitbit or Apple Watch) to acquire data on the user's exercise volume. When the user asks, "Tell me how much exercise I did today," the terminal device sends the instruction to the server device, which then acquires the exercise volume data from the healthcare device. Based on the data, the system generates health advice and communicates it to the user via the terminal device.
[1311] Furthermore, it also has a function to provide product inventory information based on the user's purchase history. When a user asks, "How much shampoo is in stock?", the server means checks the purchase history in the database and searches for inventory information. If the stock is low, it suggests replenishing the product, and when the user gives the instruction to "purchase," it places an order with the affiliated e-commerce platform and notifies the terminal means of confirmation of the order completion.
[1312] Finally, an emotion recognition device is used to analyze the user's emotional state and provide appropriate feedback in real time. If the user feels stressed, the system will suggest, "Would you like me to play some music to help you relax?" This allows for more personalized responses.
[1313] (Example of a prompt)
[1314] "Please explain the process for registering profile information. The information to enter is name, age, and gender."
[1315] "Please explain the process for providing weather information based on your current location."
[1316] "Please explain the process for capturing exercise data and providing health advice to users."
[1317] "Please explain the process of suggesting products based on a user's purchase history and placing an order."
[1318] "Please explain your process for recognizing user emotions and providing appropriate feedback."
[1319] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1320] User profile registration
[1321] Step 1:
[1322] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[1323] Input: Start the system in initial setting mode
[1324] Output: Voice message "Nice to meet you! Please register your profile."
[1325] Specific operation: The device uses a voice synthesizer to call out to the user in a familiar voice.
[1326] Step 2:
[1327] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[1328] Input: Voice data such as name, age, gender, lifestyle patterns, etc.
[1329] Output: User's voice input
[1330] Specific operation: The user speaks, "My name is Yamada Taro. I am 30 years old. I am male."
[1331] Step 3:
[1332] The device uses Google Cloud Speech-to-Text to convert input voice information into text data.
[1333] Input: User voice input
[1334] Output: User profile information converted to text data
[1335] Specific operation: The device analyzes the voice data, converts it into text, and temporarily stores it in its internal memory.
[1336] Step 4:
[1337] The terminal transmits the converted text data to the server.
[1338] Input: User profile information converted to text data
[1339] Output: Send data to the server
[1340] Specific operation: The device uses Wi-Fi or 4G / 5G communication to send text data to the server.
[1341] Step 5:
[1342] The server stores the received profile information in a MySQL database.
[1343] Input: User profile information converted to text data
[1344] Output: Message that data has been saved to the server
[1345] Specific operation: The server analyzes the text data sent and stores it in the corresponding field in the database.
[1346] Weather information provided
[1347] Step 1:
[1348] The user asks the device, "What's the weather like today?"
[1349] Input: Say "What's the weather like today?"
[1350] Output: User's voice input
[1351] Specific operation: The user asks a question into the device, and the device's microphone picks up the voice.
[1352] Step 2:
[1353] The device uses voice recognition to convert the question into text data and send it to the server.
[1354] Input: Say "What's the weather like today?"
[1355] Output: Question content converted into text data
[1356] Specific operation: The device converts the voice into text data and sends it to the server using Wi-Fi or 4G / 5G communication.
[1357] Step 3:
[1358] The server obtains the user's location information and sends a request to a weather information API (e.g., OpenWeatherMap).
[1359] Input: Question content converted into text data
[1360] Output: Request sent to the weather API
[1361] Specific operation: The server retrieves the user's location information from a database or GPS information, and uses that location information to send a request to the weather information API.
[1362] Step 4:
[1363] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1364] Input: Weather information returned from the API (JSON format)
[1365] Output: Weather information text formatted for user
[1366] Specific operation: The server analyzes weather information and generates text data such as "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1367] Step 5:
[1368] The server returns the generated text data to the terminal.
[1369] Input: Weather information text formatted for the user
[1370] Output: Send text data to the terminal
[1371] Specific operation: The server sends the generated text data to the terminal.
[1372] Step 6:
[1373] The terminal synthesizes the received text data into voice and conveys it to the user.
[1374] Input: Weather information converted into text data
[1375] Output: Voiced weather information
[1376] Specific operation: The device uses Google Text-to-Speech to convert the text data into speech and tells the user, "Today is sunny. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1377] Providing healthcare advice
[1378] Step 1:
[1379] The user instructs the terminal, "Tell me how much exercise I did today."
[1380] Input: Say "How much exercise did I do today?"
[1381] Output: User's voice input
[1382] Specific operation: The user speaks to the device, "Tell me how much exercise I did today."
[1383] Step 2:
[1384] The device uses voice recognition to convert instructions into text data and send it to the server.
[1385] Input: Say "How much exercise did I do today?"
[1386] Output: User instructions converted to text data
[1387] Specific operation: The device converts the voice into text data and sends it to the server via Wi-Fi or 4G / 5G.
[1388] Step 3:
[1389] The server obtains data from healthcare devices through an API.
[1390] Input: User instructions converted into text data
[1391] Output: Sending a request to the API and retrieving exercise data
[1392] Specific operation: The server obtains the user's exercise data using the API of a healthcare device (e.g., Fitbit).
[1393] Step 4:
[1394] The server analyzes the acquired exercise data and generates advice for the user.
[1395] Input: Exercise data obtained from healthcare devices
[1396] Output: Text data generation of advice
[1397] Specific operation: The server analyzes the exercise data and generates advice text such as, "Today's exercise amount is 5,000 steps. You are halfway to your goal of 10,000 steps."
[1398] Step 5:
[1399] The server returns the generated advice text to the terminal.
[1400] Input: Text data of advice
[1401] Output: Send advice to terminal
[1402] Specific operation: The server sends the generated text data to the terminal.
[1403] Step 6:
[1404] The terminal converts the text data into speech and conveys it to the user.
[1405] Input: Text data of advice
[1406] Output: Spoken advice
[1407] Specific operation: The device uses Google Text-to-Speech to convert the text into audio and tells the user, "You've taken 5,000 steps today. You're halfway to your goal of 10,000 steps."
[1408] Product proposals and orders
[1409] Step 1:
[1410] The user asks the terminal, "How much shampoo is in stock?"
[1411] Input: Speak "What shampoo is in stock?"
[1412] Output: User's voice input
[1413] Specific operation: The user speaks to the terminal, "How much shampoo is in stock?"
[1414] Step 2:
[1415] The device uses voice recognition to convert the question into text data.
[1416] Input: Speak "What shampoo is in stock?"
[1417] Output: Question content converted into text data
[1418] Specific operation: The device converts the voice into text data.
[1419] Step 3:
[1420] The terminal transmits the text data to the server.
[1421] Input: Question content converted into text data
[1422] Output: Send data to the server
[1423] Specific operation: The device sends text data to the server using Wi-Fi or 4G / 5G.
[1424] Step 4:
[1425] The server checks the user's purchasing history in a database and searches for inventory information.
[1426] Input: Question content converted into text data
[1427] Output: Stock information search results
[1428] Specific operation: The server retrieves the user's purchase history from the database and searches for inventory information.
[1429] Step 5:
[1430] The server generates text data suggesting replenishment when stock is low or uncertain.
[1431] Input: Search results for inventory information
[1432] Output: Text data of inventory replenishment proposal
[1433] Specific operation: The server generates text data such as "Stock is low. Would you like to purchase?"
[1434] Step 6:
[1435] The server returns the generated text data to the terminal.
[1436] Input: Text data of inventory replenishment proposal
[1437] Output: Sending data to the terminal
[1438] Specific operation: The server sends the generated text data to the terminal.
[1439] Step 7:
[1440] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[1441] Input: Text data of inventory replenishment proposal
[1442] Output: Vocalized refill suggestions
[1443] Specific operation: The device uses Google Text-to-Speech to communicate with the user aloud.
[1444] Step 8:
[1445] The user instructs the terminal to "purchase."
[1446] Input: Say "Buy"
[1447] Output: User's voice input
[1448] Specific action: The user says "purchase."
[1449] Step 9:
[1450] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[1451] Input: Say "Buy"
[1452] Output: User instructions converted to text data
[1453] Specific operation: The device converts the user's voice into text data and sends it to the server.
[1454] Step 10:
[1455] The server receives the user's instructions and issues an order through the API of the affiliated online store.
[1456] Input: User instructions converted into text data
[1457] Output: Place order to online store
[1458] Specific operation: The server calls the online store's API and places an order.
[1459] Step 11:
[1460] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[1461] Input: Order completion notification from online store
[1462] Output: Order completion notification to terminal
[1463] Specific operation: The server receives confirmation of the order completion and notifies the terminal.
[1464] Step 12:
[1465] The terminal will report to the user via voice, "Your shampoo order has been completed."
[1466] Input: Notification of order completion
[1467] Output: Vocalized report
[1468] Specific operation: The device uses Google Text-to-Speech to report the completion of the order via voice.
[1469] Implementing the Emotion Engine
[1470] Step 1:
[1471] The terminal acquires emotional data from voice and facial expressions during a conversation with the user.
[1472] Input: User's voice and facial expression data
[1473] Output: Emotion data
[1474] Specific operation: The device uses an emotion recognition device to collect the user's voice and facial expression data.
[1475] Step 2:
[1476] The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[1477] Input: User's voice and facial expression data
[1478] Output: Emotional state converted into text data
[1479] Specific operation: The device analyzes the emotional data and generates text data that indicates the emotional state, such as "The user is feeling stressed."
[1480] Step 3:
[1481] The terminal transmits the generated text data of the emotional state to the server.
[1482] Input: Emotional state expressed as text data
[1483] Output: Send data to the server
[1484] Specific operation: The emotion data generated by the device is sent to the server.
[1485] Step 4:
[1486] The server generates appropriate feedback and advice for the user based on the received emotional data.
[1487] Input: Emotional state expressed as text data
[1488] Output: Text data of feedback and advice
[1489] Specific behavior: The server uses the emotional data to generate feedback and advice such as "Shall I play some music to help you relax?"
[1490] Step 5:
[1491] The server returns the generated feedback and advice text data to the terminal.
[1492] Input: Text data for feedback and advice
[1493] Output: Sending data to the terminal
[1494] Specific operation: The server sends the generated feedback to the device.
[1495] Step 6:
[1496] The device synthesizes feedback and advice into voice and conveys it to the user.
[1497] Input: Text data for feedback and advice
[1498] Output: Spoken feedback and advice
[1499] What it does: Your device will use Google Text-to-Speech to suggest aloud, "Would you like me to play some music to help you relax?"
[1500] The above is a specific description of the processing of the system.
[1501] (Application example 2)
[1502] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1503] Conventional systems provide limited information and services to users, and lack personalized support that takes emotion recognition into account. Furthermore, providing real-time, emotion-aware support is difficult for customer service in brick-and-mortar stores. The present invention aims to solve these problems and realize a system that provides interactive services tailored to the user's emotional state.
[1504] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user; communication means for communicating with the server means and transmitting profile information entered by the user to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; voice output means for synthesizing the acquired weather information and providing it to the user; emotion recognition means for acquiring and analyzing user emotion data and generating appropriate feedback and advice; and voice output means for synthesizing the feedback and advice generated by the emotion recognition means and providing it to the user. This enables the provision of highly personalized information and service support based on the user's emotional state.
[1505] The "server means" is a central system that manages user profile information and emotional data, collects and analyzes external information, and generates appropriate feedback.
[1506] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and that provides a voice interface with the user.
[1507] The "communication means" is a means for transmitting and receiving data between the server means and the terminal means.
[1508] A "location information device" is a device for acquiring the current location of a user.
[1509] The "weather information acquisition means" is a means for acquiring weather information from an external weather information providing service based on the acquired location information.
[1510] The "voice output means" is a means for synthesizing acquired or generated information into voice and conveying it to the user.
[1511] The "emotion recognition means" is a means for acquiring and analyzing the user's emotional data to generate appropriate feedback and advice.
[1512] "Feedback" refers to responses or advice to the user that are generated based on the user's emotional data and other information.
[1513] "Personalization" refers to providing information and services that are customized according to a user's individual profile and emotional state.
[1514] The system for implementing this invention provides personalized information based on a user's profile information and emotion recognition. Specifically, it is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and an emotion recognition means.
[1515] First, the server manages user profile information and collects and analyzes external information. The server includes a database system for storing profile information and emotional data. Furthermore, the server analyzes data using Natural Language Processing (NLP) APIs and speech recognition APIs from Azure and Google Cloud Services. Based on the results of this analysis, the server generates appropriate feedback and advice for the user.
[1516] The terminal is a device that provides a voice interface with the user and is equipped with a microphone, a voice recognition device, and a voice synthesis device. When the user inputs a question or instruction by voice, it is converted into text data and sent to the server. The converted text data is analyzed by the server and appropriate feedback is generated. The generated feedback is sent back to the terminal and conveyed to the user via voice output means.
[1517] The communication means is used to transmit and receive data between the server and the terminal. This communication is performed in real time, and the information required by the user is provided promptly.
[1518] The location information device is used to acquire the user's current location, and based on the acquired location information, the weather information acquisition means acquires weather information from an external weather information service, allowing the user to know the weather information for their current location in real time.
[1519] The emotion recognition means acquires and analyzes the user's emotion data to generate appropriate feedback and advice. This emotion recognition is used to understand the user's emotional state from voice input and text data and provide the optimal response to the user. For example, if the user is feeling stressed, it is possible to suggest relaxing music.
[1520] The following are specific examples of embodiments:
[1521] When a user speaks into their smartphone, "What's today's promotion?", the speech is converted into text data and sent to the server. The server analyzes the text data and generates promotion information, taking into account the user's emotional state. The device then synthesizes the generated promotion information into voice and provides it to the user.
[1522] Examples of such prompts include "Determine whether the user is interested in sale days and provide appropriate promotional information" and "Determine the emotion of the following text:- Text: {user's voice input}- Emotion: happy, sad, angry, neutral."
[1523] This allows users to get the information they need in real time within a physical store and receive personalized support based on their emotional state.
[1524] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1525] Step 1:
[1526] The user speaks a question or command into the device. For example, they might say, "What are today's promotions?" The voice input is picked up by the device's microphone.
[1527] Step 2:
[1528] The terminal analyzes the voice input with a voice recognition device and converts it into text data. After this text data is generated, it is sent to the server. The input is voice data and the output is text data.
[1529] Step 3:
[1530] The server analyzes the received text data using a Natural Language Processing (NLP) API. During the analysis, it also extracts the question content and emotional part. This analysis reveals the user's question content and emotional state. The input is text data, and the output is the analysis result.
[1531] Step 4:
[1532] The server generates an appropriate response based on the analysis results and profile information. For example, if the question is about a promotion, it generates promotion information taking into account inventory information and cross-selling information. The response also changes depending on the emotional state. The input is the analysis results and profile data, and the output is the response data.
[1533] Step 5:
[1534] The server then personalizes the generated response data according to the user's emotional state and sends the response in text format to the terminal. The input is the response data, and the output is the personalized text data.
[1535] Step 6:
[1536] The terminal synthesizes the received personalized text data using a voice output means and conveys the speech to the user. For example, it might say, "Today is sale day! Your favorite product is 20% off!" The input is personalized text data, and the output is voice data.
[1537] Step 7:
[1538] After receiving a response from the device, the user acts based on the content. For example, a user receives information about a sale day and decides which specific product to purchase. This step does not provide any feedback to the system, but it may affect future interactions. The input is voice data, and the output is the user's actions.
[1539] This allows users to receive personalized information and responses in real time that are tailored to their emotions.
[1540] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1541] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1542] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1543] [Third embodiment]
[1544] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1545] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1546] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1547] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1548] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1549] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1550] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1551] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1552] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1553] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1554] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1555] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1556] The present invention is a system that manages profile information about users and provides a variety of information and support, and is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, an audio output means, and a healthcare device.
[1557] User profile registration
[1558] (overview)
[1559] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[1560] (Processing flow)
[1561] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[1562] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[1563] 3. The device uses voice recognition to convert the input information into text data.
[1564] 4. The device sends the converted text data to the server.
[1565] 5. The server stores the received profile information in a database.
[1566] Weather information provided
[1567] (overview)
[1568] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[1569] (Processing flow)
[1570] 1. The user asks the device, "What's the weather like today?"
[1571] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[1572] 3. The server obtains the user's location information and sends a request to the weather information API.
[1573] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1574] 5. The server returns the generated text data to the terminal.
[1575] 6. The device synthesizes the received text data into voice and tells the user, for example, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1576] Providing healthcare advice
[1577] (overview)
[1578] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[1579] (Processing flow)
[1580] 1. The user instructs the device, "Tell me how much exercise I did today."
[1581] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[1582] 3. The server obtains data from the healthcare device through the API.
[1583] 4. The server analyzes the acquired exercise data and generates advice for the user.
[1584] 5. The server returns the generated advice text to the terminal.
[1585] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[1586] Product proposals and orders
[1587] (overview)
[1588] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[1589] (Processing flow)
[1590] 1. The user asks the terminal, "How much shampoo is in stock?"
[1591] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[1592] 3. The server checks the user's purchasing history in the database and searches for inventory information.
[1593] 4. The server generates text data suggesting replenishment if stock is low or uncertain.
[1594] 5. The server returns the generated text data to the terminal.
[1595] 6. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[1596] 7. The user instructs the device to "purchase."
[1597] 8. The device uses voice recognition to convert the instructions into text data and send it to the server.
[1598] 9. The server receives the user's instructions and issues an order to the affiliated online store via API.
[1599] 10. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[1600] 11. The terminal will report to the user by voice, "Your shampoo order has been completed."
[1601] In this way, the system of the present invention enriches the user's life through a series of operations and provides multifunctional support with a user-friendly interface.
[1602] The processing flow will be explained below.
[1603] User profile registration
[1604] Step 1:
[1605] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[1606] Step 2:
[1607] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[1608] Step 3:
[1609] The device uses voice recognition to convert the input information into text data.
[1610] Step 4:
[1611] The terminal transmits the converted text data to the server.
[1612] Step 5:
[1613] The server stores the received profile information in a database.
[1614] Weather information provided
[1615] Step 1:
[1616] The user asks the device, "What's the weather like today?"
[1617] Step 2:
[1618] The device uses voice recognition to convert the question into text data.
[1619] Step 3:
[1620] The terminal transmits the text data to the server.
[1621] Step 4:
[1622] The server obtains the user's location information and sends a request to the weather information API.
[1623] Step 5:
[1624] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1625] Step 6:
[1626] The server returns the generated text data to the terminal.
[1627] Step 7:
[1628] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1629] Providing healthcare advice
[1630] Step 1:
[1631] The user instructs the terminal, "Tell me how much exercise I did today."
[1632] Step 2:
[1633] The device uses voice recognition to convert instructions into text data.
[1634] Step 3:
[1635] The terminal transmits the text data to the server.
[1636] Step 4:
[1637] The server obtains data from healthcare devices through an API.
[1638] Step 5:
[1639] The server analyzes the acquired exercise data and generates advice for the user.
[1640] Step 6:
[1641] The server returns the generated advice text to the terminal.
[1642] Step 7:
[1643] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[1644] Product proposals and orders
[1645] Step 1:
[1646] The user asks the terminal, "How much shampoo is in stock?"
[1647] Step 2:
[1648] The device uses voice recognition to convert questions into text data.
[1649] Step 3:
[1650] The terminal transmits the text data to the server.
[1651] Step 4:
[1652] The server checks the user's purchasing history in a database and searches for inventory information.
[1653] Step 5:
[1654] The server generates text data suggesting replenishment when stock is low or uncertain.
[1655] Step 6:
[1656] The server returns the generated text data to the terminal.
[1657] Step 7:
[1658] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[1659] Step 8:
[1660] The user instructs the terminal to "purchase."
[1661] Step 9:
[1662] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[1663] Step 10:
[1664] The server receives the user's instructions and issues an order to the affiliated online store via API.
[1665] Step 11:
[1666] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[1667] Step 12:
[1668] The terminal will report to the user via voice, "Your shampoo order has been completed."
[1669] Example 1
[1670] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1671] Conventional voice interface systems have difficulty effectively managing user profile information, location information, physical activity data, and purchase history, and providing a variety of information and support. Furthermore, the complexity of the system when integrating multiple functions and improving the user experience have also been issues.
[1672] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1673] In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with users; communication means for communicating with the server means and transmitting the profile information input by the users to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; and voice output means for synthesizing the acquired weather information and providing it to the user. This makes it possible to integrate and manage the user's profile information, location information, exercise data, and purchase history, and to provide information and support efficiently and effectively.
[1674] The "server means" is a device or system for managing profile information about users and processing and storing data in cooperation with various databases.
[1675] The "terminal means" is a device or system that includes a microphone, a voice recognition device, and a voice synthesis device, provides a voice interface with the user, and collects input data from the user.
[1676] "Communication means" refers to a device or system for transmitting and receiving data between terminal means and server means.
[1677] A "location information device" is a device for acquiring a user's current location, and mainly uses location information technology such as GPS.
[1678] The "weather information acquisition means" is a device or system for acquiring the latest weather information from an external weather information providing means based on the acquired user's location information.
[1679] The "audio output means" is a device or system that converts acquired weather information and various data into audio and provides it to the user.
[1680] A "healthcare device" is a device for acquiring a user's exercise data, and primarily includes fitness trackers and smartwatches.
[1681] The "online store communication means" is a device or system for issuing an order to the online store based on the user's purchase instructions.
[1682] MODE FOR CARRYING OUT THE INVENTION
[1683] The present invention is a system that manages a user's profile information, location information, exercise data, and purchase history, and provides a variety of information and support based on this data. The system is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and a healthcare device.
[1684] User profile registration
[1685] overview
[1686] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[1687] Detailed Description
[1688] The device enters initial setup mode and starts a voice prompt saying, "Nice to meet you! Please register your profile." The user uses the microphone to input information such as their name, age, gender, and lifestyle patterns by voice. The device uses a voice recognition engine (for example, Google Voice Recognition API) to convert the voice into text data. The converted text data is sent to the server using a communication method. The server stores the received data in a database.
[1689] Specific examples
[1690] For example, if a user speaks "My name is Yamada Taro and I'm 30 years old," the device converts it into text and the server stores it in a database. An example of a prompt is "Write a program that inputs a user's name and age and stores that information in a database."
[1691] Weather information provided
[1692] overview
[1693] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[1694] Detailed Description
[1695] When a user asks a device, "What's the weather like today?", the device uses its voice recognition function to convert the question into text data and sends it to the server. The server uses a location information device to obtain the user's current location and sends a request to a weather information providing API (for example, OpenWeatherMap API). The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand. The generated text data is sent back to the device and conveyed to the user using the voice synthesis function.
[1696] Specific examples
[1697] When the user asks, "What's the weather like today?", the device responds, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees." An example of a prompt is, "Write a program that obtains weather information based on the user's location and relays it to them in voice."
[1698] Providing healthcare advice
[1699] overview
[1700] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[1701] Detailed Description
[1702] When a user instructs the device to "tell me how much exercise I did today," the device uses voice recognition to convert the instruction into text data and send it to the server. The server acquires the exercise data via the healthcare device, analyzes it, and generates advice. The advice is then sent back to the device, converted into voice, and conveyed to the user.
[1703] Specific examples
[1704] When the user asks, "How much exercise did you do today?", the device responds, "You've done 5,000 steps today. You're halfway to your goal of 10,000 steps." An example of a prompt is, "Write a program that obtains a user's exercise data and provides advice based on that data."
[1705] Product proposals and orders
[1706] overview
[1707] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[1708] Detailed Description
[1709] When a user asks the terminal, "What shampoo is in stock?", the terminal uses voice recognition to convert the question into text data and sends it to the server. The server checks the purchase history database to find inventory information. If inventory is low or unclear, the server generates a replenishment suggestion and sends it back to the terminal. The terminal notifies the user of this by voice, receives the user's purchase instructions, and issues an order.
[1710] Specific examples
[1711] For example, if a user asks, "How much shampoo is in stock?" and the terminal responds, "Your shampoo is low in stock. Would you like to buy more?", and the user responds by saying, "Purchase," the terminal reports, "Your shampoo order is complete." An example of a prompt sentence is, "Write a program that provides inventory information based on the user's purchasing history and places an order for the product if necessary."
[1712] In this way, the system of the present invention provides multifunctional services through detailed and specific processes to support the user's life.
[1713] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1714] User profile registration
[1715] Step 1:
[1716] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[1717] Specific behavior: The device plays a message aloud, prompting the user to prepare to start typing.
[1718] Input: Command to start the initial setting mode on the terminal side
[1719] Output: Start of voice prompt
[1720] Step 2:
[1721] The user follows the instructions on the device and enters their information by voice, for example, saying, "My name is Taro Yamada and I'm 30 years old."
[1722] Specific operation: The user verbally transmits their profile information through a microphone.
[1723] Input: Voice input (user profile information)
[1724] Output: Audio data
[1725] Step 3:
[1726] The device uses a voice recognition engine to convert the input voice data into text data, using the Google Voice Recognition API.
[1727] Specific operation: Converts voice data into text in real time.
[1728] Input: Audio data
[1729] Output: Text data
[1730] Step 4:
[1731] The device sends the converted text data to the server via a communication method, using an AWS S3 bucket.
[1732] Specific operation: Text data is transferred to the server using a communication protocol.
[1733] Input: Text data
[1734] Output: Send data to the server
[1735] Step 5:
[1736] The server stores the received profile information in a database, using a MySQL database.
[1737] Specific behavior: Saves and validates data.
[1738] Input: Text data
[1739] Output: Save information to a database
[1740] Weather information provided
[1741] Step 1:
[1742] The user asks the device, "What's the weather like today?" and inputs the request by voice.
[1743] Specific operation: The user verbally instructs the device that he or she wants to know weather information.
[1744] Input: Voice input (weather information request)
[1745] Output: Audio data
[1746] Step 2:
[1747] The device uses its voice recognition function to convert the question into text data and send it to the server, using the Google Voice Recognition API.
[1748] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[1749] Input: Audio data
[1750] Output: Text data (weather information request)
[1751] Step 3:
[1752] The server uses the location device to obtain the user's current location and sends a request to the OpenWeatherMap API.
[1753] Specific operation: Obtains location information and sends a request to an external API.
[1754] Input: Text data, location information
[1755] Output: Weather information request
[1756] Step 4:
[1757] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1758] What it does: Parses weather information and formats it for the user.
[1759] Input: Weather information data
[1760] Output: Text data (analyzed weather information)
[1761] Step 5:
[1762] The server returns the generated text data to the terminal, which converts it into voice and conveys it to the user.
[1763] Specific operation: Text data is synthesized into speech and information is provided to the user.
[1764] Input: Text data (analyzed weather information)
[1765] Output: Audio output (weather information)
[1766] Providing healthcare advice
[1767] Step 1:
[1768] The user instructs the device, "Tell me how much exercise I did today." The request is input by voice.
[1769] Specific operation: The user gives voice instructions to the terminal that he / she wants to know the exercise amount data.
[1770] Input: Voice input (request for exercise data)
[1771] Output: Audio data
[1772] Step 2:
[1773] The device uses voice recognition to convert instructions into text data and send it to the server.
[1774] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[1775] Input: Audio data
[1776] Output: Text data (momentum data request)
[1777] Step 3:
[1778] The server obtains exercise data via the healthcare device through the FitBit API.
[1779] Specific behavior: Uses an external API to obtain momentum data.
[1780] Input: Request data
[1781] Output: Momentum data
[1782] Step 4:
[1783] The server analyzes the acquired exercise data and generates advice for the user.
[1784] Specific operation: Analyze the momentum data and generate advice as text data.
[1785] Input: Momentum data
[1786] Output: Text data (advice)
[1787] Step 5:
[1788] The server returns the generated text data of the advice to the terminal, which converts it into voice and conveys it to the user.
[1789] Specific operation: Text data is synthesized into speech and information is provided to the user.
[1790] Input: Text data (advice)
[1791] Output: Audio output (advice)
[1792] Product proposals and orders
[1793] Step 1:
[1794] The user asks the terminal, "What shampoo is in stock?" and inputs the request by voice.
[1795] Specific operation: The user verbally instructs the terminal that he / she wants to know inventory information.
[1796] Input: Voice input (request for stock information)
[1797] Output: Audio data
[1798] Step 2:
[1799] The device uses voice recognition to convert the question into text data and send it to the server.
[1800] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[1801] Input: Audio data
[1802] Output: Text data (request for inventory information)
[1803] Step 3:
[1804] The server checks the user's purchasing history in a database and searches for inventory information.
[1805] Specific operation: Retrieve purchase history from the database and check inventory information.
[1806] Input: Text data (request for inventory information)
[1807] Output: Inventory information
[1808] Step 4:
[1809] The server generates replenishment suggestions when inventory is low or uncertain.
[1810] Specific operation: Generate replenishment suggestions as text data based on inventory information.
[1811] Input: Inventory information
[1812] Output: Text data (supplement proposal)
[1813] Step 5:
[1814] The server returns the generated text data of the proposal to the terminal, which converts it into voice and conveys it to the user.
[1815] Specific operation: Text data is synthesized into speech and information is provided to the user.
[1816] Input: Text data (supplement proposal)
[1817] Output: Audio output (supplement suggestion)
[1818] Step 6:
[1819] The user instructs the terminal to "purchase," and the terminal converts the instruction into text using voice recognition and sends it to the server.
[1820] Specific operation: Converts voice input into text data and sends it to the server.
[1821] Input: Voice input (purchase instructions)
[1822] Output: Text data (purchase instructions)
[1823] Step 7:
[1824] The server receives the user's instructions and places an order with the affiliated online store.
[1825] What it does: Places an order using the online store's API.
[1826] Input: Text data (purchase instructions)
[1827] Output: Order data
[1828] Step 8:
[1829] The server receives confirmation that the order has been completed and sends a notification to the terminal, which then tells the user, "Your shampoo order has been completed."
[1830] Specific operation: The notification data is synthesized into voice and the information is provided to the user.
[1831] Input: Order completion notification
[1832] Output: Audio output (order completion notification)
[1833] In this way, by performing specific input, data processing, and output at each step, multifunctional services are realized for users.
[1834] (Application example 1)
[1835] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1836] In modern society, users lead busy lives and require timely recommendations for optimal products and services based on weather information and health status to facilitate smooth purchasing behavior. It is also necessary to improve the quality of life by managing this information in an integrated manner and providing appropriate advice to users. However, conventional systems have had difficulty integrating various data, such as user profiles, location information, weather information, purchase history, and health data, to provide consistent services to users.
[1837] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1838] In this invention, the server includes server means for managing profile information about users, terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting the profile information entered by the user to the server means, a location information device for acquiring user location information, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information and providing it to the user, proposal generation means for proposing optimal products and services based on the acquired weather information and user preference information, and health support means for proposing optimal health foods and lifestyle-related products based on the user's purchase history and health data. This allows users to obtain information necessary for their daily lives in a centralized manner and improve their quality of life.
[1839] The "server means" is a central computer device that manages user profile information, purchase history, health data, etc., and integrates various data to provide optimal services to users.
[1840] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and provides a voice interface with the user.
[1841] The "communication means" is an interface for the terminal means to send and receive data to and from the server means.
[1842] A "location information device" is a device for obtaining the current location of a user.
[1843] The "weather information acquisition means" is a means for acquiring the latest weather information from an external weather information providing service based on the user's location information.
[1844] The "audio output means" is a device that synthesizes acquired information into voice and provides it to the user in the form of voice.
[1845] The "proposal generating means" is a means for proposing optimal products and services based on the acquired weather information and user preference information.
[1846] "Health support tools" are tools for suggesting optimal health foods and lifestyle-related products based on a user's purchasing history and health data.
[1847] This invention is a system that provides various information and support based on user profile information. The system is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a suggestion generation means, and a health support means.
[1848] Hardware and software used
[1849] 1. Server means: A central computer device (e.g., a cloud server) that manages user profile information, purchase history, and health data.
[1850] 2. Terminal means: A device equipped with a microphone, a voice recognition device, and a voice synthesis device (examples include smartphones, smart glasses, and head-mounted displays).
[1851] 3. Communication means: An interface that transmits and receives data between terminal means and server means (example: Internet communication).
[1852] 4. Location information device: A device that obtains the user's current location (example: GPS).
[1853] 5. Weather information acquisition method: A method for acquiring the latest weather information from an external weather information service (e.g., weather API).
[1854] 6. Voice output means: A device that provides information to the user by voice synthesis (example: speaker).
[1855] 7. Proposal generation means: A means for proposing optimal products and services based on weather information and user preference information (specific example: proposal generation algorithm).
[1856] 8. Health support tools: Tools that suggest optimal health foods and lifestyle-related products based on a user's purchasing history and health data (example: health advice algorithm).
[1857] Data processing and calculation
[1858] Registration of profile information: The user enters his / her profile information by voice through the terminal means, and the voice recognition device converts the information into text data and transmits it to the server means.
[1859] Obtaining weather information: Based on the user's location information obtained by the location information device, the weather information providing means obtains the latest weather information, and the audio output means provides the information by audio.
[1860] Proposal generation: The proposal generator integrates weather information and user preference information to suggest the most suitable products and services to the user.
[1861] Health support: Health support tools analyze users' purchasing history and health data to recommend optimal health foods and lifestyle-related products.
[1862] Examples of concrete examples and prompts
[1863] Example 1: "When a user asks their device, 'What's the weather like today?' the system obtains their location, retrieves weather information from a weather API, and provides, via voice synthesis, the answer, 'It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees.'"
[1864] Example prompt: "Suggest restaurants that deliver to a given user based on their location and weather."
[1865] Example 2: "When a user asks the device, 'What do you recommend for lunch today?' the system will suggest the best restaurant and dish based on weather information and the user's preferences."
[1866] Example prompt: "Based on the user's healthcare data, please recommend a healthy meal for them."
[1867] In this way, the system integrates multiple data sets, provides users with the information and services they need in a unified manner, and improves their quality of life.
[1868] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1869] System processing flow
[1870] Step 1:
[1871] The user inputs voice into the terminal means. The user asks questions such as "What's the weather like today?" or "What do you recommend for lunch today?" This voice input is captured by the microphone in the terminal means.
[1872] Step 2:
[1873] The terminal means converts the user's voice input into text data using a voice recognition device, and the converted text data is sent to the server means for further processing, where the voice data is processed as text data.
[1874] Step 3:
[1875] The server analyzes the received text data. For example, it analyzes the text "Tell me the weather today" and recognizes that the user's intention is to obtain weather information. Based on the analysis result, it obtains the user's current location from the location information device.
[1876] Step 4:
[1877] The server means uses the acquired location information to send a request to the weather information API, which acquires the latest weather information from an external weather information service based on the location information.
[1878] Step 5:
[1879] The server means receives and analyzes the weather information returned from the weather information providing API. Based on the analysis results, the server means generates text data for a voice response. For example, the generated text data is "It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees."
[1880] Step 6:
[1881] The generated text data is transmitted from the server means to the terminal means. Specifically, the server means transfers the text data for speech synthesis to the terminal means.
[1882] Step 7:
[1883] The terminal means converts the transmitted text data into voice using a voice synthesizer. This provides the user with a specific weather forecast. The terminal means outputs a voice saying, "It's sunny today. The maximum temperature will be 25 degrees, and the minimum temperature will be 15 degrees."
[1884] Step 8:
[1885] If the user wants to ask more detailed information or another question, they can input the question into the terminal again by voice. For example, they can ask, "What do you recommend for lunch today?" In this case, the process starts again from step 1.
[1886] Detailed System Operation
[1887] 1. Capturing voice input: The user makes a voice input, which is captured by the microphone of the terminal means.
[1888] Input: User's voice
[1889] Output: Captured audio data
[1890] 2. Speech recognition: The captured voice data is analyzed by a voice recognizer and converted into text data.
[1891] Input: Audio data
[1892] Output: Text data
[1893] 3. Analysis of text data: The server receives the text data, analyzes it, and recognizes the user's intention.
[1894] Input: Text data
[1895] Output: User request information
[1896] 4. Obtaining location information: The server means obtains the user's current location from the location information device.
[1897] Input: User information
[1898] Output: Current location data
[1899] 5. Obtaining weather information: The server sends a request to the weather information API to obtain the latest weather information.
[1900] Input: Current location data
[1901] Output: Weather information
[1902] 6. Generation of voice response: The server means generates text data for a voice response based on the weather information.
[1903] Input: Weather information
[1904] Output: Text data
[1905] 7. Sending text data: The server means sends the generated text data to the terminal means.
[1906] Input: Text data
[1907] Output: Text data transferred to the terminal means
[1908] 8. Speech synthesis: The terminal means synthesizes text data into speech and outputs it as speech.
[1909] Input: Text data
[1910] Output: Audio output
[1911] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1912] The present invention is a system that manages user profile information and provides a variety of information and support, and is particularly capable of interacting with the user while taking their emotions into consideration. The system is primarily composed of a server, a stuffed toy-type terminal, communication means, a location information device, weather information acquisition means, audio output means, a healthcare device, and an emotion engine.
[1913] User profile registration
[1914] (overview)
[1915] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[1916] (Processing flow)
[1917] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[1918] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[1919] 3. The device uses voice recognition to convert the input information into text data.
[1920] 4. The device sends the converted text data to the server.
[1921] 5. The server stores the received profile information in a database.
[1922] Weather information provided
[1923] (overview)
[1924] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[1925] (Processing flow)
[1926] 1. The user asks the device, "What's the weather like today?"
[1927] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[1928] 3. The server obtains the user's location information and sends a request to the weather information API.
[1929] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1930] 5. The server returns the generated text data to the terminal.
[1931] 6. The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1932] Providing healthcare advice
[1933] (overview)
[1934] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[1935] (Processing flow)
[1936] 1. The user instructs the device, "Tell me how much exercise I did today."
[1937] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[1938] 3. The server obtains data from the healthcare device through the API.
[1939] 4. The server analyzes the acquired exercise data and generates advice for the user.
[1940] 5. The server returns the generated advice text to the terminal.
[1941] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[1942] Product proposals and orders
[1943] (overview)
[1944] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[1945] (Processing flow)
[1946] 1. The user asks the terminal, "How much shampoo is in stock?"
[1947] 2. The device uses voice recognition to convert the question into text data.
[1948] 3. The device sends the text data to the server.
[1949] 4. The server checks the user's purchasing history in the database and searches for inventory information.
[1950] 5. The server generates text data suggesting replenishment if stock is low or uncertain.
[1951] 6. The server returns the generated text data to the terminal.
[1952] 7. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[1953] 8. The user instructs the device to "purchase."
[1954] 9. The device uses voice recognition to convert the instructions into text data and send it to the server.
[1955] 10. The server receives the user's instructions and issues an order to the affiliated online store via API.
[1956] 11. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[1957] 12. The terminal will report to the user by voice, "Your shampoo order has been completed."
[1958] Implementing the Emotion Engine
[1959] (overview)
[1960] The system uses an emotion engine to recognize the user's emotions and interact with them based on those emotions, allowing for more personalized responses.
[1961] (Processing flow)
[1962] 1. The device acquires emotional data from voice and facial expressions while interacting with the user.
[1963] 2. The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[1964] 3. The device sends the generated emotional state text data to the server.
[1965] 4. The server generates appropriate feedback and advice for the user based on the received emotional data.
[1966] 5. The server sends the generated feedback and advice text data back to the device.
[1967] 6. The device synthesizes feedback and advice into voice and conveys it to the user.
[1968] Example: If the user is feeling stressed, the device might suggest, "Would you like me to play some music to help you calm down?"
[1969] In this way, by incorporating an emotion engine, it becomes possible to respond according to the user's emotional state, resulting in a system that provides more friendly and personalized services.
[1970] The processing flow will be explained below.
[1971] User profile registration
[1972] Step 1:
[1973] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[1974] Step 2:
[1975] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[1976] Step 3:
[1977] The device uses voice recognition to convert the input information into text data.
[1978] Step 4:
[1979] The terminal transmits the converted text data to the server.
[1980] Step 5:
[1981] The server stores the received profile information in a database.
[1982] Weather information provided
[1983] Step 1:
[1984] The user asks the device, "What's the weather like today?"
[1985] Step 2:
[1986] The device uses voice recognition to convert the question into text data.
[1987] Step 3:
[1988] The terminal transmits the text data to the server.
[1989] Step 4:
[1990] The server obtains the user's location information and sends a request to the weather information API.
[1991] Step 5:
[1992] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[1993] Step 6:
[1994] The server returns the generated text data to the terminal.
[1995] Step 7:
[1996] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[1997] Providing healthcare advice
[1998] Step 1:
[1999] The user instructs the terminal, "Tell me how much exercise I did today."
[2000] Step 2:
[2001] The device uses voice recognition to convert instructions into text data.
[2002] Step 3:
[2003] The terminal transmits the text data to the server.
[2004] Step 4:
[2005] The server obtains data from healthcare devices through an API.
[2006] Step 5:
[2007] The server analyzes the acquired exercise data and generates advice for the user.
[2008] Step 6:
[2009] The server returns the generated advice text data to the terminal.
[2010] Step 7:
[2011] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[2012] Product proposals and orders
[2013] Step 1:
[2014] The user asks the terminal, "How much shampoo is in stock?"
[2015] Step 2:
[2016] The device uses voice recognition to convert questions into text data.
[2017] Step 3:
[2018] The terminal transmits the text data to the server.
[2019] Step 4:
[2020] The server checks the user's purchasing history in a database and searches for inventory information.
[2021] Step 5:
[2022] The server generates text data suggesting replenishment when stock is low or uncertain.
[2023] Step 6:
[2024] The server returns the generated text data to the terminal.
[2025] Step 7:
[2026] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[2027] Step 8:
[2028] The user instructs the terminal to "purchase."
[2029] Step 9:
[2030] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[2031] Step 10:
[2032] The server receives the user's instructions and issues an order to the affiliated online store via API.
[2033] Step 11:
[2034] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[2035] Step 12:
[2036] The terminal will report to the user via voice, "Your shampoo order has been completed."
[2037] Implementing the Emotion Engine
[2038] Step 1:
[2039] During a conversation with the user, the device uses a microphone and camera to acquire emotional data from voice and facial expressions.
[2040] Step 2:
[2041] The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[2042] Step 3:
[2043] The terminal transmits the generated text data of the emotional state to the server.
[2044] Step 4:
[2045] The server generates appropriate feedback and advice for the user based on the received emotional data.
[2046] Step 5:
[2047] The server returns the generated feedback and advice text data to the terminal.
[2048] Step 6:
[2049] The device synthesizes feedback and advice into voice and conveys it to the user.
[2050] Example: If the user is feeling stressed, the device might suggest, "Would you like me to play some music to help you calm down?"
[2051] In this way, by incorporating an emotion engine, the system is able to respond according to the user's emotional state, providing a more friendly and personalized service.
[2052] Example 2
[2053] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2054] In order to meet today's diverse and changing user needs, systems that can properly manage user profile information and provide personalized services based on that information are necessary. However, conventional systems have had difficulty comprehensively managing users' emotions, health status, and product purchase history, and providing feedback and advice in real time. To solve these problems, systems with more advanced recognition and response capabilities are required.
[2055] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2056] In this invention, the server includes means for managing profile information about users, stuffed toy-type terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting profile information entered by the user to the server means, a location information device for acquiring location information of the user, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information into voice and providing it to the user, emotion recognition means having an emotion recognition device and acquiring emotion data from the user's voice and facial expressions and transmitting it to the server, and server means for analyzing the emotion data and generating appropriate feedback and advice. This enables personalized responses according to the user's emotional state and health condition, as well as real-time feedback and advice.
[2057] The "server means" is a computer processing device that manages user profile information, purchase history, emotional data, etc., and has data analysis and feedback generation functions.
[2058] The "stuffed toy type terminal means" is a stuffed toy type device that incorporates a microphone, a voice recognition device, and a voice synthesis device and provides a voice interface with the user.
[2059] "Communication means" refers to a means for transmitting and receiving data between a server and a terminal means, and includes wireless communication technologies such as Wi-Fi, 4G, and 5G.
[2060] A "location information device" is a device that obtains a user's current location and uses GPS or other location identification technology.
[2061] The "weather information acquisition means" is a means for using weather information acquired from an external weather information providing service.
[2062] The "audio output means" is a means for converting acquired information into audio and providing it to the user, and utilizes voice synthesis technology.
[2063] The "emotion recognition means" is a means for acquiring and analyzing emotion data from the user's voice and facial expressions.
[2064] A "healthcare device" is a device for acquiring data on the amount of exercise performed by a user and has a measurement function for health management.
[2065] "Online Store Communication Means" means a means for communicating with the e-commerce platform and placing an order.
[2066] The present invention is a system that manages user profile information and provides a variety of information and support. In particular, the present invention provides a system that enables interaction that takes into account the user's emotions. The system mainly consists of a server means, a stuffed toy-type terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a health care device, and an emotion recognition device.
[2067] The overall operation of the system proceeds as follows: First, the user registers profile information using the stuffed toy-type terminal means. The terminal means uses a built-in microphone and voice recognition device to convert the user's voice input into text data and transmits that data to the server means. The server means stores the received data in a local database. Through this procedure, the user's profile information is managed.
[2068] Next, when the user asks, "What's the weather like today?", the terminal means uses a voice recognition device to convert the question into text data and sends it to the server means. The server means acquires the user's location information and sends a request to a weather information service (e.g., OpenWeatherMap API) to acquire the latest weather information. The acquired weather information is analyzed, converted into a format that is easy for the user to understand, and sent to the terminal means. Finally, the terminal means uses a voice synthesis device to convey the text data to the user, thereby providing the weather information.
[2069] The system also works with healthcare devices (e.g., Fitbit or Apple Watch) to acquire data on the user's exercise volume. When the user asks, "Tell me how much exercise I did today," the terminal device sends the instruction to the server device, which then acquires the exercise volume data from the healthcare device. Based on the data, the system generates health advice and communicates it to the user via the terminal device.
[2070] Furthermore, it also has a function to provide product inventory information based on the user's purchase history. When a user asks, "How much shampoo is in stock?", the server means checks the purchase history in the database and searches for inventory information. If the stock is low, it suggests replenishing the product, and when the user gives the instruction to "purchase," it places an order with the affiliated e-commerce platform and notifies the terminal means of confirmation of the order completion.
[2071] Finally, an emotion recognition device is used to analyze the user's emotional state and provide appropriate feedback in real time. If the user feels stressed, the system will suggest, "Would you like me to play some music to help you relax?" This allows for more personalized responses.
[2072] (Example of a prompt)
[2073] "Please explain the process for registering profile information. The information to enter is name, age, and gender."
[2074] "Please explain the process for providing weather information based on your current location."
[2075] "Please explain the process for capturing exercise data and providing health advice to users."
[2076] "Please explain the process of suggesting products based on a user's purchase history and placing an order."
[2077] "Please explain your process for recognizing user emotions and providing appropriate feedback."
[2078] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2079] User profile registration
[2080] Step 1:
[2081] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[2082] Input: Start the system in initial setting mode
[2083] Output: Voice message "Nice to meet you! Please register your profile."
[2084] Specific operation: The device uses a voice synthesizer to call out to the user in a familiar voice.
[2085] Step 2:
[2086] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[2087] Input: Voice data such as name, age, gender, lifestyle patterns, etc.
[2088] Output: User's voice input
[2089] Specific operation: The user speaks, "My name is Yamada Taro. I am 30 years old. I am male."
[2090] Step 3:
[2091] The device uses Google Cloud Speech-to-Text to convert input voice information into text data.
[2092] Input: User voice input
[2093] Output: User profile information converted to text data
[2094] Specific operation: The device analyzes the voice data, converts it into text, and temporarily stores it in its internal memory.
[2095] Step 4:
[2096] The terminal transmits the converted text data to the server.
[2097] Input: User profile information converted to text data
[2098] Output: Send data to the server
[2099] Specific operation: The device uses Wi-Fi or 4G / 5G communication to send text data to the server.
[2100] Step 5:
[2101] The server stores the received profile information in a MySQL database.
[2102] Input: User profile information converted to text data
[2103] Output: Message that data has been saved to the server
[2104] Specific operation: The server analyzes the text data sent and stores it in the corresponding field in the database.
[2105] Weather information provided
[2106] Step 1:
[2107] The user asks the device, "What's the weather like today?"
[2108] Input: Say "What's the weather like today?"
[2109] Output: User's voice input
[2110] Specific operation: The user asks a question into the device, and the device's microphone picks up the voice.
[2111] Step 2:
[2112] The device uses voice recognition to convert the question into text data and send it to the server.
[2113] Input: Say "What's the weather like today?"
[2114] Output: Question content converted into text data
[2115] Specific operation: The device converts the voice into text data and sends it to the server using Wi-Fi or 4G / 5G communication.
[2116] Step 3:
[2117] The server obtains the user's location information and sends a request to a weather information API (e.g., OpenWeatherMap).
[2118] Input: Question content converted into text data
[2119] Output: Request sent to the weather API
[2120] Specific operation: The server retrieves the user's location information from a database or GPS information, and uses that location information to send a request to the weather information API.
[2121] Step 4:
[2122] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[2123] Input: Weather information returned from the API (JSON format)
[2124] Output: Weather information text formatted for user
[2125] Specific operation: The server analyzes weather information and generates text data such as "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[2126] Step 5:
[2127] The server returns the generated text data to the terminal.
[2128] Input: Weather information text formatted for the user
[2129] Output: Send text data to the terminal
[2130] Specific operation: The server sends the generated text data to the terminal.
[2131] Step 6:
[2132] The terminal synthesizes the received text data into voice and conveys it to the user.
[2133] Input: Weather information converted into text data
[2134] Output: Voiced weather information
[2135] Specific operation: The device uses Google Text-to-Speech to convert the text data into speech and tells the user, "Today is sunny. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[2136] Providing healthcare advice
[2137] Step 1:
[2138] The user instructs the terminal, "Tell me how much exercise I did today."
[2139] Input: Say "How much exercise did I do today?"
[2140] Output: User's voice input
[2141] Specific operation: The user speaks to the device, "Tell me how much exercise I did today."
[2142] Step 2:
[2143] The device uses voice recognition to convert instructions into text data and send it to the server.
[2144] Input: Say "How much exercise did I do today?"
[2145] Output: User instructions converted to text data
[2146] Specific operation: The device converts the voice into text data and sends it to the server via Wi-Fi or 4G / 5G.
[2147] Step 3:
[2148] The server obtains data from healthcare devices through an API.
[2149] Input: User instructions converted into text data
[2150] Output: Sending a request to the API and retrieving exercise data
[2151] Specific operation: The server obtains the user's exercise data using the API of a healthcare device (e.g., Fitbit).
[2152] Step 4:
[2153] The server analyzes the acquired exercise data and generates advice for the user.
[2154] Input: Exercise data obtained from healthcare devices
[2155] Output: Text data generation of advice
[2156] Specific operation: The server analyzes the exercise data and generates advice text such as, "Today's exercise amount is 5,000 steps. You are halfway to your goal of 10,000 steps."
[2157] Step 5:
[2158] The server returns the generated advice text to the terminal.
[2159] Input: Text data of advice
[2160] Output: Send advice to terminal
[2161] Specific operation: The server sends the generated text data to the terminal.
[2162] Step 6:
[2163] The terminal converts the text data into speech and conveys it to the user.
[2164] Input: Text data of advice
[2165] Output: Spoken advice
[2166] Specific operation: The device uses Google Text-to-Speech to convert the text into audio and tells the user, "You've taken 5,000 steps today. You're halfway to your goal of 10,000 steps."
[2167] Product proposals and orders
[2168] Step 1:
[2169] The user asks the terminal, "How much shampoo is in stock?"
[2170] Input: Speak "What shampoo is in stock?"
[2171] Output: User's voice input
[2172] Specific operation: The user speaks to the terminal, "How much shampoo is in stock?"
[2173] Step 2:
[2174] The device uses voice recognition to convert the question into text data.
[2175] Input: Speak "What shampoo is in stock?"
[2176] Output: Question content converted into text data
[2177] Specific operation: The device converts the voice into text data.
[2178] Step 3:
[2179] The terminal transmits the text data to the server.
[2180] Input: Question content converted into text data
[2181] Output: Send data to the server
[2182] Specific operation: The device sends text data to the server using Wi-Fi or 4G / 5G.
[2183] Step 4:
[2184] The server checks the user's purchasing history in a database and searches for inventory information.
[2185] Input: Question content converted into text data
[2186] Output: Stock information search results
[2187] Specific operation: The server retrieves the user's purchase history from the database and searches for inventory information.
[2188] Step 5:
[2189] The server generates text data suggesting replenishment when stock is low or uncertain.
[2190] Input: Search results for inventory information
[2191] Output: Text data of inventory replenishment proposal
[2192] Specific operation: The server generates text data such as "Stock is low. Would you like to purchase?"
[2193] Step 6:
[2194] The server returns the generated text data to the terminal.
[2195] Input: Text data of inventory replenishment proposal
[2196] Output: Sending data to the terminal
[2197] Specific operation: The server sends the generated text data to the terminal.
[2198] Step 7:
[2199] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[2200] Input: Text data of inventory replenishment proposal
[2201] Output: Vocalized refill suggestions
[2202] Specific operation: The device uses Google Text-to-Speech to communicate with the user aloud.
[2203] Step 8:
[2204] The user instructs the terminal to "purchase."
[2205] Input: Say "Buy"
[2206] Output: User's voice input
[2207] Specific action: The user says "purchase."
[2208] Step 9:
[2209] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[2210] Input: Say "Buy"
[2211] Output: User instructions converted to text data
[2212] Specific operation: The device converts the user's voice into text data and sends it to the server.
[2213] Step 10:
[2214] The server receives the user's instructions and issues an order through the API of the affiliated online store.
[2215] Input: User instructions converted into text data
[2216] Output: Place order to online store
[2217] Specific operation: The server calls the online store's API and places an order.
[2218] Step 11:
[2219] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[2220] Input: Order completion notification from online store
[2221] Output: Order completion notification to terminal
[2222] Specific operation: The server receives confirmation of the order completion and notifies the terminal.
[2223] Step 12:
[2224] The terminal will report to the user via voice, "Your shampoo order has been completed."
[2225] Input: Notification of order completion
[2226] Output: Vocalized report
[2227] Specific operation: The device uses Google Text-to-Speech to report the completion of the order via voice.
[2228] Implementing the Emotion Engine
[2229] Step 1:
[2230] The terminal acquires emotional data from voice and facial expressions during a conversation with the user.
[2231] Input: User's voice and facial expression data
[2232] Output: Emotion data
[2233] Specific operation: The device uses an emotion recognition device to collect the user's voice and facial expression data.
[2234] Step 2:
[2235] The device uses an emotion recognition device to analyze the acquired emotion data and generate the emotional state as text data.
[2236] Input: User's voice and facial expression data
[2237] Output: Emotional state converted into text data
[2238] Specific operation: The device analyzes the emotional data and generates text data that indicates the emotional state, such as "The user is feeling stressed."
[2239] Step 3:
[2240] The terminal transmits the generated text data of the emotional state to the server.
[2241] Input: Emotional state expressed as text data
[2242] Output: Send data to the server
[2243] Specific operation: The emotion data generated by the device is sent to the server.
[2244] Step 4:
[2245] The server generates appropriate feedback and advice for the user based on the received emotional data.
[2246] Input: Emotional state expressed as text data
[2247] Output: Text data of feedback and advice
[2248] Specific behavior: The server uses the emotional data to generate feedback and advice such as "Shall I play some music to help you relax?"
[2249] Step 5:
[2250] The server returns the generated feedback and advice text data to the terminal.
[2251] Input: Text data for feedback and advice
[2252] Output: Sending data to the terminal
[2253] Specific operation: The server sends the generated feedback to the device.
[2254] Step 6:
[2255] The device synthesizes feedback and advice into voice and conveys it to the user.
[2256] Input: Text data for feedback and advice
[2257] Output: Spoken feedback and advice
[2258] What it does: Your device will use Google Text-to-Speech to suggest aloud, "Would you like me to play some music to help you relax?"
[2259] The above is a specific description of the processing of the system.
[2260] (Application example 2)
[2261] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2262] Conventional systems provide limited information and services to users, and lack personalized support that takes emotion recognition into account. Furthermore, providing real-time, emotion-aware support is difficult for customer service in brick-and-mortar stores. The present invention aims to solve these problems and realize a system that provides interactive services tailored to the user's emotional state.
[2263] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user; communication means for communicating with the server means and transmitting profile information entered by the user to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; voice output means for synthesizing the acquired weather information and providing it to the user; emotion recognition means for acquiring and analyzing user emotion data and generating appropriate feedback and advice; and voice output means for synthesizing the feedback and advice generated by the emotion recognition means and providing it to the user. This enables the provision of highly personalized information and service support based on the user's emotional state.
[2264] The "server means" is a central system that manages user profile information and emotional data, collects and analyzes external information, and generates appropriate feedback.
[2265] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and that provides a voice interface with the user.
[2266] The "communication means" is a means for transmitting and receiving data between the server means and the terminal means.
[2267] A "location information device" is a device for acquiring the current location of a user.
[2268] The "weather information acquisition means" is a means for acquiring weather information from an external weather information providing service based on the acquired location information.
[2269] The "voice output means" is a means for synthesizing acquired or generated information into voice and conveying it to the user.
[2270] The "emotion recognition means" is a means for acquiring and analyzing the user's emotional data to generate appropriate feedback and advice.
[2271] "Feedback" refers to responses or advice to the user that are generated based on the user's emotional data and other information.
[2272] "Personalization" refers to providing information and services that are customized according to a user's individual profile and emotional state.
[2273] The system for implementing this invention provides personalized information based on a user's profile information and emotion recognition. Specifically, it is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and an emotion recognition means.
[2274] First, the server manages user profile information and collects and analyzes external information. The server includes a database system for storing profile information and emotional data. Furthermore, the server analyzes data using Natural Language Processing (NLP) APIs and speech recognition APIs from Azure and Google Cloud Services. Based on the results of this analysis, the server generates appropriate feedback and advice for the user.
[2275] The terminal is a device that provides a voice interface with the user and is equipped with a microphone, a voice recognition device, and a voice synthesis device. When the user inputs a question or instruction by voice, it is converted into text data and sent to the server. The converted text data is analyzed by the server and appropriate feedback is generated. The generated feedback is sent back to the terminal and conveyed to the user via voice output means.
[2276] The communication means is used to transmit and receive data between the server and the terminal. This communication is performed in real time, and the information required by the user is provided promptly.
[2277] The location information device is used to acquire the user's current location, and based on the acquired location information, the weather information acquisition means acquires weather information from an external weather information service, allowing the user to know the weather information for their current location in real time.
[2278] The emotion recognition means acquires and analyzes the user's emotion data to generate appropriate feedback and advice. This emotion recognition is used to understand the user's emotional state from voice input and text data and provide the optimal response to the user. For example, if the user is feeling stressed, it is possible to suggest relaxing music.
[2279] The following are specific examples of embodiments:
[2280] When a user speaks into their smartphone, "What's today's promotion?", the speech is converted into text data and sent to the server. The server analyzes the text data and generates promotion information, taking into account the user's emotional state. The device then synthesizes the generated promotion information into voice and provides it to the user.
[2281] Examples of such prompts include "Determine whether the user is interested in sale days and provide appropriate promotional information" and "Determine the emotion of the following text:- Text: {user's voice input}- Emotion: happy, sad, angry, neutral."
[2282] This allows users to get the information they need in real time within a physical store and receive personalized support based on their emotional state.
[2283] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2284] Step 1:
[2285] The user speaks a question or command into the device. For example, they might say, "What are today's promotions?" The voice input is picked up by the device's microphone.
[2286] Step 2:
[2287] The terminal analyzes the voice input with a voice recognition device and converts it into text data. After this text data is generated, it is sent to the server. The input is voice data and the output is text data.
[2288] Step 3:
[2289] The server analyzes the received text data using a Natural Language Processing (NLP) API. During the analysis, it also extracts the question content and emotional part. This analysis reveals the user's question content and emotional state. The input is text data, and the output is the analysis result.
[2290] Step 4:
[2291] The server generates an appropriate response based on the analysis results and profile information. For example, if the question is about a promotion, it generates promotion information taking into account inventory information and cross-selling information. The response also changes depending on the emotional state. The input is the analysis results and profile data, and the output is the response data.
[2292] Step 5:
[2293] The server then personalizes the generated response data according to the user's emotional state and sends the response in text format to the terminal. The input is the response data, and the output is the personalized text data.
[2294] Step 6:
[2295] The terminal synthesizes the received personalized text data using a voice output means and conveys the speech to the user. For example, it might say, "Today is sale day! Your favorite product is 20% off!" The input is personalized text data, and the output is voice data.
[2296] Step 7:
[2297] After receiving a response from the device, the user acts based on the content. For example, a user receives information about a sale day and decides which specific product to purchase. This step does not provide any feedback to the system, but it may affect future interactions. The input is voice data, and the output is the user's actions.
[2298] This allows users to receive personalized information and responses in real time that are tailored to their emotions.
[2299] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2300] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2301] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2302] [Fourth embodiment]
[2303] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2304] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2305] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2306] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2307] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2308] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2309] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2310] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2311] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2312] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2313] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2314] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2315] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2316] The present invention is a system that manages profile information about users and provides a variety of information and support, and is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, an audio output means, and a healthcare device.
[2317] User profile registration
[2318] (overview)
[2319] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[2320] (Processing flow)
[2321] 1. The device enters initial setup mode and prompts the user, "Hello! Please register your profile."
[2322] 2. The user follows the instructions on the device and enters their information (name, age, gender, lifestyle patterns, etc.) by voice.
[2323] 3. The device uses voice recognition to convert the input information into text data.
[2324] 4. The device sends the converted text data to the server.
[2325] 5. The server stores the received profile information in a database.
[2326] Weather information provided
[2327] (overview)
[2328] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[2329] (Processing flow)
[2330] 1. The user asks the device, "What's the weather like today?"
[2331] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[2332] 3. The server obtains the user's location information and sends a request to the weather information API.
[2333] 4. The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[2334] 5. The server returns the generated text data to the terminal.
[2335] 6. The device synthesizes the received text data into voice and tells the user, for example, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[2336] Providing healthcare advice
[2337] (overview)
[2338] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[2339] (Processing flow)
[2340] 1. The user instructs the device, "Tell me how much exercise I did today."
[2341] 2. The device uses voice recognition to convert the instructions into text data and send it to the server.
[2342] 3. The server obtains data from the healthcare device through the API.
[2343] 4. The server analyzes the acquired exercise data and generates advice for the user.
[2344] 5. The server returns the generated advice text to the terminal.
[2345] 6. The device converts the text data into speech and tells the user, "Today's exercise volume was 5,000 steps. You're halfway to your goal of 10,000 steps."
[2346] Product proposals and orders
[2347] (overview)
[2348] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[2349] (Processing flow)
[2350] 1. The user asks the terminal, "How much shampoo is in stock?"
[2351] 2. The device uses voice recognition to convert the question into text data and send it to the server.
[2352] 3. The server checks the user's purchasing history in the database and searches for inventory information.
[2353] 4. The server generates text data suggesting replenishment if stock is low or uncertain.
[2354] 5. The server returns the generated text data to the terminal.
[2355] 6. The device converts the text data into speech and tells the user, "Your shampoo is low in stock. Would you like to buy more?"
[2356] 7. The user instructs the device to "purchase."
[2357] 8. The device uses voice recognition to convert the instructions into text data and send it to the server.
[2358] 9. The server receives the user's instructions and issues an order to the affiliated online store via API.
[2359] 10. The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[2360] 11. The terminal will report to the user by voice, "Your shampoo order has been completed."
[2361] In this way, the system of the present invention enriches the user's life through a series of operations and provides multifunctional support with a user-friendly interface.
[2362] The processing flow will be explained below.
[2363] User profile registration
[2364] Step 1:
[2365] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[2366] Step 2:
[2367] The user follows the instructions on the device and inputs their own information (name, age, gender, lifestyle patterns, etc.) by voice.
[2368] Step 3:
[2369] The device uses voice recognition to convert the input information into text data.
[2370] Step 4:
[2371] The terminal transmits the converted text data to the server.
[2372] Step 5:
[2373] The server stores the received profile information in a database.
[2374] Weather information provided
[2375] Step 1:
[2376] The user asks the device, "What's the weather like today?"
[2377] Step 2:
[2378] The device uses voice recognition to convert the question into text data.
[2379] Step 3:
[2380] The terminal transmits the text data to the server.
[2381] Step 4:
[2382] The server obtains the user's location information and sends a request to the weather information API.
[2383] Step 5:
[2384] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[2385] Step 6:
[2386] The server returns the generated text data to the terminal.
[2387] Step 7:
[2388] The device synthesizes the received text data into voice and tells the user, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees."
[2389] Providing healthcare advice
[2390] Step 1:
[2391] The user instructs the terminal, "Tell me how much exercise I did today."
[2392] Step 2:
[2393] The device uses voice recognition to convert instructions into text data.
[2394] Step 3:
[2395] The terminal transmits the text data to the server.
[2396] Step 4:
[2397] The server obtains data from healthcare devices through an API.
[2398] Step 5:
[2399] The server analyzes the acquired exercise data and generates advice for the user.
[2400] Step 6:
[2401] The server returns the generated advice text to the terminal.
[2402] Step 7:
[2403] The device converts the text data into speech and tells the user, "Today you've taken 5,000 steps. You're halfway to your goal of 10,000 steps."
[2404] Product proposals and orders
[2405] Step 1:
[2406] The user asks the terminal, "How much shampoo is in stock?"
[2407] Step 2:
[2408] The device uses voice recognition to convert questions into text data.
[2409] Step 3:
[2410] The terminal transmits the text data to the server.
[2411] Step 4:
[2412] The server checks the user's purchasing history in a database and searches for inventory information.
[2413] Step 5:
[2414] The server generates text data suggesting replenishment when stock is low or uncertain.
[2415] Step 6:
[2416] The server returns the generated text data to the terminal.
[2417] Step 7:
[2418] The device converts the text data into speech and tells the user, "Your shampoo stock is low. Would you like to buy some?"
[2419] Step 8:
[2420] The user instructs the terminal to "purchase."
[2421] Step 9:
[2422] The terminal uses voice recognition to convert the instructions into text data and send it to the server.
[2423] Step 10:
[2424] The server receives the user's instructions and issues an order to the affiliated online store via API.
[2425] Step 11:
[2426] The server receives confirmation that the order has been completed and notifies the terminal that the order has been completed.
[2427] Step 12:
[2428] The terminal will report to the user via voice, "Your shampoo order has been completed."
[2429] Example 1
[2430] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2431] Conventional voice interface systems have difficulty effectively managing user profile information, location information, physical activity data, and purchase history, and providing a variety of information and support. Furthermore, the complexity of the system when integrating multiple functions and improving the user experience have also been issues.
[2432] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2433] In this invention, the server includes: server means for managing profile information about users; terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with users; communication means for communicating with the server means and transmitting the profile information input by the users to the server means; a location information device for acquiring user location information; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; and voice output means for synthesizing the acquired weather information and providing it to the user. This makes it possible to integrate and manage the user's profile information, location information, exercise data, and purchase history, and to provide information and support efficiently and effectively.
[2434] The "server means" is a device or system for managing profile information about users and processing and storing data in cooperation with various databases.
[2435] The "terminal means" is a device or system that includes a microphone, a voice recognition device, and a voice synthesis device, provides a voice interface with the user, and collects input data from the user.
[2436] "Communication means" refers to a device or system for transmitting and receiving data between terminal means and server means.
[2437] A "location information device" is a device for acquiring a user's current location, and mainly uses location information technology such as GPS.
[2438] The "weather information acquisition means" is a device or system for acquiring the latest weather information from an external weather information providing means based on the acquired user's location information.
[2439] The "audio output means" is a device or system that converts acquired weather information and various data into audio and provides it to the user.
[2440] A "healthcare device" is a device for acquiring a user's exercise data, and primarily includes fitness trackers and smartwatches.
[2441] The "online store communication means" is a device or system for issuing an order to the online store based on the user's purchase instructions.
[2442] MODE FOR CARRYING OUT THE INVENTION
[2443] The present invention is a system that manages a user's profile information, location information, exercise data, and purchase history, and provides a variety of information and support based on this data. The system is mainly composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, and a healthcare device.
[2444] User profile registration
[2445] overview
[2446] This system collects user profile information during initial setup and provides individual services based on that information. Users can register their profile by voice via a terminal.
[2447] Detailed Description
[2448] The device enters initial setup mode and starts a voice prompt saying, "Nice to meet you! Please register your profile." The user uses the microphone to input information such as their name, age, gender, and lifestyle patterns by voice. The device uses a voice recognition engine (for example, Google Voice Recognition API) to convert the voice into text data. The converted text data is sent to the server using a communication method. The server stores the received data in a database.
[2449] Specific examples
[2450] For example, if a user speaks "My name is Yamada Taro and I'm 30 years old," the device converts it into text and the server stores it in a database. An example of a prompt is "Write a program that inputs a user's name and age and stores that information in a database."
[2451] Weather information provided
[2452] overview
[2453] The system provides users with up-to-date weather information based on their current location, allowing them to easily check the weather before heading out.
[2454] Detailed Description
[2455] When a user asks a device, "What's the weather like today?", the device uses its voice recognition function to convert the question into text data and sends it to the server. The server uses a location information device to obtain the user's current location and sends a request to a weather information providing API (for example, OpenWeatherMap API). The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand. The generated text data is sent back to the device and conveyed to the user using the voice synthesis function.
[2456] Specific examples
[2457] When the user asks, "What's the weather like today?", the device responds, "It's sunny today. The maximum temperature is 25 degrees and the minimum temperature is 15 degrees." An example of a prompt is, "Write a program that obtains weather information based on the user's location and relays it to them in voice."
[2458] Providing healthcare advice
[2459] overview
[2460] The system acquires the user's exercise data from a healthcare device and provides health advice based on the data.
[2461] Detailed Description
[2462] When a user instructs the device to "tell me how much exercise I did today," the device uses voice recognition to convert the instruction into text data and send it to the server. The server acquires the exercise data via the healthcare device, analyzes it, and generates advice. The advice is then sent back to the device, converted into voice, and conveyed to the user.
[2463] Specific examples
[2464] When the user asks, "How much exercise did you do today?", the device responds, "You've done 5,000 steps today. You're halfway to your goal of 10,000 steps." An example of a prompt is, "Write a program that obtains a user's exercise data and provides advice based on that data."
[2465] Product proposals and orders
[2466] overview
[2467] This system provides product inventory information based on the user's purchasing history, helping the user easily purchase the products they need.
[2468] Detailed Description
[2469] When a user asks the terminal, "What shampoo is in stock?", the terminal uses voice recognition to convert the question into text data and sends it to the server. The server checks the purchase history database to find inventory information. If inventory is low or unclear, the server generates a replenishment suggestion and sends it back to the terminal. The terminal notifies the user of this by voice, receives the user's purchase instructions, and issues an order.
[2470] Specific examples
[2471] For example, if a user asks, "How much shampoo is in stock?" and the terminal responds, "Your shampoo is low in stock. Would you like to buy more?", and the user responds by saying, "Purchase," the terminal reports, "Your shampoo order is complete." An example of a prompt sentence is, "Write a program that provides inventory information based on the user's purchasing history and places an order for the product if necessary."
[2472] In this way, the system of the present invention provides multifunctional services through detailed and specific processes to support the user's life.
[2473] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2474] User profile registration
[2475] Step 1:
[2476] The device will enter initial setup mode and prompt the user, "Hello! Please register your profile."
[2477] Specific behavior: The device plays a message aloud, prompting the user to prepare to start typing.
[2478] Input: Command to start the initial setting mode on the terminal side
[2479] Output: Start of voice prompt
[2480] Step 2:
[2481] The user follows the instructions on the device and enters their information by voice, for example, saying, "My name is Taro Yamada and I'm 30 years old."
[2482] Specific operation: The user verbally transmits their profile information through a microphone.
[2483] Input: Voice input (user profile information)
[2484] Output: Audio data
[2485] Step 3:
[2486] The device uses a voice recognition engine to convert the input voice data into text data, using the Google Voice Recognition API.
[2487] Specific operation: Converts voice data into text in real time.
[2488] Input: Audio data
[2489] Output: Text data
[2490] Step 4:
[2491] The device sends the converted text data to the server via a communication method, using an AWS S3 bucket.
[2492] Specific operation: Text data is transferred to the server using a communication protocol.
[2493] Input: Text data
[2494] Output: Send data to the server
[2495] Step 5:
[2496] The server stores the received profile information in a database, using a MySQL database.
[2497] Specific behavior: Saves and validates data.
[2498] Input: Text data
[2499] Output: Save information to a database
[2500] Weather information provided
[2501] Step 1:
[2502] The user asks the device, "What's the weather like today?" and inputs the request by voice.
[2503] Specific operation: The user verbally instructs the device that he or she wants to know weather information.
[2504] Input: Voice input (weather information request)
[2505] Output: Audio data
[2506] Step 2:
[2507] The device uses its voice recognition function to convert the question into text data and send it to the server, using the Google Voice Recognition API.
[2508] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[2509] Input: Audio data
[2510] Output: Text data (weather information request)
[2511] Step 3:
[2512] The server uses the location device to obtain the user's current location and sends a request to the OpenWeatherMap API.
[2513] Specific operation: Obtains location information and sends a request to an external API.
[2514] Input: Text data, location information
[2515] Output: Weather information request
[2516] Step 4:
[2517] The server analyzes the weather information returned from the API and generates text data in a format that is easy for the user to understand.
[2518] What it does: Parses weather information and formats it for the user.
[2519] Input: Weather information data
[2520] Output: Text data (analyzed weather information)
[2521] Step 5:
[2522] The server returns the generated text data to the terminal, which converts it into voice and conveys it to the user.
[2523] Specific operation: Text data is synthesized into speech and information is provided to the user.
[2524] Input: Text data (analyzed weather information)
[2525] Output: Audio output (weather information)
[2526] Providing healthcare advice
[2527] Step 1:
[2528] The user instructs the device, "Tell me how much exercise I did today." The request is input by voice.
[2529] Specific operation: The user gives voice instructions to the terminal that he / she wants to know the exercise amount data.
[2530] Input: Voice input (request for exercise data)
[2531] Output: Audio data
[2532] Step 2:
[2533] The device uses voice recognition to convert instructions into text data and send it to the server.
[2534] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[2535] Input: Audio data
[2536] Output: Text data (momentum data request)
[2537] Step 3:
[2538] The server obtains exercise data via the healthcare device through the FitBit API.
[2539] Specific behavior: Uses an external API to obtain momentum data.
[2540] Input: Request data
[2541] Output: Momentum data
[2542] Step 4:
[2543] The server analyzes the acquired exercise data and generates advice for the user.
[2544] Specific operation: Analyze the momentum data and generate advice as text data.
[2545] Input: Momentum data
[2546] Output: Text data (advice)
[2547] Step 5:
[2548] The server returns the generated text data of the advice to the terminal, which converts it into voice and conveys it to the user.
[2549] Specific operation: Text data is synthesized into speech and information is provided to the user.
[2550] Input: Text data (advice)
[2551] Output: Audio output (advice)
[2552] Product proposals and orders
[2553] Step 1:
[2554] The user asks the terminal, "What shampoo is in stock?" and inputs the request by voice.
[2555] Specific operation: The user verbally instructs the terminal that he / she wants to know inventory information.
[2556] Input: Voice input (request for stock information)
[2557] Output: Audio data
[2558] Step 2:
[2559] The device uses voice recognition to convert the question into text data and send it to the server.
[2560] Specific operation: Converts voice data into text and sends it to the server via a communication protocol.
[2561] Input: Audio data
[2562] Output: Text data (request for inventory information)
[2563] Step 3:
[2564] The server checks the user's purchasing history in a database and searches for inventory information.
[2565] Specific operation: Retrieve purchase history from the database and check inventory information.
[2566] Input: Text data (request for inventory information)
[2567] Output: Inventory information
[2568] Step 4:
[2569] The server generates replenishment suggestions when inventory is low or uncertain.
[2570] Specific operation: Generate replenishment suggestions as text data based on inventory information.
[2571] Input: Inventory information
[2572] Output: Text data (supplement proposal)
[2573] Step 5:
[2574] The server returns the generated text data of the proposal to the terminal, which converts it into voice and conveys it to the user.
[2575] Specific operation: Text data is synthesized into speech and information is provided to the user.
[2576] Input: Text data (supplement proposal)
[2577] Output: Audio output (supplement suggestion)
[2578] Step 6:
[2579] The user instructs the terminal to "purchase," and the terminal converts the instruction into text using voice recognition and sends it to the server.
[2580] Specific operation: Converts voice input into text data and sends it to the server.
[2581] Input: Voice input (purchase instructions)
[2582] Output: Text data (purchase instructions)
[2583] Step 7:
[2584] The server receives the user's instructions and places an order with the affiliated online store.
[2585] What it does: Places an order using the online store's API.
[2586] Input: Text data (purchase instructions)
[2587] Output: Order data
[2588] Step 8:
[2589] The server receives confirmation that the order has been completed and sends a notification to the terminal, which then tells the user, "Your shampoo order has been completed."
[2590] Specific operation: The notification data is synthesized into voice and the information is provided to the user.
[2591] Input: Order completion notification
[2592] Output: Audio output (order completion notification)
[2593] In this way, by performing specific input, data processing, and output at each step, multifunctional services are realized for users.
[2594] (Application example 1)
[2595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2596] In modern society, users lead busy lives and require timely recommendations for optimal products and services based on weather information and health status to facilitate smooth purchasing behavior. It is also necessary to improve the quality of life by managing this information in an integrated manner and providing appropriate advice to users. However, conventional systems have had difficulty integrating various data, such as user profiles, location information, weather information, purchase history, and health data, to provide consistent services to users.
[2597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2598] In this invention, the server includes server means for managing profile information about users, terminal means having a microphone, a voice recognition device, and a voice synthesis device and providing a voice interface with the user, communication means for communicating with the server means and transmitting the profile information entered by the user to the server means, a location information device for acquiring user location information, weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information, voice output means for synthesizing the acquired weather information and providing it to the user, proposal generation means for proposing optimal products and services based on the acquired weather information and user preference information, and health support means for proposing optimal health foods and lifestyle-related products based on the user's purchase history and health data. This allows users to obtain information necessary for their daily lives in a centralized manner and improve their quality of life.
[2599] The "server means" is a central computer device that manages user profile information, purchase history, health data, etc., and integrates various data to provide optimal services to users.
[2600] The "terminal means" is a device that includes a microphone, a voice recognition device, and a voice synthesis device, and provides a voice interface with the user.
[2601] The "communication means" is an interface for the terminal means to send and receive data to and from the server means.
[2602] A "location information device" is a device for obtaining the current location of a user.
[2603] The "weather information acquisition means" is a means for acquiring the latest weather information from an external weather information providing service based on the user's location information.
[2604] The "audio output means" is a device that synthesizes acquired information into voice and provides it to the user in the form of voice.
[2605] The "proposal generating means" is a means for proposing optimal products and services based on the acquired weather information and user preference information.
[2606] "Health support tools" are tools for suggesting optimal health foods and lifestyle-related products based on a user's purchasing history and health data.
[2607] This invention is a system that provides various information and support based on user profile information. The system is composed of a server means, a terminal means, a communication means, a location information device, a weather information acquisition means, a voice output means, a suggestion generation means, and a health support means.
[2608] Hardware and software used
[2609] 1. Server means: A central computer device (e.g., a cloud server) that manages user profile information, purchase history, and health data.
[2610] 2. Terminal means: A device equipped with a microphone, a voice recognition device, and a voice synthesis device (examples include smartphones, smart glasses, and head-mounted displays).
[2611] 3. Communication means: An interface that transmits and receives data between terminal means and server means (example: Internet communication).
[2612] 4. Location information device: A device that obtains the user's current location (example: GPS).
[2613] 5. Weather information acquisition method: A method for acquiring the latest weather information from an external weather information service (e.g., weather API).
[2614] 6. Voice output means: A device that provides information to the user by voice synthesis (example: speaker).
[2615] 7. Proposal generation means: A means for proposing optimal products and services based on weather information and user preference information (specific example: proposal generation algorithm).
[2616] 8. Health support tools: Tools that suggest optimal health foods and lifestyle-related products based on a user's purchasing history and health data (example: health advice algorithm).
[2617] Data processing and calculation
[2618] Registration of profile information: The user enters his / her profile information by voice through the terminal means, and the voice recognition device converts the information into text data and transmits it to the server means.
[2619] Obtaining weather information: Based on the user's location information obtained by the location information device, the weather information providing means obtains the latest weather information, and the audio output means provides the information by audio.
[2620] Proposal generation: The proposal generator integrates weather information and user preference information to suggest the most suitable products and services to the user.
[2621] Health support: Health support tools analyze users' purchasing history and health data to recommend optimal health foods and lifestyle-related products.
[2622] Examples of concrete examples and prompts
[2623] Example 1: "When a user asks their device, 'What's the weather like today?' the system obtains their location, retrieves weather information from a weather API, and provides, via voice synthesis, the answer, 'It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees.'"
[2624] Example prompt: "Suggest restaurants that deliver to a given user based on their location and weather."
[2625] Example 2: "When a user asks the device, 'What do you recommend for lunch today?' the system will suggest the best restaurant and dish based on weather information and the user's preferences."
[2626] Example prompt: "Based on the user's healthcare data, please recommend a healthy meal for them."
[2627] In this way, the system integrates multiple data sets, provides users with the information and services they need in a unified manner, and improves their quality of life.
[2628] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2629] System processing flow
[2630] Step 1:
[2631] The user inputs voice into the terminal means. The user asks questions such as "What's the weather like today?" or "What do you recommend for lunch today?" This voice input is captured by the microphone in the terminal means.
[2632] Step 2:
[2633] The terminal means converts the user's voice input into text data using a voice recognition device, and the converted text data is sent to the server means for further processing, where the voice data is processed as text data.
[2634] Step 3:
[2635] The server analyzes the received text data. For example, it analyzes the text "Tell me the weather today" and recognizes that the user's intention is to obtain weather information. Based on the analysis result, it obtains the user's current location from the location information device.
[2636] Step 4:
[2637] The server means uses the acquired location information to send a request to the weather information API, which acquires the latest weather information from an external weather information service based on the location information.
[2638] Step 5:
[2639] The server means receives and analyzes the weather information returned from the weather information providing API. Based on the analysis results, the server means generates text data for a voice response. For example, the generated text data is "It's sunny today. The maximum temperature is 25 degrees, and the minimum temperature is 15 degrees."
[2640] Step 6:
[2641] The generated text data is transmitted from the server means to the terminal means. Specifically, the server means transfers the text data for speech synthesis to the terminal means.
[2642] Step 7:
[2643] The terminal means converts the transmitted text data into voice using a voice synthesizer. This provides the user with a specific weath...
Claims
1. server means for managing profile information about users; a stuffed toy-type terminal means having a microphone, a voice recognition device, and a voice synthesis device, and providing a voice interface with a user; a communication means for communicating with the server means and transmitting the profile information input by the user to the server means; a location information device for acquiring location information of a user; weather information acquisition means for acquiring weather information from an external weather information providing means based on the acquired location information; a voice output means for synthesizing the acquired weather information into voice and providing it to the user; A system including:
2. The system according to claim 1, wherein the stuffed toy-type terminal means further comprises a communication means for communicating with a healthcare device to acquire the user's exercise data, and the server means generates advice based on the acquired exercise data and provides it to the user through the audio output means.
3. The system according to claim 1, wherein the server means has a database for managing user purchasing history, provides product inventory information based on the purchasing history through a stuffed toy-type terminal means, and further includes an online store communication means for issuing orders to the online store based on purchasing instructions from the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A