system
The system addresses the inefficiencies of existing action planning systems by converting voice input to text, analyzing user intent, and refining suggestions based on feedback, ensuring high-quality and intuitive activity planning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Existing systems require users to manually search for and evaluate action suggestions, which is time-consuming and often results in inconsistent quality, lacking intuitive and voice-based interfaces for planning enjoyable time with family and friends.
A system that receives voice input, converts it to text, analyzes the data to generate action suggestions, and provides feedback-based refinement, using a terminal and server with voice recognition, data analysis, and generative AI to offer personalized and efficient planning.
Enables users to quickly and intuitively plan enjoyable activities by providing high-quality suggestions that accurately reflect their intentions and preferences, enhancing user experience through a user-friendly interface.
Smart Images

Figure 2026063769000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, it is very important to improve the quality of time spent with family and friends. However, many people spend a lot of time when deciding on the next action. As a result, the enjoyable time cannot often be utilized to the maximum. Also, even if there is a system that proposes actions, there are often problems such as the lack of an intuitive and friendly interface, making it difficult for users to use easily.
Means for Solving the Problems
[0005] This invention provides a system that receives voice input from a user, converts it into text data, and analyzes it. Based on the analyzed data, it generates multiple action suggestions and presents the user with the most suitable suggestion. It also has a function to receive feedback from the user and generate and present suggestions again based on that feedback. Furthermore, by transmitting location information along with the received voice input and providing detailed information about the suggestions, the system enables the user to easily and quickly decide on their next action. This system is housed in a dedicated terminal and is designed to be used by many people by being installed in places where people gather. In addition, its user-friendly language and interface make the user's conversation natural and enjoyable, and increase their anticipation for the next action.
[0006] "Voice input" refers to instructions or information that a user gives to a system verbally.
[0007] "Text data" refers to data obtained by converting voice input into written information, and is the data that the system analyzes.
[0008] "Analyzing" refers to data processing performed in order to understand the meaning and intent behind the received data.
[0009] "Action suggestions" refer to specific action plans or destinations that the system presents to the user.
[0010] "To present" refers to the act of showing or conveying information generated by a system to a user.
[0011] "Feedback" refers to the reactions and opinions that users give to the information they are presented with.
[0012] "Location information" refers to data that indicates the user's current geographical location.
[0013] "Detailed information" refers to additional information and explanations related to the suggested action, provided to support the user's actions.
[0014] A "terminal" refers to the hardware device in which this system is housed.
[0015] "Approachable language" refers to the selection and expression of words in a format that is easy for users to use and understand. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention provides a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Specific embodiments of this system are described in detail below.
[0038] System Configuration
[0039] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[0040] Program Processing Description
[0041] 1. User input reception:
[0042] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0043] The device uses a speech recognition module to convert voice input into text data.
[0044] The device sends the converted text data and the user's location information to the server.
[0045] 2. Data analysis and proposal generation:
[0046] The server analyzes the text data it receives to understand the user's intent.
[0047] The server uses generative AI to generate multiple action suggestions, taking into account location information, season, weather information, and past suggestion history.
[0048] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[0049] 3. Presentation to the user:
[0050] The terminal converts the received suggestion into speech using a speech synthesis module.
[0051] The device presents the generated audio data to the user (e.g., "How about a drive along the Shonan coast?").
[0052] 4. Accepting feedback:
[0053] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0054] The device then uses the speech recognition module again to convert the voice feedback into text data.
[0055] The terminal sends the converted feedback to the server.
[0056] 5. Reanalysis and proposals (if necessary):
[0057] The server analyzes the feedback and generates new suggestions based on the new conditions.
[0058] The server sends a revised suggestion to the terminal, and the terminal again presents the suggestion to the user via voice (e.g., "How about Hayama Beach?").
[0059] Specific example
[0060] When a user asks the device, "Where should we go today?", the device converts the voice input into text and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations, then selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device informs the user of this. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[0061] This system is designed to allow users to make decisions about their next actions with ease, using friendly language and an intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with confidence.
[0062] The following describes the processing flow.
[0063] Step 1:
[0064] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0065] The terminal accepts voice input and activates the voice recognition module.
[0066] Step 2:
[0067] The device converts voice input into text data.
[0068] The converted text data is reviewed to determine if it is a valid request.
[0069] Step 3:
[0070] The terminal transmits the converted text data and the user's location information to the server via the communication module.
[0071] Step 4:
[0072] The server processes the received text data using a parsing module to understand the user's intent (e.g., understanding the intent "I want to go somewhere where I can see the ocean").
[0073] Step 5:
[0074] The server references the user's location information, current season, weather information, and past suggestion history, and uses generative AI to generate multiple action suggestions.
[0075] Step 6:
[0076] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[0077] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[0078] Step 7:
[0079] The server sends the selected proposal as text data to the terminal.
[0080] Step 8:
[0081] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[0082] The device generates audio data which is then presented to the user through the speaker (e.g., "How about a drive along the Shonan coast?").
[0083] Step 9:
[0084] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0085] The device restarts its speech recognition module and converts the user's feedback into text data.
[0086] Step 10:
[0087] The device sends the converted feedback back to the server.
[0088] Step 11:
[0089] The server analyzes the feedback and generates multiple action suggestions again based on the new conditions.
[0090] Select the best option from the revised proposals (e.g., choose "Hayama Beach").
[0091] Step 12:
[0092] The server sends the re-selected proposals to the terminal as text data.
[0093] Step 13:
[0094] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[0095] Through these steps, users can quickly and intuitively decide on their next course of action.
[0096] (Example 1)
[0097] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] Conventional action planning systems have the problem of requiring users to search for, evaluate, and judge information themselves, which is time-consuming. Furthermore, the quality of suggestions is often inconsistent, making it difficult to accurately reflect the user's intentions. In addition, there is a lack of voice-based interfaces, hindering the intuitive and rapid development of action plans.
[0099] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0100] In this invention, the server includes means for analyzing text data and understanding the user's intent, means for generating multiple action suggestions using generative artificial intelligence based on the analyzed data, and means for selecting the optimal action suggestion and transmitting it along with detailed information. This makes it possible to accurately reflect the user's intent and provide high-quality action suggestions quickly and intuitively.
[0101] "Means for receiving voice input" refers to a combination of hardware and software for capturing the user's voice as a digital signal.
[0102] "Means for converting voice input into text data" refers to speech recognition technology that analyzes captured voice signals and converts them into corresponding text data.
[0103] "Means for transmitting location information to a server" refers to a communication module that transmits location information, which identifies the user's current location, to a server.
[0104] "Methods for analyzing text data and understanding user intent" refers to the process of analyzing text data using natural language processing technology to grasp the intent behind user requests and questions.
[0105] "Means for generating multiple action suggestions using generative artificial intelligence" refers to an algorithm that uses artificial intelligence technology to create multiple action options based on user requests.
[0106] The "means for selecting the most suitable action proposal and sending it along with detailed information" refers to a function that selects the most appropriate action proposal from those generated and sends its detailed information (e.g., access methods and local conditions) to the terminal.
[0107] "Means for converting generated action suggestions into speech and presenting them to the user" refers to speech synthesis technology and output devices that convert text data into speech data and present suggestions to the user visually and aurally.
[0108] "A means of receiving user feedback, converting audio into text data, and sending it to a server" refers to a process that receives user feedback as voice input, converts it into text, and transmits it to a server.
[0109] "A means of generating new suggestions based on received feedback and presenting them in audio format" refers to a process of generating new action suggestions based on user feedback, converting them back into audio data, and presenting them to the user.
[0110] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[0111] System Configuration
[0112] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[0113] Hardware and software to be used
[0114] The device includes a microphone for receiving voice input, a speech recognition module (e.g., Google® Cloud Speech-to-Text API, Amazon Transcribe, etc.) for converting speech to text data, and a GPS module for obtaining the user's location information. A communication module is used to send the converted text data and location information to the server. On the server side, natural language processing technology (e.g., BERT, GPT-3®, etc.) is used as a data analysis module to understand the user's intent. Furthermore, generative AI (e.g., OpenAI® GPT-3, etc.) is used to generate multiple action suggestions. After the server selects the optimal action suggestion, it sends it to the device along with detailed information (such as how to access it and the local weather). The device uses a speech synthesis module (e.g., Google Text-to-Speech API, Amazon Polly, etc.) to convert the suggested content into speech and present it to the user.
[0115] Specific example
[0116] As a concrete example, suppose a user asks their device, "Where should we go today?" The device converts the voice input into text data and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations. From these, it selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device communicates this to the user. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[0117] As described above, this system is designed to allow users to make decisions about their next actions in an enjoyable way through its user-friendly language and intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with peace of mind.
[0118] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0119] Step 1:
[0120] User input reception
[0121] The user speaks into the device to ask a question or make a request (e.g., "Where should we go today?"). The device's microphone collects the voice data and uses it as input for a speech recognition module (such as Google Cloud Speech-to-Text API or Amazon Transcribe). The speech recognition module converts the voice data into text data, generating a text-based input. This text data includes the user's intent or inquiry.
[0122] Step 2:
[0123] Acquiring and transmitting location information
[0124] The device uses a GPS module to determine the user's current location. The acquired location information is sent to the server along with text data. Data communication is conducted using the HTTPS protocol during transmission, ensuring data security.
[0125] Step 3:
[0126] Data Analysis
[0127] The server analyzes the received text data and location information. It uses natural language processing models (e.g., BERT or GPT-3) to analyze the text data and understand the user's intent. This process extracts keywords and contextual information from the text data. Based on the analysis results, the server generates data that reflects the user's wishes and requests.
[0128] Step 4:
[0129] Generating action proposals
[0130] The server generates multiple action suggestions based on the analysis results using a generative AI (e.g., OpenAI's GPT-3). The server also takes into account external data such as the user's location, season, and weather information (e.g., using the OpenWeatherMap API). This improves the accuracy and appropriateness of the suggestions. The generative AI generates multiple candidate suggestions and adds detailed information to each suggestion.
[0131] Step 5:
[0132] Selecting and sending the most suitable proposal
[0133] The server uses evaluation criteria (e.g., distance, weather conditions, event information, etc.) to select the most suitable action from the generated suggestions. It then sends the best suggestion and its detailed information (such as access methods) to the terminal. The server uses asynchronous communication to transmit data quickly and efficiently.
[0134] Step 6:
[0135] Presentation to the user
[0136] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Google Text-to-Speech API or Amazon Polly). The generated speech data is then presented to the user through the speaker (e.g., "How about a drive to Shonan Beach?"). Simultaneously, displaying the suggestions in text format on the screen is also being considered.
[0137] Step 7:
[0138] Accepting feedback
[0139] The user provides voice feedback on the suggested location (e.g., "Is there a place a little closer?"). The device uses a voice recognition module to convert the voice feedback into text data. This text data is then sent back to the server for further analysis.
[0140] Step 8:
[0141] Generation and presentation of revised proposals
[0142] The server analyzes the feedback and generates new suggestions based on the new conditions. Generative AI is used again in this re-suggestion to create new, appropriate candidates. The server sends detailed information about the re-suggestion to the terminal, which then presents it to the user again (e.g., "How about Hayama Beach?"). This allows the system to repeat suggestions until the user is satisfied.
[0143] Through this series of processes, users can intuitively and efficiently plan their actions.
[0144] (Application Example 1)
[0145] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0146] In modern society, efficiently and intuitively planning time with family and friends is crucial for users to make the most of their limited time. However, users often have to go through the trouble of translating their desires into concrete and appropriate action suggestions, and obtaining detailed information to judge the appropriateness of those suggestions. Furthermore, the lack of means to properly make reservations or orders for the suggested actions can hinder the smooth execution of those suggestions. There is a need for a system that can solve this problem.
[0147] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0148] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data and understanding the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, and means for making reservations or orders related to the suggested actions. This allows the user to easily receive action suggestions through voice input, provide feedback to obtain the most suitable suggestion, and smoothly transition to actual actions based on the suggestions.
[0149] "Means for receiving voice input" refers to hardware or software that receives voice data and converts it into a format that can be processed within the system.
[0150] "Means for converting received voice input into text data" refers to software or hardware that uses speech recognition technology to convert a user's voice input into corresponding text data.
[0151] "Means for analyzing converted text data and understanding user intent" refers to software or hardware that uses natural language processing techniques to analyze the content of text data and understand what the user is asking for.
[0152] "Means for generating multiple action suggestions based on analyzed data" refers to an algorithm for generating multiple action options for the user based on the analysis results, and the hardware or software that executes it.
[0153] "Means for presenting generated action suggestions" refers to devices and software that present generated action options to the user, either visually or audibly.
[0154] "Means for receiving user feedback" refers to the hardware and software used to receive and incorporate user responses to suggestions into the system.
[0155] "Means for generating and presenting new suggestions based on received feedback" refers to algorithms and devices that generate and present new action suggestions based on user feedback.
[0156] "Means of making reservations or orders related to suggested actions" refers to online services for making restaurant reservations or ordering goods based on actions suggested within the system, as well as the software and hardware for executing them.
[0157] "Means of transmitting location information" refers to a GPS module and communication module used to determine the user's current location and transmit it to a server.
[0158] This invention provides a system for users to plan meals efficiently and intuitively. The system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[0159] System Configuration
[0160] This system includes means for receiving voice input, means for converting voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, means for making reservations or orders related to the suggested actions, and means for transmitting location information.
[0161] Processing flow
[0162] 1. Accepting voice input
[0163] Users input their meal requests by voice into a device such as a smartphone. For example, they might say, "I'd like to eat some delicious pizza nearby."
[0164] The device is equipped with a microphone and accepts voice input.
[0165] 2. Text conversion of audio data
[0166] The device uses a speech recognition module (e.g., Google Speech Recognition) to convert the audio data into text data.
[0167] 3. Acquisition of location information
[0168] The device uses a GPS module to obtain the user's current location.
[0169] 4. Sending data
[0170] The device sends the acquired text data and location information to the server. A communication module is used.
[0171] 5. Data Analysis and Proposal Generation
[0172] The server analyzes text data and uses natural language processing techniques to understand the user's intent.
[0173] The server uses generative AI (e.g., GPT-3) to generate multiple action suggestions (e.g., restaurant or cafe selection) considering the user's intent, location, season, weather information, etc.
[0174] 6. Presentation of Action Proposals
[0175] The server sends the generated action suggestion to the device. The device then presents the suggestion to the user via text or voice (e.g., "I'd like to recommend John's Pizza nearby. Would you like to make a reservation?").
[0176] 7. Accepting Feedback
[0177] Users provide feedback on the suggested locations, such as "Is there a closer location?"
[0178] The device converts the feedback into text data and sends it to the server.
[0179] 8. Reanalysis and Re-proposal
[0180] The server analyzes the feedback and generates new suggestions based on the new conditions.
[0181] The revised suggestion is sent to the device, and the device presents the suggestion to the user again (e.g., "How about 'Erika's Pizza,' which is a 5-minute walk from there?").
[0182] 9. Execute reservation / order
[0183] If the user agrees to the suggestions, they can make restaurant reservations or order delivery from their device. The server then sends the reservation and order information to the online service to complete the process.
[0184] Specific example
[0185] When a user speaks to a smartphone app saying, "I want to eat some delicious pizza nearby," the app converts the speech to text, retrieves the user's location information, and sends it to a server. The server uses generative AI (e.g., GPT-3) to suggest the best pizza restaurant based on the user's location and preferences. One option presented is "John's Pizza - 5 minutes away," and the app responds with a voice message saying, "I'd like to introduce you to a nearby 'John's Pizza.' Would you like to make a reservation?"
[0186] Example of a prompt
[0187] User: "I want to eat some delicious pizza nearby."
[0188] system:
[0189] Speech recognition: Converts speech to text → "I want to eat some delicious pizza nearby."
[0190] Location information acquisition: Obtain current location from GPS.
[0191] Server: Receives text and location information, and generates suggestions using generative AI → "John's Pizza"
[0192] Response: Voice response from the speech generation module → "I'd like to recommend a nearby 'John's Pizza'. Would you like to make a reservation?"
[0193] Thus, by using the system of the present invention, users can not only easily perform voice input and receive appropriate suggestions, but also smoothly carry out actions based on those suggestions.
[0194] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0195] Step 1:
[0196] The user inputs their dining requests via voice. Using their smartphone's microphone, the user might say, "I want to eat some delicious pizza nearby." This voice data becomes the input.
[0197] Step 2:
[0198] The device converts the voice data into text data. A speech recognition module (e.g., Google Speech Recognition) analyzes the voice and outputs the text data "I want to eat some delicious pizza nearby." This text data becomes the input for the next processing step.
[0199] Step 3:
[0200] The device acquires its current location information. Using a GPS module, it measures the user's current location and obtains latitude and longitude information. This location information becomes the input for the next processing step.
[0201] Step 4:
[0202] The terminal sends text data and location information to the server. The text data and location information are passed to the server via a communication module. This text data and location information becomes the input for the next processing step.
[0203] Step 5:
[0204] The server analyzes the text data to understand the user's intent. Using natural language processing techniques, it analyzes the received text data and outputs the user's intent, "I want recommendations for nearby pizza restaurants." This intent becomes the input for the next processing step.
[0205] Step 6:
[0206] The server generates multiple action suggestions based on the user's intent and location information. Using a generative AI (e.g., GPT-3), it generates several pizza restaurant candidates, taking into account the user's intent, location information, and additional information such as weather and time of day. These action suggestions become the input for the next processing step.
[0207] Step 7:
[0208] The server selects the most suitable action suggestion and sends it to the terminal along with detailed information. The terminal then selects the most appropriate restaurant from the multiple suggestions and sends it along with its details (address, opening hours, menu, etc.). This becomes the input for the next processing step.
[0209] Step 8:
[0210] The device presents the user with an action suggestion. Using a speech synthesis module that converts text data into speech, it suggests to the user verbally, "I'd like to recommend a nearby 'John's Pizza.' Would you like to make a reservation?" This suggestion becomes the input for the next processing step.
[0211] Step 9:
[0212] The user provides feedback. They give voice feedback on the suggested location, such as "Is there a closer location?" This voice data becomes the input for the next processing step.
[0213] Step 10:
[0214] The device converts the user's voice feedback into text data. The voice recognition module is used again to convert the feedback into text data. This feedback becomes the input for the next processing step.
[0215] Step 11:
[0216] The terminal sends feedback text data to the server. The feedback text data is passed to the server via the communication module. This feedback becomes the input for the next processing step.
[0217] Step 12:
[0218] The server analyzes the feedback and generates new action suggestions. The generative AI is used again to generate new candidates based on the feedback and new conditions. These new action suggestions become the input for the next processing step.
[0219] Step 13:
[0220] The server sends a revised suggestion to the terminal, which then presents the suggestion to the user again. It then makes another voice suggestion, such as, "How about 'Erika's Pizza,' which is a 5-minute walk from there?" This process is repeated until the user is satisfied.
[0221] Step 14:
[0222] If the user agrees to the final action suggestion, they can make a reservation or order delivery from the relevant restaurant using their device. The device uses a communication module to send the reservation or order information to the server, which then transmits this information to the online service to complete the process. This reservation or order becomes the final output.
[0223] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0224] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, the invention aims to provide users with more personalized suggestions by incorporating an emotion engine that recognizes user emotions. Specific embodiments of this system are described in detail below.
[0225] System Configuration
[0226] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[0227] Program Processing Description
[0228] 1. User input reception:
[0229] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0230] The terminal accepts voice input and activates the voice recognition module.
[0231] 2. Emotional analysis of voice data:
[0232] The device sends voice input to the emotion engine.
[0233] The emotion engine analyzes the audio data and extracts emotional states (e.g., joy, excitement, fatigue, etc.).
[0234] The extracted emotion data is sent back to the device.
[0235] 3. Convert to text data:
[0236] The device converts voice input into text data.
[0237] The converted text data is reviewed to determine if it is a valid request.
[0238] 4. Sending data to the server:
[0239] The device sends the converted text data, user location information, and sentiment data to the server.
[0240] 5. Data analysis and proposal generation:
[0241] The server analyzes the text data it receives to understand the user's intent.
[0242] The server uses generative AI to generate multiple action suggestions, taking into account location information, current season, weather information, sentiment data, and past suggestion history.
[0243] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[0244] 6. Presentation to the user:
[0245] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[0246] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[0247] 7. Accepting feedback:
[0248] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0249] The device then uses the speech recognition module again to convert the voice feedback into text data.
[0250] 8. Data retransmission and reanalysis:
[0251] The device then sends the converted feedback and sentiment data back to the server.
[0252] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[0253] The server selects the most suitable proposal from the revised proposals and sends the selected proposal to the terminal.
[0254] 9. Final presentation:
[0255] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[0256] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[0257] The following describes the processing flow.
[0258] System Configuration
[0259] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[0260] Program processing steps
[0261] Step 1:
[0262] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0263] The terminal accepts voice input and activates the voice recognition module.
[0264] Step 2:
[0265] The device sends voice data to the emotion engine.
[0266] The emotion engine analyzes the audio data and extracts the user's emotional state (e.g., joy, excitement, fatigue, etc.).
[0267] The extracted emotion data is sent back to the device.
[0268] Step 3:
[0269] The device converts voice input into text data.
[0270] Review the converted text data to determine if it is a valid request.
[0271] Step 4:
[0272] The terminal transmits the converted text data, the user's location information, and the emotion data to the server via the communication module.
[0273] Step 5:
[0274] The server analyzes the received text data to understand the user's intention.
[0275] The server utilizes the generative AI based on the location information, current season, weather information, emotion data, and past proposal history to generate multiple action proposals.
[0276] Step 6:
[0277] The server selects the optimal one from the generated action proposals (e.g., select "Enoshima Coast").
[0278] Summarize the selected proposal content and its detailed information (access method, highlights, weather forecast, etc.).
[0279] Step 7:
[0280] The server transmits the selected proposal content to the terminal as text data.
[0281] Step 8:
[0282] The terminal converts the received proposal content into audio data using the speech synthesis module.
[0283] The terminal presents the generated audio data to the user through the speaker (e.g., "How about a drive to Enoshima Coast?").
[0284] Step 9:
[0285] The user provides feedback on the presented proposal (e.g., "Isn't there a closer place?").
[0286] The terminal uses the speech recognition module again to convert the voice feedback into text data.
[0287] Step ۱۰:
[0288] The terminal sends the converted feedback and emotion data back to the server.
[0289] Step ۱۱:
[0290] The server analyzes the feedback and emotion data and generates multiple action proposals again based on the new conditions.
[0291] The server selects the optimal one from the re-proposals (e.g., selects "Hayama Coast").
[0292] Step ۱۲:
[0293] The server sends the re-selected proposal content to the terminal as text data,
[0294] Step ۱۳:
[0295] The terminal uses the speech synthesis module to convert the received proposal content into voice data and presents it to the user (e.g., "How about Hayama Coast?").
[0296] Through these steps, the user can quickly and intuitively decide on the next action. By using the emotion engine, personalized proposals that suit the user's mood can be made, further improving user satisfaction.
[0297] (Example 2)
[0298] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0299] Conventional behavior suggestion systems often fail to adequately consider users' emotional states and offer only uniform suggestions, resulting in low user satisfaction. Therefore, there is a need to develop a system that provides individually optimized suggestions based on the user's emotional state, thereby increasing user satisfaction. Furthermore, the ability to quickly respond to received feedback and offer revised suggestions is also required.
[0300] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for performing sentiment analysis on received voice data, means for extracting the user's emotional state based on the sentiment analysis, and means for analyzing the converted text data and sentiment data to understand the user's intent. This enables highly personalized suggestions based on sentiment analysis.
[0301] "Voice input" is a method of input that allows users to communicate questions or requests to a system using their voice.
[0302] "Text data" refers to data obtained by analyzing voice input and converting it into text information.
[0303] "Emotional analysis" is the process of extracting a user's emotional state (e.g., joy, excitement, fatigue, etc.) from audio data.
[0304] "Emotional data" refers to data that indicates the user's emotional state, obtained as a result of emotion analysis.
[0305] "Understanding user intent" means analyzing converted text and sentiment data to grasp what the user wants.
[0306] A "generative AI model" is an artificial intelligence algorithm for generating appropriate action proposals based on given data.
[0307] An "action proposal" refers to ideas or plans for proposing actions or activities suitable for the user.
[0308] "Feedback" means that the user responds with opinions or requests regarding the proposal to the system.
[0309] The present invention is a system for efficiently and intuitively planning enjoyable time spent by the user with family and friends. This system includes a dedicated terminal and is assumed to be installed at a place where people gather. Furthermore, the present invention aims to provide more personalized proposals to the user by combining an emotion engine that recognizes the user's emotions.
[0310] The system configuration consists of a terminal that accepts voice input, a server that analyzes data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, and the server is equipped with a data analysis module and a generative AI model. The emotion engine extracts emotions from the user's voice data and adjusts the action proposal based on the analysis result.
[0311] First, the system starts operating when the user inputs a question or request by voice to the terminal. For example, a voice input such as "Where should we go to play today?" is made. The terminal receives the voice, activates the voice recognition module (e.g., Google Cloud Speech-to-Text API), and converts the voice into text data. This conversion may take several seconds.
[0312] Next, the terminal uses an emotion engine (e.g., IBM Watson® Tone Analyzer) to analyze the voice data and extract the user's emotional state. The extracted emotion data is sent back to the terminal and then sent to the server along with the converted text data. The server receives this data and uses natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to analyze the user's intent.
[0313] The server further analyzes location information, current season, weather information, and past behavioral history, and uses a generative AI model (e.g., OpenAI GPT-3) to generate multiple action suggestions. It selects the most suitable suggestion from the generated suggestions and sends it back to the terminal. The terminal converts the received suggestion into audio data using a speech synthesis module (e.g., Amazon Polly) and presents it to the user through its speaker.
[0314] The user provides feedback on the presented suggestion (e.g., "Is there a place a little closer?"). The device uses the speech recognition module again to convert the voice feedback into text data and sends it back to the server. The server analyzes the feedback and sentiment data, generates multiple action suggestions again based on the new conditions, and sends the best suggestion back to the device. Finally, the device uses the speech synthesis module to convert the received suggestion into voice data and presents it to the user.
[0315] As a concrete example, the following prompt sentences could be input into the generation AI model:
[0316] Prompt: "Please suggest some family-friendly outing destinations. Users are happy and prefer nearby locations."
[0317] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[0318] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0319] Step 1:
[0320] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0321] The device accepts voice input and activates a speech recognition module (e.g., Google Cloud Speech-to-Text API). The voice data is converted into text data. The input is the user's voice data, and the output is text data.
[0322] Step 2:
[0323] The terminal sends the converted text data to an emotion engine (e.g., IBM Watson Tone Analyzer). The emotion engine analyzes the audio data and extracts the user's emotional state. The input is audio data, and the output is emotion data.
[0324] Step 3:
[0325] The device verifies the sentiment data and text data returned from the sentiment engine. It performs text verification to determine if the request is appropriate. The input is sentiment data and text data, and the output is the verified text data and sentiment data.
[0326] Step 4:
[0327] The device sends verified text data, user location information (e.g., GPS data), and sentiment data to the server. This communication typically uses the HTTP or HTTPS protocol. Inputs are text data, location information, and sentiment data, while output is a notification to the server that the transmission is complete.
[0328] Step 5:
[0329] The server analyzes the received text data using natural language processing algorithms (e.g., SpaCy or NLTK) to understand the user's intent. The input is text data, and the output is user intent data.
[0330] Step 6:
[0331] The server considers the user's location, current season, weather information, and sentiment data, and also accesses a database of past behavioral history. It generates multiple action suggestions using a generative AI model (e.g., OpenAI GPT-3). The input is location information, season, weather information, sentiment data, and past behavioral history, and the output is a list of action suggestions.
[0332] Step 7:
[0333] The server uses a machine learning algorithm to select the most suitable action from the generated suggestions. The selected action suggestion, along with detailed information, is then sent to the terminal. The input is a list of action suggestions, and the output is the optimal action suggestion and its details.
[0334] Step 8:
[0335] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Amazon Polly). This data is then presented to the user through the speaker. The input consists of optimal action suggestions and detailed information, while the output is speech data.
[0336] Step 9:
[0337] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?"). The device then uses its speech recognition module again to convert the voice feedback into text data. The input is the user's voice data, and the output is the text data of the feedback.
[0338] Step 10:
[0339] The device then resends the converted feedback and sentiment data to the server. The inputs are the feedback text data and sentiment data, and the output is a notification that the transmission to the server is complete.
[0340] Step 11:
[0341] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions. The input is feedback text data and sentiment data, and the output is a list of new action suggestions.
[0342] Step 12:
[0343] The server selects the most suitable action from the regenerated suggestions and sends the selected suggestion to the terminal. The input is a list of new action suggestions, and the output is the re-selected action suggestion.
[0344] Step 13:
[0345] The terminal receives the suggested content again, converts it into speech data using a speech synthesis module, and presents it to the user. The input is the re-selected action suggestion, and the output is the speech data.
[0346] (Application Example 2)
[0347] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0348] Conventional behavior suggestion systems have difficulty taking into account the user's emotional state and detailed location information, resulting in suggestions that are not sufficiently personalized. Furthermore, the process of regenerating suggestions based on user voice feedback is inefficient, making it difficult to increase user satisfaction.
[0349] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for extracting the emotional state and adjusting the action suggestions based on the extracted emotional data, means for optimizing the action suggestions generated based on the user's current location information, emotional data, and past suggestion history, means for presenting the generated action suggestions, means for receiving feedback from the user, and means for generating and presenting suggestions again based on the received feedback. This makes it possible to provide personalized action suggestions that take into account the user's emotional state and location information, thereby improving user satisfaction.
[0350] A "means for receiving voice input" refers to a module for acquiring voice input from a user as digital data.
[0351] "Means for converting voice input into text data" refers to a program that converts acquired voice data into text information.
[0352] "Means of analyzing text data and understanding user intent" refers to algorithms that analyze converted text data to understand the content of user utterances and requests.
[0353] "A means of generating multiple action suggestions based on analyzed data" refers to a program that generates multiple actions or choices that a user should take, based on the results of analyzing text data.
[0354] "Means for extracting emotional states" refers to an analysis engine that automatically extracts a user's emotions from audio and video data.
[0355] "Means for adjusting action suggestions based on extracted emotional data" refers to a program that modifies generated action suggestions to a more appropriate form according to the user's emotional state.
[0356] "Means for transmitting location information" refers to a module that acquires data on the user's current location and sends it to an analysis server.
[0357] "Means for optimizing action suggestions generated based on the user's current location information, emotional data, and past suggestion history" refers to a program that generates optimal action suggestions by considering multiple factors such as the user's location information, emotional state, and past suggestion history.
[0358] "Means for presenting generated action suggestions" refers to output devices such as displays and speech synthesis modules, as well as programs, for providing generated action suggestions to the user.
[0359] "Means for receiving user feedback" refers to a module that receives responses and opinions from users regarding suggestions they have made, in the form of audio or text.
[0360] "A means of generating and presenting new suggestions based on received feedback" refers to a program that analyzes feedback received from users, generates new action suggestions based on that analysis, and presents them to the user again.
[0361] "Means of providing detailed information" refers to modules that provide users with background information and additional explanations for the generated action suggestions.
[0362] This invention provides a system that allows users to efficiently receive action suggestions in physical stores using voice input. The specific configuration and operation of this system are described in detail below.
[0363] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module, a communication module, and an emotion recognition camera. The server includes a data analysis module, a generative AI model, and a database.
[0364] Specific hardware and software examples include the following:
[0365] Speech recognition module: Google Speech-to-Text API
[0366] Emotion analysis engine: Microsoft® Azure® Emotion API
[0367] Generative AI: OpenAI GPT Model
[0368] Communication module: HTTP / HTTPS
[0369] Data Analysis Module: Python Scripts and SQL Databases
[0370] System operation
[0371] 1. Accepting voice input:
[0372] The user initiates voice input by speaking into a smart terminal installed in the store. The user's voice is captured through the microphone and transmitted to the system as digital data.
[0373] 2. Text conversion and analysis of audio data:
[0374] The device converts the acquired audio data into text data using the Google Speech-to-Text API. The converted text data is sent to the server, where a data analysis module analyzes the user's intent.
[0375] 3. Extraction of emotional states:
[0376] The device's built-in emotion recognition camera and audio data are used to extract the user's emotional state via the Microsoft Azure Emotion API. This emotional data is then sent to a server.
[0377] 4. Generating action proposals:
[0378] The server uses the OpenAI GPT model to generate multiple action suggestions based on analyzed text data, sentiment data, user location information, and past suggestion history. These action suggestions are adjusted according to the user's emotional state.
[0379] 5. Presenting action proposals:
[0380] The generated action suggestions are sent from the server to the terminal and presented to the user using the terminal's display or speech synthesis module. A concrete example is, "There's a new autumn book fair going on at the bookstore on the third floor."
[0381] 6. Receiving user feedback and making revisions:
[0382] When a user provides feedback on a suggested solution, the device accepts voice input again and performs re-analysis. Based on the feedback, the server regenerates multiple action suggestions under new conditions and presents them to the user again. An example of a specific prompt is: "The user is asking, 'Are there any stores closer?' Their emotional state is excited, and their current location is the first floor of a shopping mall. Please tell me about recommended stores and events."
[0383] This allows the system to provide personalized action suggestions that take into account the user's emotional state and location, thereby improving user satisfaction.
[0384] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0385] Step 1:
[0386] The user initiates voice input. The user speaks into the smart device, and the voice input is acquired through the device's microphone. The input is the user's voice data, and the output is digital voice data.
[0387] Step 2:
[0388] The device converts the audio data into text data. The acquired audio data is converted into text data using the Google Speech-to-Text API. The input is audio data, and the output is text data.
[0389] Step 3:
[0390] The terminal sends text and audio data to the server. The terminal sends the converted text data and the acquired audio data to the server. The input is text and audio data, and the output is the data sent to the server.
[0391] Step 4:
[0392] The server analyzes text data to understand the user's intent. The data analysis module analyzes text data to understand the user's requests and intent. The input is text data, and the output is the analysis result.
[0393] Step 5:
[0394] The server extracts emotional data. The audio data is analyzed for emotional state using the Microsoft Azure Emotion API, and emotional data is extracted. The input is audio data, and the output is emotional data.
[0395] Step 6:
[0396] The server generates action suggestions. The server uses the OpenAI GPT model to generate action suggestions, taking into account analysis results, sentiment data, user location information, and past suggestion history. The inputs are analysis results, sentiment data, location information, and past suggestion history, and the output is multiple action suggestions.
[0397] Step 7:
[0398] The server optimizes the action suggestions. Using generative AI and a database, it optimizes the generated action suggestions according to the user's emotional state. The input is the action suggestion, and the output is the optimized action suggestion.
[0399] Step 8:
[0400] The server sends optimized action suggestions to the terminal, which then presents them to the user. The input is the optimized action suggestions, and the output is the data sent to the terminal.
[0401] Step 9:
[0402] The device presents action suggestions. Using the device's display and speech synthesis module, the generated action suggestions are presented to the user visually and audibly. The input is optimized action suggestions, and the output is the presentation of suggestions to the user.
[0403] Step 10:
[0404] The user provides feedback on the proposal. The user inputs the feedback by voice, which is captured through the device's microphone. The input is feedback voice data, and the output is digital voice data.
[0405] Step 11:
[0406] The device converts the feedback audio data into text. The acquired feedback audio data is then converted back into text data using the Google Speech-to-Text API. The input is the feedback audio data, and the output is text data.
[0407] Step 12:
[0408] The server re-parses the feedback text data and sentiment data. The terminal sends the converted feedback text data and sentiment data to the server, which then re-parses it. The input is the feedback text data and sentiment data, and the output is the re-parsed result.
[0409] Step 13:
[0410] The server generates action suggestions again. Based on the re-analysis results, it generates action suggestions again using the OpenAI GPT model according to the new conditions. The input is the re-analysis results, and the output is the regenerated action suggestions.
[0411] Step 14:
[0412] The server optimizes the regenerated action suggestions and sends them to the terminal. The terminal then presents them to the user. The input is the regenerated action suggestions, and the output is the data sent to the terminal.
[0413] Step 15:
[0414] The terminal presents the action suggestion again. The action suggestion is presented to the user again using a display or speech synthesis module. The input is the regenerated action suggestion, and the output is the presentation of the revised suggestion to the user.
[0415] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0416] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0417] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0418] [Second Embodiment]
[0419] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0420] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0421] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0422] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0423] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0424] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0425] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0426] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0427] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0428] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0429] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0430] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0431] This invention provides a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Specific embodiments of this system are described in detail below.
[0432] System Configuration
[0433] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[0434] Program Processing Description
[0435] 1. User input reception:
[0436] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0437] The device uses a speech recognition module to convert voice input into text data.
[0438] The device sends the converted text data and the user's location information to the server.
[0439] 2. Data analysis and proposal generation:
[0440] The server analyzes the text data it receives to understand the user's intent.
[0441] The server uses generative AI to generate multiple action suggestions, taking into account location information, season, weather information, and past suggestion history.
[0442] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[0443] 3. Presentation to the user:
[0444] The terminal converts the received suggestion into speech using a speech synthesis module.
[0445] The device presents the generated audio data to the user (e.g., "How about a drive along the Shonan coast?").
[0446] 4. Accepting feedback:
[0447] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0448] The device then uses the speech recognition module again to convert the voice feedback into text data.
[0449] The terminal sends the converted feedback to the server.
[0450] 5. Reanalysis and proposals (if necessary):
[0451] The server analyzes the feedback and generates new suggestions based on the new conditions.
[0452] The server sends a revised suggestion to the terminal, and the terminal again presents the suggestion to the user via voice (e.g., "How about Hayama Beach?").
[0453] Specific example
[0454] When a user asks the device, "Where should we go today?", the device converts the voice input into text and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations, then selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device informs the user of this. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[0455] This system is designed to allow users to make decisions about their next actions with ease, using friendly language and an intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with confidence.
[0456] The following describes the processing flow.
[0457] Step 1:
[0458] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0459] The terminal accepts voice input and activates the voice recognition module.
[0460] Step 2:
[0461] The device converts voice input into text data.
[0462] The converted text data is reviewed to determine if it is a valid request.
[0463] Step 3:
[0464] The terminal transmits the converted text data and the user's location information to the server via the communication module.
[0465] Step 4:
[0466] The server processes the received text data using a parsing module to understand the user's intent (e.g., understanding the intent "I want to go somewhere where I can see the ocean").
[0467] Step 5:
[0468] The server references the user's location information, current season, weather information, and past suggestion history, and uses generative AI to generate multiple action suggestions.
[0469] Step 6:
[0470] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[0471] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[0472] Step 7:
[0473] The server sends the selected proposal as text data to the terminal.
[0474] Step 8:
[0475] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[0476] The device generates audio data which is then presented to the user through the speaker (e.g., "How about a drive along the Shonan coast?").
[0477] Step 9:
[0478] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0479] The device restarts its speech recognition module and converts the user's feedback into text data.
[0480] Step 10:
[0481] The device sends the converted feedback back to the server.
[0482] Step 11:
[0483] The server analyzes the feedback and generates multiple action suggestions again based on the new conditions.
[0484] Select the best option from the revised proposals (e.g., choose "Hayama Beach").
[0485] Step 12:
[0486] The server sends the re-selected proposals to the terminal as text data.
[0487] Step 13:
[0488] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[0489] Through these steps, users can quickly and intuitively decide on their next course of action.
[0490] (Example 1)
[0491] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0492] Conventional action planning systems have the problem of requiring users to search for, evaluate, and judge information themselves, which is time-consuming. Furthermore, the quality of suggestions is often inconsistent, making it difficult to accurately reflect the user's intentions. In addition, there is a lack of voice-based interfaces, hindering the intuitive and rapid development of action plans.
[0493] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0494] In this invention, the server includes means for analyzing text data and understanding the user's intent, means for generating multiple action suggestions using generative artificial intelligence based on the analyzed data, and means for selecting the optimal action suggestion and transmitting it along with detailed information. This makes it possible to accurately reflect the user's intent and provide high-quality action suggestions quickly and intuitively.
[0495] "Means for receiving voice input" refers to a combination of hardware and software for capturing the user's voice as a digital signal.
[0496] "Means for converting voice input into text data" refers to speech recognition technology that analyzes captured voice signals and converts them into corresponding text data.
[0497] "Means for transmitting location information to a server" refers to a communication module that transmits location information, which identifies the user's current location, to a server.
[0498] "Methods for analyzing text data and understanding user intent" refers to the process of analyzing text data using natural language processing technology to grasp the intent behind user requests and questions.
[0499] "Means for generating multiple action suggestions using generative artificial intelligence" refers to an algorithm that uses artificial intelligence technology to create multiple action options based on user requests.
[0500] The "means for selecting the most suitable action proposal and sending it along with detailed information" refers to a function that selects the most appropriate action proposal from those generated and sends its detailed information (e.g., access methods and local conditions) to the terminal.
[0501] "Means for converting generated action suggestions into speech and presenting them to the user" refers to speech synthesis technology and output devices that convert text data into speech data and present suggestions to the user visually and aurally.
[0502] "A means of receiving user feedback, converting audio into text data, and sending it to a server" refers to a process that receives user feedback as voice input, converts it into text, and transmits it to a server.
[0503] "A means of generating new suggestions based on received feedback and presenting them in audio format" refers to a process of generating new action suggestions based on user feedback, converting them back into audio data, and presenting them to the user.
[0504] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[0505] System Configuration
[0506] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[0507] Hardware and software to be used
[0508] The device includes a microphone for receiving voice input, a speech recognition module (e.g., Google Cloud Speech-to-Text API, Amazon Transcribe, etc.) for converting speech to text data, and a GPS module for obtaining the user's location information. A communication module is used to send the converted text data and location information to the server. On the server side, natural language processing technology (e.g., BERT, GPT-3, etc.) is used as a data analysis module to understand the user's intent. Furthermore, generative AI (e.g., OpenAI's GPT-3, etc.) is used to generate multiple action suggestions. After the server selects the optimal action suggestion, it sends it to the device along with detailed information (such as how to access it and the local weather). The device uses a speech synthesis module (e.g., Google Text-to-Speech API, Amazon Polly, etc.) to convert the suggested content into speech and present it to the user.
[0509] Specific example
[0510] As a concrete example, suppose a user asks their device, "Where should we go today?" The device converts the voice input into text data and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations. From these, it selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device communicates this to the user. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[0511] As described above, this system is designed to allow users to make decisions about their next actions in an enjoyable way through its user-friendly language and intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with peace of mind.
[0512] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0513] Step 1:
[0514] User input reception
[0515] The user speaks into the device to ask a question or make a request (e.g., "Where should we go today?"). The device's microphone collects the voice data and uses it as input for a speech recognition module (such as Google Cloud Speech-to-Text API or Amazon Transcribe). The speech recognition module converts the voice data into text data, generating a text-based input. This text data includes the user's intent or inquiry.
[0516] Step 2:
[0517] Acquiring and transmitting location information
[0518] The device uses a GPS module to determine the user's current location. The acquired location information is sent to the server along with text data. Data communication is conducted using the HTTPS protocol during transmission, ensuring data security.
[0519] Step 3:
[0520] Data Analysis
[0521] The server analyzes the received text data and location information. It uses natural language processing models (e.g., BERT or GPT-3) to analyze the text data and understand the user's intent. This process extracts keywords and contextual information from the text data. Based on the analysis results, the server generates data that reflects the user's wishes and requests.
[0522] Step 4:
[0523] Generating action proposals
[0524] The server generates multiple action suggestions based on the analysis results using a generative AI (e.g., OpenAI's GPT-3). The server also takes into account external data such as the user's location, season, and weather information (e.g., using the OpenWeatherMap API). This improves the accuracy and appropriateness of the suggestions. The generative AI generates multiple candidate suggestions and adds detailed information to each suggestion.
[0525] Step 5:
[0526] Selecting and sending the most suitable proposal
[0527] The server uses evaluation criteria (e.g., distance, weather conditions, event information, etc.) to select the most suitable action from the generated suggestions. It then sends the best suggestion and its detailed information (such as access methods) to the terminal. The server uses asynchronous communication to transmit data quickly and efficiently.
[0528] Step 6:
[0529] Presentation to the user
[0530] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Google Text-to-Speech API or Amazon Polly). The generated speech data is then presented to the user through the speaker (e.g., "How about a drive to Shonan Beach?"). Simultaneously, displaying the suggestions in text format on the screen is also being considered.
[0531] Step 7:
[0532] Accepting feedback
[0533] The user provides voice feedback on the suggested location (e.g., "Is there a place a little closer?"). The device uses a voice recognition module to convert the voice feedback into text data. This text data is then sent back to the server for further analysis.
[0534] Step 8:
[0535] Generation and presentation of revised proposals
[0536] The server analyzes the feedback and generates new suggestions based on the new conditions. Generative AI is used again in this re-suggestion to create new, appropriate candidates. The server sends detailed information about the re-suggestion to the terminal, which then presents it to the user again (e.g., "How about Hayama Beach?"). This allows the system to repeat suggestions until the user is satisfied.
[0537] Through this series of processes, users can intuitively and efficiently plan their actions.
[0538] (Application Example 1)
[0539] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0540] In modern society, efficiently and intuitively planning time with family and friends is crucial for users to make the most of their limited time. However, users often have to go through the trouble of translating their desires into concrete and appropriate action suggestions, and obtaining detailed information to judge the appropriateness of those suggestions. Furthermore, the lack of means to properly make reservations or orders for the suggested actions can hinder the smooth execution of those suggestions. There is a need for a system that can solve this problem.
[0541] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0542] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data and understanding the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, and means for making reservations or orders related to the suggested actions. This allows the user to easily receive action suggestions through voice input, provide feedback to obtain the most suitable suggestion, and smoothly transition to actual actions based on the suggestions.
[0543] "Means for receiving voice input" refers to hardware or software that receives voice data and converts it into a format that can be processed within the system.
[0544] "Means for converting received voice input into text data" refers to software or hardware that uses speech recognition technology to convert a user's voice input into corresponding text data.
[0545] "Means for analyzing converted text data and understanding user intent" refers to software or hardware that uses natural language processing techniques to analyze the content of text data and understand what the user is asking for.
[0546] "Means for generating multiple action suggestions based on analyzed data" refers to an algorithm for generating multiple action options for the user based on the analysis results, and the hardware or software that executes it.
[0547] "Means for presenting generated action suggestions" refers to devices and software that present generated action options to the user, either visually or audibly.
[0548] "Means for receiving user feedback" refers to the hardware and software used to receive and incorporate user responses to suggestions into the system.
[0549] "Means for generating and presenting new suggestions based on received feedback" refers to algorithms and devices that generate and present new action suggestions based on user feedback.
[0550] "Means of making reservations or orders related to suggested actions" refers to online services for making restaurant reservations or ordering goods based on actions suggested within the system, as well as the software and hardware for executing them.
[0551] "Means of transmitting location information" refers to a GPS module and communication module used to determine the user's current location and transmit it to a server.
[0552] This invention provides a system for users to plan meals efficiently and intuitively. The system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[0553] System Configuration
[0554] This system includes means for receiving voice input, means for converting voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, means for making reservations or orders related to the suggested actions, and means for transmitting location information.
[0555] Processing flow
[0556] 1. Accepting voice input
[0557] Users input their meal requests by voice into a device such as a smartphone. For example, they might say, "I'd like to eat some delicious pizza nearby."
[0558] The device is equipped with a microphone and accepts voice input.
[0559] 2. Text conversion of audio data
[0560] The device uses a speech recognition module (e.g., Google Speech Recognition) to convert the audio data into text data.
[0561] 3. Acquisition of location information
[0562] The device uses a GPS module to obtain the user's current location.
[0563] 4. Sending data
[0564] The device sends the acquired text data and location information to the server. A communication module is used.
[0565] 5. Data Analysis and Proposal Generation
[0566] The server analyzes text data and uses natural language processing techniques to understand the user's intent.
[0567] The server uses generative AI (e.g., GPT-3) to generate multiple action suggestions (e.g., restaurant or cafe selection) considering the user's intent, location, season, weather information, etc.
[0568] 6. Presentation of Action Proposals
[0569] The server sends the generated action suggestion to the device. The device then presents the suggestion to the user via text or voice (e.g., "I'd like to recommend John's Pizza nearby. Would you like to make a reservation?").
[0570] 7. Accepting Feedback
[0571] Users provide feedback on the suggested locations, such as "Is there a closer location?"
[0572] The device converts the feedback into text data and sends it to the server.
[0573] 8. Reanalysis and Re-proposal
[0574] The server analyzes the feedback and generates new suggestions based on the new conditions.
[0575] The revised suggestion is sent to the device, and the device presents the suggestion to the user again (e.g., "How about 'Erika's Pizza,' which is a 5-minute walk from there?").
[0576] 9. Execute reservation / order
[0577] If the user agrees to the suggestions, they can make restaurant reservations or order delivery from their device. The server then sends the reservation and order information to the online service to complete the process.
[0578] Specific example
[0579] When a user speaks to a smartphone app saying, "I want to eat some delicious pizza nearby," the app converts the speech to text, retrieves the user's location information, and sends it to a server. The server uses generative AI (e.g., GPT-3) to suggest the best pizza restaurant based on the user's location and preferences. One option presented is "John's Pizza - 5 minutes away," and the app responds with a voice message saying, "I'd like to introduce you to a nearby 'John's Pizza.' Would you like to make a reservation?"
[0580] Example of a prompt
[0581] User: "I want to eat some delicious pizza nearby."
[0582] system:
[0583] Speech recognition: Converts speech to text → "I want to eat some delicious pizza nearby."
[0584] Location information acquisition: Obtain current location from GPS.
[0585] Server: Receives text and location information, and generates suggestions using generative AI → "John's Pizza"
[0586] Response: Voice response from the speech generation module → "I'd like to recommend a nearby 'John's Pizza'. Would you like to make a reservation?"
[0587] Thus, by using the system of the present invention, users can not only easily perform voice input and receive appropriate suggestions, but also smoothly carry out actions based on those suggestions.
[0588] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0589] Step 1:
[0590] The user inputs their dining requests via voice. Using their smartphone's microphone, the user might say, "I want to eat some delicious pizza nearby." This voice data becomes the input.
[0591] Step 2:
[0592] The device converts the voice data into text data. A speech recognition module (e.g., Google Speech Recognition) analyzes the voice and outputs the text data "I want to eat some delicious pizza nearby." This text data becomes the input for the next processing step.
[0593] Step 3:
[0594] The device acquires its current location information. Using a GPS module, it measures the user's current location and obtains latitude and longitude information. This location information becomes the input for the next processing step.
[0595] Step 4:
[0596] The terminal sends text data and location information to the server. The text data and location information are passed to the server via a communication module. This text data and location information becomes the input for the next processing step.
[0597] Step 5:
[0598] The server analyzes the text data to understand the user's intent. Using natural language processing techniques, it analyzes the received text data and outputs the user's intent, "I want recommendations for nearby pizza restaurants." This intent becomes the input for the next processing step.
[0599] Step 6:
[0600] The server generates multiple action suggestions based on the user's intent and location information. Using a generative AI (e.g., GPT-3), it generates several pizza restaurant candidates, taking into account the user's intent, location information, and additional information such as weather and time of day. These action suggestions become the input for the next processing step.
[0601] Step 7:
[0602] The server selects the most suitable action suggestion and sends it to the terminal along with detailed information. The terminal then selects the most appropriate restaurant from the multiple suggestions and sends it along with its details (address, opening hours, menu, etc.). This becomes the input for the next processing step.
[0603] Step 8:
[0604] The device presents the user with an action suggestion. Using a speech synthesis module that converts text data into speech, it suggests to the user verbally, "I'd like to recommend a nearby 'John's Pizza.' Would you like to make a reservation?" This suggestion becomes the input for the next processing step.
[0605] Step 9:
[0606] The user provides feedback. They give voice feedback on the suggested location, such as "Is there a closer location?" This voice data becomes the input for the next processing step.
[0607] Step 10:
[0608] The device converts the user's voice feedback into text data. The voice recognition module is used again to convert the feedback into text data. This feedback becomes the input for the next processing step.
[0609] Step 11:
[0610] The terminal sends feedback text data to the server. The feedback text data is passed to the server via the communication module. This feedback becomes the input for the next processing step.
[0611] Step 12:
[0612] The server analyzes the feedback and generates new action suggestions. The generative AI is used again to generate new candidates based on the feedback and new conditions. These new action suggestions become the input for the next processing step.
[0613] Step 13:
[0614] The server sends a revised suggestion to the terminal, which then presents the suggestion to the user again. It then makes another voice suggestion, such as, "How about 'Erika's Pizza,' which is a 5-minute walk from there?" This process is repeated until the user is satisfied.
[0615] Step 14:
[0616] If the user agrees to the final action suggestion, they can make a reservation or order delivery from the relevant restaurant using their device. The device uses a communication module to send the reservation or order information to the server, which then transmits this information to the online service to complete the process. This reservation or order becomes the final output.
[0617] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0618] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, the invention aims to provide users with more personalized suggestions by incorporating an emotion engine that recognizes user emotions. Specific embodiments of this system are described in detail below.
[0619] System Configuration
[0620] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[0621] Program Processing Description
[0622] 1. User input reception:
[0623] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0624] The terminal accepts voice input and activates the voice recognition module.
[0625] 2. Emotional analysis of voice data:
[0626] The device sends voice input to the emotion engine.
[0627] The emotion engine analyzes the audio data and extracts emotional states (e.g., joy, excitement, fatigue, etc.).
[0628] The extracted emotion data is sent back to the device.
[0629] 3. Convert to text data:
[0630] The device converts voice input into text data.
[0631] The converted text data is reviewed to determine if it is a valid request.
[0632] 4. Sending data to the server:
[0633] The device sends the converted text data, user location information, and sentiment data to the server.
[0634] 5. Data analysis and proposal generation:
[0635] The server analyzes the text data it receives to understand the user's intent.
[0636] The server uses generative AI to generate multiple action suggestions, taking into account location information, current season, weather information, sentiment data, and past suggestion history.
[0637] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[0638] 6. Presentation to the user:
[0639] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[0640] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[0641] 7. Accepting feedback:
[0642] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0643] The device then uses the speech recognition module again to convert the voice feedback into text data.
[0644] 8. Data retransmission and reanalysis:
[0645] The device then sends the converted feedback and sentiment data back to the server.
[0646] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[0647] The server selects the most suitable proposal from the revised proposals and sends the selected proposal to the terminal.
[0648] 9. Final presentation:
[0649] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[0650] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[0651] The following describes the processing flow.
[0652] System Configuration
[0653] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[0654] Program processing steps
[0655] Step 1:
[0656] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0657] The terminal accepts voice input and activates the voice recognition module.
[0658] Step 2:
[0659] The device sends voice data to the emotion engine.
[0660] The emotion engine analyzes the audio data and extracts the user's emotional state (e.g., joy, excitement, fatigue, etc.).
[0661] The extracted emotion data is sent back to the device.
[0662] Step 3:
[0663] The device converts voice input into text data.
[0664] Review the converted text data to determine if it is a valid request.
[0665] Step 4:
[0666] The device transmits the converted text data, user location information, and sentiment data to the server via a communication module.
[0667] Step 5:
[0668] The server analyzes the text data it receives to understand the user's intent.
[0669] The server uses generative AI to generate multiple action suggestions based on location information, current season, weather information, sentiment data, and past suggestion history.
[0670] Step 6:
[0671] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[0672] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[0673] Step 7:
[0674] The server sends the selected proposal as text data to the terminal.
[0675] Step 8:
[0676] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[0677] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[0678] Step 9:
[0679] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0680] The device then uses the speech recognition module again to convert the voice feedback into text data.
[0681] Step 10:
[0682] The device then sends the converted feedback and sentiment data back to the server.
[0683] Step 11:
[0684] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[0685] The server selects the best option from the alternative suggestions (e.g., "Hayama Beach").
[0686] Step 12:
[0687] The server sends the re-selected proposals to the terminal as text data.
[0688] Step 13:
[0689] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[0690] Through these steps, users can quickly and intuitively decide on their next course of action. By using an emotion engine, personalized suggestions that align with the user's feelings become possible, further improving user satisfaction.
[0691] (Example 2)
[0692] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0693] Conventional behavior suggestion systems often fail to adequately consider users' emotional states and offer only uniform suggestions, resulting in low user satisfaction. Therefore, there is a need to develop a system that provides individually optimized suggestions based on the user's emotional state, thereby increasing user satisfaction. Furthermore, the ability to quickly respond to received feedback and offer revised suggestions is also required.
[0694] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for performing sentiment analysis on received voice data, means for extracting the user's emotional state based on the sentiment analysis, and means for analyzing the converted text data and sentiment data to understand the user's intent. This enables highly personalized suggestions based on sentiment analysis.
[0695] "Voice input" is a method of input that allows users to communicate questions or requests to a system using their voice.
[0696] "Text data" refers to data obtained by analyzing voice input and converting it into text information.
[0697] "Emotional analysis" is the process of extracting a user's emotional state (e.g., joy, excitement, fatigue, etc.) from audio data.
[0698] "Emotional data" refers to data that indicates the user's emotional state, obtained as a result of emotion analysis.
[0699] "Understanding user intent" means analyzing converted text and sentiment data to grasp what the user wants.
[0700] A "generative AI model" is an artificial intelligence algorithm that generates appropriate action suggestions based on given data.
[0701] "Action suggestions" refer to ideas and plans for proposing actions and activities that are appropriate for the user.
[0702] "Feedback" refers to users providing feedback, opinions, and requests regarding suggestions to the system.
[0703] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, the invention aims to provide users with more personalized suggestions.
[0704] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI model. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[0705] First, the system starts working when the user inputs a question or request via voice into the device. For example, a voice input might be, "Where should we go today?" The device receives the voice input, activates a speech recognition module (such as the Google Cloud Speech-to-Text API), and converts the voice into text data. This conversion may take several seconds.
[0706] Next, the terminal uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the voice data and extract the user's emotional state. The extracted emotion data is sent back to the terminal and then transmitted to the server along with the converted text data. The server receives this data and uses natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to analyze the user's intent.
[0707] The server further analyzes location information, current season, weather information, and past behavioral history, and uses a generative AI model (e.g., OpenAI GPT-3) to generate multiple action suggestions. It selects the most suitable suggestion from the generated suggestions and sends it back to the terminal. The terminal converts the received suggestion into audio data using a speech synthesis module (e.g., Amazon Polly) and presents it to the user through its speaker.
[0708] The user provides feedback on the presented suggestion (e.g., "Is there a place a little closer?"). The device uses the speech recognition module again to convert the voice feedback into text data and sends it back to the server. The server analyzes the feedback and sentiment data, generates multiple action suggestions again based on the new conditions, and sends the best suggestion back to the device. Finally, the device uses the speech synthesis module to convert the received suggestion into voice data and presents it to the user.
[0709] As a concrete example, the following prompt sentences could be input into the generation AI model:
[0710] Prompt: "Please suggest some family-friendly outing destinations. Users are happy and prefer nearby locations."
[0711] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[0712] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0713] Step 1:
[0714] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0715] The device accepts voice input and activates a speech recognition module (e.g., Google Cloud Speech-to-Text API). The voice data is converted into text data. The input is the user's voice data, and the output is text data.
[0716] Step 2:
[0717] The terminal sends the converted text data to an emotion engine (e.g., IBM Watson Tone Analyzer). The emotion engine analyzes the audio data and extracts the user's emotional state. The input is audio data, and the output is emotion data.
[0718] Step 3:
[0719] The device verifies the sentiment data and text data returned from the sentiment engine. It performs text verification to determine if the request is appropriate. The input is sentiment data and text data, and the output is the verified text data and sentiment data.
[0720] Step 4:
[0721] The device sends verified text data, user location information (e.g., GPS data), and sentiment data to the server. This communication typically uses the HTTP or HTTPS protocol. Inputs are text data, location information, and sentiment data, while output is a notification to the server that the transmission is complete.
[0722] Step 5:
[0723] The server analyzes the received text data using natural language processing algorithms (e.g., SpaCy or NLTK) to understand the user's intent. The input is text data, and the output is user intent data.
[0724] Step 6:
[0725] The server considers the user's location, current season, weather information, and sentiment data, and also accesses a database of past behavioral history. It generates multiple action suggestions using a generative AI model (e.g., OpenAI GPT-3). The input is location information, season, weather information, sentiment data, and past behavioral history, and the output is a list of action suggestions.
[0726] Step 7:
[0727] The server uses a machine learning algorithm to select the most suitable action from the generated suggestions. The selected action suggestion, along with detailed information, is then sent to the terminal. The input is a list of action suggestions, and the output is the optimal action suggestion and its details.
[0728] Step 8:
[0729] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Amazon Polly). This data is then presented to the user through the speaker. The input consists of optimal action suggestions and detailed information, while the output is speech data.
[0730] Step 9:
[0731] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?"). The device then uses its speech recognition module again to convert the voice feedback into text data. The input is the user's voice data, and the output is the text data of the feedback.
[0732] Step 10:
[0733] The device then resends the converted feedback and sentiment data to the server. The inputs are the feedback text data and sentiment data, and the output is a notification that the transmission to the server is complete.
[0734] Step 11:
[0735] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions. The input is feedback text data and sentiment data, and the output is a list of new action suggestions.
[0736] Step 12:
[0737] The server selects the most suitable action from the regenerated suggestions and sends the selected suggestion to the terminal. The input is a list of new action suggestions, and the output is the re-selected action suggestion.
[0738] Step 13:
[0739] The terminal receives the suggested content again, converts it into speech data using a speech synthesis module, and presents it to the user. The input is the re-selected action suggestion, and the output is the speech data.
[0740] (Application Example 2)
[0741] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0742] Conventional behavior suggestion systems have difficulty taking into account the user's emotional state and detailed location information, resulting in suggestions that are not sufficiently personalized. Furthermore, the process of regenerating suggestions based on user voice feedback is inefficient, making it difficult to increase user satisfaction.
[0743] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for extracting the emotional state and adjusting the action suggestions based on the extracted emotional data, means for optimizing the action suggestions generated based on the user's current location information, emotional data, and past suggestion history, means for presenting the generated action suggestions, means for receiving feedback from the user, and means for generating and presenting suggestions again based on the received feedback. This makes it possible to provide personalized action suggestions that take into account the user's emotional state and location information, thereby improving user satisfaction.
[0744] A "means for receiving voice input" refers to a module for acquiring voice input from a user as digital data.
[0745] "Means for converting voice input into text data" refers to a program that converts acquired voice data into text information.
[0746] "Means of analyzing text data and understanding user intent" refers to algorithms that analyze converted text data to understand the content of user utterances and requests.
[0747] "A means of generating multiple action suggestions based on analyzed data" refers to a program that generates multiple actions or choices that a user should take, based on the results of analyzing text data.
[0748] "Means for extracting emotional states" refers to an analysis engine that automatically extracts a user's emotions from audio and video data.
[0749] "Means for adjusting action suggestions based on extracted emotional data" refers to a program that modifies generated action suggestions to a more appropriate form according to the user's emotional state.
[0750] "Means for transmitting location information" refers to a module that acquires data on the user's current location and sends it to an analysis server.
[0751] "Means for optimizing action suggestions generated based on the user's current location information, emotional data, and past suggestion history" refers to a program that generates optimal action suggestions by considering multiple factors such as the user's location information, emotional state, and past suggestion history.
[0752] "Means for presenting generated action suggestions" refers to output devices such as displays and speech synthesis modules, as well as programs, for providing generated action suggestions to the user.
[0753] "Means for receiving user feedback" refers to a module that receives responses and opinions from users regarding suggestions they have made, in the form of audio or text.
[0754] "A means of generating and presenting new suggestions based on received feedback" refers to a program that analyzes feedback received from users, generates new action suggestions based on that analysis, and presents them to the user again.
[0755] "Means of providing detailed information" refers to modules that provide users with background information and additional explanations for the generated action suggestions.
[0756] This invention provides a system that allows users to efficiently receive action suggestions in physical stores using voice input. The specific configuration and operation of this system are described in detail below.
[0757] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module, a communication module, and an emotion recognition camera. The server includes a data analysis module, a generative AI model, and a database.
[0758] Specific hardware and software examples include the following:
[0759] Speech recognition module: Google Speech-to-Text API
[0760] Emotion analysis engine: Microsoft Azure Emotion API
[0761] Generative AI: OpenAI GPT Model
[0762] Communication module: HTTP / HTTPS
[0763] Data Analysis Module: Python Scripts and SQL Databases
[0764] System operation
[0765] 1. Accepting voice input:
[0766] The user initiates voice input by speaking into a smart terminal installed in the store. The user's voice is captured through the microphone and transmitted to the system as digital data.
[0767] 2. Text conversion and analysis of audio data:
[0768] The device converts the acquired audio data into text data using the Google Speech-to-Text API. The converted text data is sent to the server, where a data analysis module analyzes the user's intent.
[0769] 3. Extraction of emotional states:
[0770] The device's built-in emotion recognition camera and audio data are used to extract the user's emotional state via the Microsoft Azure Emotion API. This emotional data is then sent to a server.
[0771] 4. Generating action proposals:
[0772] The server uses the OpenAI GPT model to generate multiple action suggestions based on analyzed text data, sentiment data, user location information, and past suggestion history. These action suggestions are adjusted according to the user's emotional state.
[0773] 5. Presenting action proposals:
[0774] The generated action suggestions are sent from the server to the terminal and presented to the user using the terminal's display or speech synthesis module. A concrete example is, "There's a new autumn book fair going on at the bookstore on the third floor."
[0775] 6. Receiving user feedback and making revisions:
[0776] When a user provides feedback on a suggested solution, the device accepts voice input again and performs re-analysis. Based on the feedback, the server regenerates multiple action suggestions under new conditions and presents them to the user again. An example of a specific prompt is: "The user is asking, 'Are there any stores closer?' Their emotional state is excited, and their current location is the first floor of a shopping mall. Please tell me about recommended stores and events."
[0777] This allows the system to provide personalized action suggestions that take into account the user's emotional state and location, thereby improving user satisfaction.
[0778] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0779] Step 1:
[0780] The user initiates voice input. The user speaks into the smart device, and the voice input is acquired through the device's microphone. The input is the user's voice data, and the output is digital voice data.
[0781] Step 2:
[0782] The device converts the audio data into text data. The acquired audio data is converted into text data using the Google Speech-to-Text API. The input is audio data, and the output is text data.
[0783] Step 3:
[0784] The terminal sends text and audio data to the server. The terminal sends the converted text data and the acquired audio data to the server. The input is text and audio data, and the output is the data sent to the server.
[0785] Step 4:
[0786] The server analyzes text data to understand the user's intent. The data analysis module analyzes text data to understand the user's requests and intent. The input is text data, and the output is the analysis result.
[0787] Step 5:
[0788] The server extracts emotional data. The audio data is analyzed for emotional state using the Microsoft Azure Emotion API, and emotional data is extracted. The input is audio data, and the output is emotional data.
[0789] Step 6:
[0790] The server generates action suggestions. The server uses the OpenAI GPT model to generate action suggestions, taking into account analysis results, sentiment data, user location information, and past suggestion history. The inputs are analysis results, sentiment data, location information, and past suggestion history, and the output is multiple action suggestions.
[0791] Step 7:
[0792] The server optimizes the action suggestions. Using generative AI and a database, it optimizes the generated action suggestions according to the user's emotional state. The input is the action suggestion, and the output is the optimized action suggestion.
[0793] Step 8:
[0794] The server sends optimized action suggestions to the terminal, which then presents them to the user. The input is the optimized action suggestions, and the output is the data sent to the terminal.
[0795] Step 9:
[0796] The device presents action suggestions. Using the device's display and speech synthesis module, the generated action suggestions are presented to the user visually and audibly. The input is optimized action suggestions, and the output is the presentation of suggestions to the user.
[0797] Step 10:
[0798] The user provides feedback on the proposal. The user inputs the feedback by voice, which is captured through the device's microphone. The input is feedback voice data, and the output is digital voice data.
[0799] Step 11:
[0800] The device converts the feedback audio data into text. The acquired feedback audio data is then converted back into text data using the Google Speech-to-Text API. The input is the feedback audio data, and the output is text data.
[0801] Step 12:
[0802] The server re-parses the feedback text data and sentiment data. The terminal sends the converted feedback text data and sentiment data to the server, which then re-parses it. The input is the feedback text data and sentiment data, and the output is the re-parsed result.
[0803] Step 13:
[0804] The server generates action suggestions again. Based on the re-analysis results, it generates action suggestions again using the OpenAI GPT model according to the new conditions. The input is the re-analysis results, and the output is the regenerated action suggestions.
[0805] Step 14:
[0806] The server optimizes the regenerated action suggestions and sends them to the terminal. The terminal then presents them to the user. The input is the regenerated action suggestions, and the output is the data sent to the terminal.
[0807] Step 15:
[0808] The terminal presents the action suggestion again. The action suggestion is presented to the user again using a display or speech synthesis module. The input is the regenerated action suggestion, and the output is the presentation of the revised suggestion to the user.
[0809] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0810] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0811] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0812] [Third Embodiment]
[0813] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0814] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0815] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0816] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0817] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0818] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0819] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0820] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0821] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0822] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0823] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0824] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0825] This invention provides a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Specific embodiments of this system are described in detail below.
[0826] System Configuration
[0827] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[0828] Program Processing Description
[0829] 1. User input reception:
[0830] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0831] The device uses a speech recognition module to convert voice input into text data.
[0832] The device sends the converted text data and the user's location information to the server.
[0833] 2. Data analysis and proposal generation:
[0834] The server analyzes the text data it receives to understand the user's intent.
[0835] The server uses generative AI to generate multiple action suggestions, taking into account location information, season, weather information, and past suggestion history.
[0836] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[0837] 3. Presentation to the user:
[0838] The terminal converts the received suggestion into speech using a speech synthesis module.
[0839] The device presents the generated audio data to the user (e.g., "How about a drive along the Shonan coast?").
[0840] 4. Accepting feedback:
[0841] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0842] The device then uses the speech recognition module again to convert the voice feedback into text data.
[0843] The terminal sends the converted feedback to the server.
[0844] 5. Reanalysis and proposals (if necessary):
[0845] The server analyzes the feedback and generates new suggestions based on the new conditions.
[0846] The server sends a revised suggestion to the terminal, and the terminal again presents the suggestion to the user via voice (e.g., "How about Hayama Beach?").
[0847] Specific example
[0848] When a user asks the device, "Where should we go today?", the device converts the voice input into text and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations, then selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device informs the user of this. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[0849] This system is designed to allow users to make decisions about their next actions with ease, using friendly language and an intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with confidence.
[0850] The following describes the processing flow.
[0851] Step 1:
[0852] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[0853] The terminal accepts voice input and activates the voice recognition module.
[0854] Step 2:
[0855] The device converts voice input into text data.
[0856] The converted text data is reviewed to determine if it is a valid request.
[0857] Step 3:
[0858] The terminal transmits the converted text data and the user's location information to the server via the communication module.
[0859] Step 4:
[0860] The server processes the received text data using a parsing module to understand the user's intent (e.g., understanding the intent "I want to go somewhere where I can see the ocean").
[0861] Step 5:
[0862] The server references the user's location information, current season, weather information, and past suggestion history, and uses generative AI to generate multiple action suggestions.
[0863] Step 6:
[0864] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[0865] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[0866] Step 7:
[0867] The server sends the selected proposal as text data to the terminal.
[0868] Step 8:
[0869] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[0870] The device generates audio data which is then presented to the user through the speaker (e.g., "How about a drive along the Shonan coast?").
[0871] Step 9:
[0872] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[0873] The device restarts its speech recognition module and converts the user's feedback into text data.
[0874] Step 10:
[0875] The device sends the converted feedback back to the server.
[0876] Step 11:
[0877] The server analyzes the feedback and generates multiple action suggestions again based on the new conditions.
[0878] Select the best option from the revised proposals (e.g., choose "Hayama Beach").
[0879] Step 12:
[0880] The server sends the re-selected proposals to the terminal as text data.
[0881] Step 13:
[0882] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[0883] Through these steps, users can quickly and intuitively decide on their next course of action.
[0884] (Example 1)
[0885] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0886] Conventional action planning systems have the problem of requiring users to search for, evaluate, and judge information themselves, which is time-consuming. Furthermore, the quality of suggestions is often inconsistent, making it difficult to accurately reflect the user's intentions. In addition, there is a lack of voice-based interfaces, hindering the intuitive and rapid development of action plans.
[0887] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0888] In this invention, the server includes means for analyzing text data and understanding the user's intent, means for generating multiple action suggestions using generative artificial intelligence based on the analyzed data, and means for selecting the optimal action suggestion and transmitting it along with detailed information. This makes it possible to accurately reflect the user's intent and provide high-quality action suggestions quickly and intuitively.
[0889] "Means for receiving voice input" refers to a combination of hardware and software for capturing the user's voice as a digital signal.
[0890] "Means for converting voice input into text data" refers to speech recognition technology that analyzes captured voice signals and converts them into corresponding text data.
[0891] "Means for transmitting location information to a server" refers to a communication module that transmits location information, which identifies the user's current location, to a server.
[0892] "Methods for analyzing text data and understanding user intent" refers to the process of analyzing text data using natural language processing technology to grasp the intent behind user requests and questions.
[0893] "Means for generating multiple action suggestions using generative artificial intelligence" refers to an algorithm that uses artificial intelligence technology to create multiple action options based on user requests.
[0894] The "means for selecting the most suitable action proposal and sending it along with detailed information" refers to a function that selects the most appropriate action proposal from those generated and sends its detailed information (e.g., access methods and local conditions) to the terminal.
[0895] "Means for converting generated action suggestions into speech and presenting them to the user" refers to speech synthesis technology and output devices that convert text data into speech data and present suggestions to the user visually and aurally.
[0896] "A means of receiving user feedback, converting audio into text data, and sending it to a server" refers to a process that receives user feedback as voice input, converts it into text, and transmits it to a server.
[0897] "A means of generating new suggestions based on received feedback and presenting them in audio format" refers to a process of generating new action suggestions based on user feedback, converting them back into audio data, and presenting them to the user.
[0898] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[0899] System Configuration
[0900] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[0901] Hardware and software to be used
[0902] The device includes a microphone for receiving voice input, a speech recognition module (e.g., Google Cloud Speech-to-Text API, Amazon Transcribe, etc.) for converting speech to text data, and a GPS module for obtaining the user's location information. A communication module is used to send the converted text data and location information to the server. On the server side, natural language processing technology (e.g., BERT, GPT-3, etc.) is used as a data analysis module to understand the user's intent. Furthermore, generative AI (e.g., OpenAI's GPT-3, etc.) is used to generate multiple action suggestions. After the server selects the optimal action suggestion, it sends it to the device along with detailed information (such as how to access it and the local weather). The device uses a speech synthesis module (e.g., Google Text-to-Speech API, Amazon Polly, etc.) to convert the suggested content into speech and present it to the user.
[0903] Specific example
[0904] As a concrete example, suppose a user asks their device, "Where should we go today?" The device converts the voice input into text data and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations. From these, it selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device communicates this to the user. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[0905] As described above, this system is designed to allow users to make decisions about their next actions in an enjoyable way through its user-friendly language and intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with peace of mind.
[0906] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0907] Step 1:
[0908] User input reception
[0909] The user speaks into the device to ask a question or make a request (e.g., "Where should we go today?"). The device's microphone collects the voice data and uses it as input for a speech recognition module (such as Google Cloud Speech-to-Text API or Amazon Transcribe). The speech recognition module converts the voice data into text data, generating a text-based input. This text data includes the user's intent or inquiry.
[0910] Step 2:
[0911] Acquiring and transmitting location information
[0912] The device uses a GPS module to determine the user's current location. The acquired location information is sent to the server along with text data. Data communication is conducted using the HTTPS protocol during transmission, ensuring data security.
[0913] Step 3:
[0914] Data Analysis
[0915] The server analyzes the received text data and location information. It uses natural language processing models (e.g., BERT or GPT-3) to analyze the text data and understand the user's intent. This process extracts keywords and contextual information from the text data. Based on the analysis results, the server generates data that reflects the user's wishes and requests.
[0916] Step 4:
[0917] Generating action proposals
[0918] The server generates multiple action suggestions based on the analysis results using a generative AI (e.g., OpenAI's GPT-3). The server also takes into account external data such as the user's location, season, and weather information (e.g., using the OpenWeatherMap API). This improves the accuracy and appropriateness of the suggestions. The generative AI generates multiple candidate suggestions and adds detailed information to each suggestion.
[0919] Step 5:
[0920] Selecting and sending the most suitable proposal
[0921] The server uses evaluation criteria (e.g., distance, weather conditions, event information, etc.) to select the most suitable action from the generated suggestions. It then sends the best suggestion and its detailed information (such as access methods) to the terminal. The server uses asynchronous communication to transmit data quickly and efficiently.
[0922] Step 6:
[0923] Presentation to the user
[0924] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Google Text-to-Speech API or Amazon Polly). The generated speech data is then presented to the user through the speaker (e.g., "How about a drive to Shonan Beach?"). Simultaneously, displaying the suggestions in text format on the screen is also being considered.
[0925] Step 7:
[0926] Accepting feedback
[0927] The user provides voice feedback on the suggested location (e.g., "Is there a place a little closer?"). The device uses a voice recognition module to convert the voice feedback into text data. This text data is then sent back to the server for further analysis.
[0928] Step 8:
[0929] Generation and presentation of revised proposals
[0930] The server analyzes the feedback and generates new suggestions based on the new conditions. Generative AI is used again in this re-suggestion to create new, appropriate candidates. The server sends detailed information about the re-suggestion to the terminal, which then presents it to the user again (e.g., "How about Hayama Beach?"). This allows the system to repeat suggestions until the user is satisfied.
[0931] Through this series of processes, users can intuitively and efficiently plan their actions.
[0932] (Application Example 1)
[0933] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0934] In modern society, efficiently and intuitively planning time with family and friends is crucial for users to make the most of their limited time. However, users often have to go through the trouble of translating their desires into concrete and appropriate action suggestions, and obtaining detailed information to judge the appropriateness of those suggestions. Furthermore, the lack of means to properly make reservations or orders for the suggested actions can hinder the smooth execution of those suggestions. There is a need for a system that can solve this problem.
[0935] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0936] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data and understanding the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, and means for making reservations or orders related to the suggested actions. This allows the user to easily receive action suggestions through voice input, provide feedback to obtain the most suitable suggestion, and smoothly transition to actual actions based on the suggestions.
[0937] "Means for receiving voice input" refers to hardware or software that receives voice data and converts it into a format that can be processed within the system.
[0938] "Means for converting received voice input into text data" refers to software or hardware that uses speech recognition technology to convert a user's voice input into corresponding text data.
[0939] "Means for analyzing converted text data and understanding user intent" refers to software or hardware that uses natural language processing techniques to analyze the content of text data and understand what the user is asking for.
[0940] "Means for generating multiple action suggestions based on analyzed data" refers to an algorithm for generating multiple action options for the user based on the analysis results, and the hardware or software that executes it.
[0941] "Means for presenting generated action suggestions" refers to devices and software that present generated action options to the user, either visually or audibly.
[0942] "Means for receiving user feedback" refers to the hardware and software used to receive and incorporate user responses to suggestions into the system.
[0943] "Means for generating and presenting new suggestions based on received feedback" refers to algorithms and devices that generate and present new action suggestions based on user feedback.
[0944] "Means of making reservations or orders related to suggested actions" refers to online services for making restaurant reservations or ordering goods based on actions suggested within the system, as well as the software and hardware for executing them.
[0945] "Means of transmitting location information" refers to a GPS module and communication module used to determine the user's current location and transmit it to a server.
[0946] This invention provides a system for users to plan meals efficiently and intuitively. The system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[0947] System Configuration
[0948] This system includes means for receiving voice input, means for converting voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, means for making reservations or orders related to the suggested actions, and means for transmitting location information.
[0949] Processing flow
[0950] 1. Accepting voice input
[0951] Users input their meal requests by voice into a device such as a smartphone. For example, they might say, "I'd like to eat some delicious pizza nearby."
[0952] The device is equipped with a microphone and accepts voice input.
[0953] 2. Text conversion of audio data
[0954] The device uses a speech recognition module (e.g., Google Speech Recognition) to convert the audio data into text data.
[0955] 3. Acquisition of location information
[0956] The device uses a GPS module to obtain the user's current location.
[0957] 4. Sending data
[0958] The device sends the acquired text data and location information to the server. A communication module is used.
[0959] 5. Data Analysis and Proposal Generation
[0960] The server analyzes text data and uses natural language processing techniques to understand the user's intent.
[0961] The server uses generative AI (e.g., GPT-3) to generate multiple action suggestions (e.g., restaurant or cafe selection) considering the user's intent, location, season, weather information, etc.
[0962] 6. Presentation of Action Proposals
[0963] The server sends the generated action suggestion to the device. The device then presents the suggestion to the user via text or voice (e.g., "I'd like to recommend John's Pizza nearby. Would you like to make a reservation?").
[0964] 7. Accepting Feedback
[0965] Users provide feedback on the suggested locations, such as "Is there a closer location?"
[0966] The device converts the feedback into text data and sends it to the server.
[0967] 8. Reanalysis and Re-proposal
[0968] The server analyzes the feedback and generates new suggestions based on the new conditions.
[0969] The revised suggestion is sent to the device, and the device presents the suggestion to the user again (e.g., "How about 'Erika's Pizza,' which is a 5-minute walk from there?").
[0970] 9. Execute reservation / order
[0971] If the user agrees to the suggestions, they can make restaurant reservations or order delivery from their device. The server then sends the reservation and order information to the online service to complete the process.
[0972] Specific example
[0973] When a user speaks to a smartphone app saying, "I want to eat some delicious pizza nearby," the app converts the speech to text, retrieves the user's location information, and sends it to a server. The server uses generative AI (e.g., GPT-3) to suggest the best pizza restaurant based on the user's location and preferences. One option presented is "John's Pizza - 5 minutes away," and the app responds with a voice message saying, "I'd like to introduce you to a nearby 'John's Pizza.' Would you like to make a reservation?"
[0974] Example of a prompt
[0975] User: "I want to eat some delicious pizza nearby."
[0976] system:
[0977] Speech recognition: Converts speech to text → "I want to eat some delicious pizza nearby."
[0978] Location information acquisition: Obtain current location from GPS.
[0979] Server: Receives text and location information, and generates suggestions using generative AI → "John's Pizza"
[0980] Response: Voice response from the speech generation module → "I'd like to recommend a nearby 'John's Pizza'. Would you like to make a reservation?"
[0981] Thus, by using the system of the present invention, users can not only easily perform voice input and receive appropriate suggestions, but also smoothly carry out actions based on those suggestions.
[0982] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0983] Step 1:
[0984] The user inputs their dining requests via voice. Using their smartphone's microphone, the user might say, "I want to eat some delicious pizza nearby." This voice data becomes the input.
[0985] Step 2:
[0986] The device converts the voice data into text data. A speech recognition module (e.g., Google Speech Recognition) analyzes the voice and outputs the text data "I want to eat some delicious pizza nearby." This text data becomes the input for the next processing step.
[0987] Step 3:
[0988] The device acquires its current location information. Using a GPS module, it measures the user's current location and obtains latitude and longitude information. This location information becomes the input for the next processing step.
[0989] Step 4:
[0990] The terminal sends text data and location information to the server. The text data and location information are passed to the server via a communication module. This text data and location information becomes the input for the next processing step.
[0991] Step 5:
[0992] The server analyzes the text data to understand the user's intent. Using natural language processing techniques, it analyzes the received text data and outputs the user's intent, "I want recommendations for nearby pizza restaurants." This intent becomes the input for the next processing step.
[0993] Step 6:
[0994] The server generates multiple action suggestions based on the user's intent and location information. Using a generative AI (e.g., GPT-3), it generates several pizza restaurant candidates, taking into account the user's intent, location information, and additional information such as weather and time of day. These action suggestions become the input for the next processing step.
[0995] Step 7:
[0996] The server selects the most suitable action suggestion and sends it to the terminal along with detailed information. The terminal then selects the most appropriate restaurant from the multiple suggestions and sends it along with its details (address, opening hours, menu, etc.). This becomes the input for the next processing step.
[0997] Step 8:
[0998] The device presents the user with an action suggestion. Using a speech synthesis module that converts text data into speech, it suggests to the user verbally, "I'd like to recommend a nearby 'John's Pizza.' Would you like to make a reservation?" This suggestion becomes the input for the next processing step.
[0999] Step 9:
[1000] The user provides feedback. They give voice feedback on the suggested location, such as "Is there a closer location?" This voice data becomes the input for the next processing step.
[1001] Step 10:
[1002] The device converts the user's voice feedback into text data. The voice recognition module is used again to convert the feedback into text data. This feedback becomes the input for the next processing step.
[1003] Step 11:
[1004] The terminal sends feedback text data to the server. The feedback text data is passed to the server via the communication module. This feedback becomes the input for the next processing step.
[1005] Step 12:
[1006] The server analyzes the feedback and generates new action suggestions. The generative AI is used again to generate new candidates based on the feedback and new conditions. These new action suggestions become the input for the next processing step.
[1007] Step 13:
[1008] The server sends a revised suggestion to the terminal, which then presents the suggestion to the user again. It then makes another voice suggestion, such as, "How about 'Erika's Pizza,' which is a 5-minute walk from there?" This process is repeated until the user is satisfied.
[1009] Step 14:
[1010] If the user agrees to the final action suggestion, they can make a reservation or order delivery from the relevant restaurant using their device. The device uses a communication module to send the reservation or order information to the server, which then transmits this information to the online service to complete the process. This reservation or order becomes the final output.
[1011] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1012] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, the invention aims to provide users with more personalized suggestions by incorporating an emotion engine that recognizes user emotions. Specific embodiments of this system are described in detail below.
[1013] System Configuration
[1014] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[1015] Program Processing Description
[1016] 1. User input reception:
[1017] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1018] The terminal accepts voice input and activates the voice recognition module.
[1019] 2. Emotional analysis of voice data:
[1020] The device sends voice input to the emotion engine.
[1021] The emotion engine analyzes the audio data and extracts emotional states (e.g., joy, excitement, fatigue, etc.).
[1022] The extracted emotion data is sent back to the device.
[1023] 3. Convert to text data:
[1024] The device converts voice input into text data.
[1025] The converted text data is reviewed to determine if it is a valid request.
[1026] 4. Sending data to the server:
[1027] The device sends the converted text data, user location information, and sentiment data to the server.
[1028] 5. Data analysis and proposal generation:
[1029] The server analyzes the text data it receives to understand the user's intent.
[1030] The server uses generative AI to generate multiple action suggestions, taking into account location information, current season, weather information, sentiment data, and past suggestion history.
[1031] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[1032] 6. Presentation to the user:
[1033] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[1034] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[1035] 7. Accepting feedback:
[1036] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[1037] The device then uses the speech recognition module again to convert the voice feedback into text data.
[1038] 8. Data retransmission and reanalysis:
[1039] The device then sends the converted feedback and sentiment data back to the server.
[1040] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[1041] The server selects the most suitable proposal from the revised proposals and sends the selected proposal to the terminal.
[1042] 9. Final presentation:
[1043] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[1044] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[1045] The following describes the processing flow.
[1046] System Configuration
[1047] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[1048] Program processing steps
[1049] Step 1:
[1050] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1051] The terminal accepts voice input and activates the voice recognition module.
[1052] Step 2:
[1053] The device sends voice data to the emotion engine.
[1054] The emotion engine analyzes the audio data and extracts the user's emotional state (e.g., joy, excitement, fatigue, etc.).
[1055] The extracted emotion data is sent back to the device.
[1056] Step 3:
[1057] The device converts voice input into text data.
[1058] Review the converted text data to determine if it is a valid request.
[1059] Step 4:
[1060] The device transmits the converted text data, user location information, and sentiment data to the server via a communication module.
[1061] Step 5:
[1062] The server analyzes the text data it receives to understand the user's intent.
[1063] The server uses generative AI to generate multiple action suggestions based on location information, current season, weather information, sentiment data, and past suggestion history.
[1064] Step 6:
[1065] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[1066] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[1067] Step 7:
[1068] The server sends the selected proposal as text data to the terminal.
[1069] Step 8:
[1070] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[1071] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[1072] Step 9:
[1073] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[1074] The device then uses the speech recognition module again to convert the voice feedback into text data.
[1075] Step 10:
[1076] The device then sends the converted feedback and sentiment data back to the server.
[1077] Step 11:
[1078] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[1079] The server selects the best option from the alternative suggestions (e.g., "Hayama Beach").
[1080] Step 12:
[1081] The server sends the re-selected proposals to the terminal as text data.
[1082] Step 13:
[1083] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[1084] Through these steps, users can quickly and intuitively decide on their next course of action. By using an emotion engine, personalized suggestions that align with the user's feelings become possible, further improving user satisfaction.
[1085] (Example 2)
[1086] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1087] Conventional behavior suggestion systems often fail to adequately consider users' emotional states and offer only uniform suggestions, resulting in low user satisfaction. Therefore, there is a need to develop a system that provides individually optimized suggestions based on the user's emotional state, thereby increasing user satisfaction. Furthermore, the ability to quickly respond to received feedback and offer revised suggestions is also required.
[1088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for performing sentiment analysis on received voice data, means for extracting the user's emotional state based on the sentiment analysis, and means for analyzing the converted text data and sentiment data to understand the user's intent. This enables highly personalized suggestions based on sentiment analysis.
[1089] "Voice input" is a method of input that allows users to communicate questions or requests to a system using their voice.
[1090] "Text data" refers to data obtained by analyzing voice input and converting it into text information.
[1091] "Emotional analysis" is the process of extracting a user's emotional state (e.g., joy, excitement, fatigue, etc.) from audio data.
[1092] "Emotional data" refers to data that indicates the user's emotional state, obtained as a result of emotion analysis.
[1093] "Understanding user intent" means analyzing converted text and sentiment data to grasp what the user wants.
[1094] A "generative AI model" is an artificial intelligence algorithm that generates appropriate action suggestions based on given data.
[1095] "Action suggestions" refer to ideas and plans for proposing actions and activities that are appropriate for the user.
[1096] "Feedback" refers to users providing feedback, opinions, and requests regarding suggestions to the system.
[1097] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, the invention aims to provide users with more personalized suggestions.
[1098] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI model. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[1099] First, the system starts working when the user inputs a question or request via voice into the device. For example, a voice input might be, "Where should we go today?" The device receives the voice input, activates a speech recognition module (such as the Google Cloud Speech-to-Text API), and converts the voice into text data. This conversion may take several seconds.
[1100] Next, the terminal uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the voice data and extract the user's emotional state. The extracted emotion data is sent back to the terminal and then transmitted to the server along with the converted text data. The server receives this data and uses natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to analyze the user's intent.
[1101] The server further analyzes location information, current season, weather information, and past behavioral history, and uses a generative AI model (e.g., OpenAI GPT-3) to generate multiple action suggestions. It selects the most suitable suggestion from the generated suggestions and sends it back to the terminal. The terminal converts the received suggestion into audio data using a speech synthesis module (e.g., Amazon Polly) and presents it to the user through its speaker.
[1102] The user provides feedback on the presented suggestion (e.g., "Is there a place a little closer?"). The device uses the speech recognition module again to convert the voice feedback into text data and sends it back to the server. The server analyzes the feedback and sentiment data, generates multiple action suggestions again based on the new conditions, and sends the best suggestion back to the device. Finally, the device uses the speech synthesis module to convert the received suggestion into voice data and presents it to the user.
[1103] As a concrete example, the following prompt sentences could be input into the generation AI model:
[1104] Prompt: "Please suggest some family-friendly outing destinations. Users are happy and prefer nearby locations."
[1105] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[1106] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1107] Step 1:
[1108] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1109] The device accepts voice input and activates a speech recognition module (e.g., Google Cloud Speech-to-Text API). The voice data is converted into text data. The input is the user's voice data, and the output is text data.
[1110] Step 2:
[1111] The terminal sends the converted text data to an emotion engine (e.g., IBM Watson Tone Analyzer). The emotion engine analyzes the audio data and extracts the user's emotional state. The input is audio data, and the output is emotion data.
[1112] Step 3:
[1113] The device verifies the sentiment data and text data returned from the sentiment engine. It performs text verification to determine if the request is appropriate. The input is sentiment data and text data, and the output is the verified text data and sentiment data.
[1114] Step 4:
[1115] The device sends verified text data, user location information (e.g., GPS data), and sentiment data to the server. This communication typically uses the HTTP or HTTPS protocol. Inputs are text data, location information, and sentiment data, while output is a notification to the server that the transmission is complete.
[1116] Step 5:
[1117] The server analyzes the received text data using natural language processing algorithms (e.g., SpaCy or NLTK) to understand the user's intent. The input is text data, and the output is user intent data.
[1118] Step 6:
[1119] The server considers the user's location, current season, weather information, and sentiment data, and also accesses a database of past behavioral history. It generates multiple action suggestions using a generative AI model (e.g., OpenAI GPT-3). The input is location information, season, weather information, sentiment data, and past behavioral history, and the output is a list of action suggestions.
[1120] Step 7:
[1121] The server uses a machine learning algorithm to select the most suitable action from the generated suggestions. The selected action suggestion, along with detailed information, is then sent to the terminal. The input is a list of action suggestions, and the output is the optimal action suggestion and its details.
[1122] Step 8:
[1123] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Amazon Polly). This data is then presented to the user through the speaker. The input consists of optimal action suggestions and detailed information, while the output is speech data.
[1124] Step 9:
[1125] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?"). The device then uses its speech recognition module again to convert the voice feedback into text data. The input is the user's voice data, and the output is the text data of the feedback.
[1126] Step 10:
[1127] The device then resends the converted feedback and sentiment data to the server. The inputs are the feedback text data and sentiment data, and the output is a notification that the transmission to the server is complete.
[1128] Step 11:
[1129] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions. The input is feedback text data and sentiment data, and the output is a list of new action suggestions.
[1130] Step 12:
[1131] The server selects the most suitable action from the regenerated suggestions and sends the selected suggestion to the terminal. The input is a list of new action suggestions, and the output is the re-selected action suggestion.
[1132] Step 13:
[1133] The terminal receives the suggested content again, converts it into speech data using a speech synthesis module, and presents it to the user. The input is the re-selected action suggestion, and the output is the speech data.
[1134] (Application Example 2)
[1135] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1136] Conventional behavior suggestion systems have difficulty taking into account the user's emotional state and detailed location information, resulting in suggestions that are not sufficiently personalized. Furthermore, the process of regenerating suggestions based on user voice feedback is inefficient, making it difficult to increase user satisfaction.
[1137] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for extracting the emotional state and adjusting the action suggestions based on the extracted emotional data, means for optimizing the action suggestions generated based on the user's current location information, emotional data, and past suggestion history, means for presenting the generated action suggestions, means for receiving feedback from the user, and means for generating and presenting suggestions again based on the received feedback. This makes it possible to provide personalized action suggestions that take into account the user's emotional state and location information, thereby improving user satisfaction.
[1138] A "means for receiving voice input" refers to a module for acquiring voice input from a user as digital data.
[1139] "Means for converting voice input into text data" refers to a program that converts acquired voice data into text information.
[1140] "Means of analyzing text data and understanding user intent" refers to algorithms that analyze converted text data to understand the content of user utterances and requests.
[1141] "A means of generating multiple action suggestions based on analyzed data" refers to a program that generates multiple actions or choices that a user should take, based on the results of analyzing text data.
[1142] "Means for extracting emotional states" refers to an analysis engine that automatically extracts a user's emotions from audio and video data.
[1143] "Means for adjusting action suggestions based on extracted emotional data" refers to a program that modifies generated action suggestions to a more appropriate form according to the user's emotional state.
[1144] "Means for transmitting location information" refers to a module that acquires data on the user's current location and sends it to an analysis server.
[1145] "Means for optimizing action suggestions generated based on the user's current location information, emotional data, and past suggestion history" refers to a program that generates optimal action suggestions by considering multiple factors such as the user's location information, emotional state, and past suggestion history.
[1146] "Means for presenting generated action suggestions" refers to output devices such as displays and speech synthesis modules, as well as programs, for providing generated action suggestions to the user.
[1147] "Means for receiving user feedback" refers to a module that receives responses and opinions from users regarding suggestions they have made, in the form of audio or text.
[1148] "A means of generating and presenting new suggestions based on received feedback" refers to a program that analyzes feedback received from users, generates new action suggestions based on that analysis, and presents them to the user again.
[1149] "Means of providing detailed information" refers to modules that provide users with background information and additional explanations for the generated action suggestions.
[1150] This invention provides a system that allows users to efficiently receive action suggestions in physical stores using voice input. The specific configuration and operation of this system are described in detail below.
[1151] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module, a communication module, and an emotion recognition camera. The server includes a data analysis module, a generative AI model, and a database.
[1152] Specific hardware and software examples include the following:
[1153] Speech recognition module: Google Speech-to-Text API
[1154] Emotion analysis engine: Microsoft Azure Emotion API
[1155] Generative AI: OpenAI GPT Model
[1156] Communication module: HTTP / HTTPS
[1157] Data Analysis Module: Python Scripts and SQL Databases
[1158] System operation
[1159] 1. Accepting voice input:
[1160] The user initiates voice input by speaking into a smart terminal installed in the store. The user's voice is captured through the microphone and transmitted to the system as digital data.
[1161] 2. Text conversion and analysis of audio data:
[1162] The device converts the acquired audio data into text data using the Google Speech-to-Text API. The converted text data is sent to the server, where a data analysis module analyzes the user's intent.
[1163] 3. Extraction of emotional states:
[1164] The device's built-in emotion recognition camera and audio data are used to extract the user's emotional state via the Microsoft Azure Emotion API. This emotional data is then sent to a server.
[1165] 4. Generating action proposals:
[1166] The server uses the OpenAI GPT model to generate multiple action suggestions based on analyzed text data, sentiment data, user location information, and past suggestion history. These action suggestions are adjusted according to the user's emotional state.
[1167] 5. Presenting action proposals:
[1168] The generated action suggestions are sent from the server to the terminal and presented to the user using the terminal's display or speech synthesis module. A concrete example is, "There's a new autumn book fair going on at the bookstore on the third floor."
[1169] 6. Receiving user feedback and making revisions:
[1170] When a user provides feedback on a suggested solution, the device accepts voice input again and performs re-analysis. Based on the feedback, the server regenerates multiple action suggestions under new conditions and presents them to the user again. An example of a specific prompt is: "The user is asking, 'Are there any stores closer?' Their emotional state is excited, and their current location is the first floor of a shopping mall. Please tell me about recommended stores and events."
[1171] This allows the system to provide personalized action suggestions that take into account the user's emotional state and location, thereby improving user satisfaction.
[1172] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1173] Step 1:
[1174] The user initiates voice input. The user speaks into the smart device, and the voice input is acquired through the device's microphone. The input is the user's voice data, and the output is digital voice data.
[1175] Step 2:
[1176] The device converts the audio data into text data. The acquired audio data is converted into text data using the Google Speech-to-Text API. The input is audio data, and the output is text data.
[1177] Step 3:
[1178] The terminal sends text and audio data to the server. The terminal sends the converted text data and the acquired audio data to the server. The input is text and audio data, and the output is the data sent to the server.
[1179] Step 4:
[1180] The server analyzes text data to understand the user's intent. The data analysis module analyzes text data to understand the user's requests and intent. The input is text data, and the output is the analysis result.
[1181] Step 5:
[1182] The server extracts emotional data. The audio data is analyzed for emotional state using the Microsoft Azure Emotion API, and emotional data is extracted. The input is audio data, and the output is emotional data.
[1183] Step 6:
[1184] The server generates action suggestions. The server uses the OpenAI GPT model to generate action suggestions, taking into account analysis results, sentiment data, user location information, and past suggestion history. The inputs are analysis results, sentiment data, location information, and past suggestion history, and the output is multiple action suggestions.
[1185] Step 7:
[1186] The server optimizes the action suggestions. Using generative AI and a database, it optimizes the generated action suggestions according to the user's emotional state. The input is the action suggestion, and the output is the optimized action suggestion.
[1187] Step 8:
[1188] The server sends optimized action suggestions to the terminal, which then presents them to the user. The input is the optimized action suggestions, and the output is the data sent to the terminal.
[1189] Step 9:
[1190] The device presents action suggestions. Using the device's display and speech synthesis module, the generated action suggestions are presented to the user visually and audibly. The input is optimized action suggestions, and the output is the presentation of suggestions to the user.
[1191] Step 10:
[1192] The user provides feedback on the proposal. The user inputs the feedback by voice, which is captured through the device's microphone. The input is feedback voice data, and the output is digital voice data.
[1193] Step 11:
[1194] The device converts the feedback audio data into text. The acquired feedback audio data is then converted back into text data using the Google Speech-to-Text API. The input is the feedback audio data, and the output is text data.
[1195] Step 12:
[1196] The server re-parses the feedback text data and sentiment data. The terminal sends the converted feedback text data and sentiment data to the server, which then re-parses it. The input is the feedback text data and sentiment data, and the output is the re-parsed result.
[1197] Step 13:
[1198] The server generates action suggestions again. Based on the re-analysis results, it generates action suggestions again using the OpenAI GPT model according to the new conditions. The input is the re-analysis results, and the output is the regenerated action suggestions.
[1199] Step 14:
[1200] The server optimizes the regenerated action suggestions and sends them to the terminal. The terminal then presents them to the user. The input is the regenerated action suggestions, and the output is the data sent to the terminal.
[1201] Step 15:
[1202] The terminal presents the action suggestion again. The action suggestion is presented to the user again using a display or speech synthesis module. The input is the regenerated action suggestion, and the output is the presentation of the revised suggestion to the user.
[1203] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1204] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1205] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1206] [Fourth Embodiment]
[1207] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1208] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1209] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1210] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1211] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1212] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1213] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1214] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1215] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1216] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1217] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1218] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1219] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1220] This invention provides a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Specific embodiments of this system are described in detail below.
[1221] System Configuration
[1222] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[1223] Program Processing Description
[1224] 1. User input reception:
[1225] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1226] The device uses a speech recognition module to convert voice input into text data.
[1227] The device sends the converted text data and the user's location information to the server.
[1228] 2. Data analysis and proposal generation:
[1229] The server analyzes the text data it receives to understand the user's intent.
[1230] The server uses generative AI to generate multiple action suggestions, taking into account location information, season, weather information, and past suggestion history.
[1231] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[1232] 3. Presentation to the user:
[1233] The terminal converts the received suggestion into speech using a speech synthesis module.
[1234] The device presents the generated audio data to the user (e.g., "How about a drive along the Shonan coast?").
[1235] 4. Accepting feedback:
[1236] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[1237] The device then uses the speech recognition module again to convert the voice feedback into text data.
[1238] The terminal sends the converted feedback to the server.
[1239] 5. Reanalysis and proposals (if necessary):
[1240] The server analyzes the feedback and generates new suggestions based on the new conditions.
[1241] The server sends a revised suggestion to the terminal, and the terminal again presents the suggestion to the user via voice (e.g., "How about Hayama Beach?").
[1242] Specific example
[1243] When a user asks the device, "Where should we go today?", the device converts the voice input into text and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations, then selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device informs the user of this. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[1244] This system is designed to allow users to make decisions about their next actions with ease, using friendly language and an intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with confidence.
[1245] The following describes the processing flow.
[1246] Step 1:
[1247] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1248] The terminal accepts voice input and activates the voice recognition module.
[1249] Step 2:
[1250] The device converts voice input into text data.
[1251] The converted text data is reviewed to determine if it is a valid request.
[1252] Step 3:
[1253] The terminal transmits the converted text data and the user's location information to the server via the communication module.
[1254] Step 4:
[1255] The server processes the received text data using a parsing module to understand the user's intent (e.g., understanding the intent "I want to go somewhere where I can see the ocean").
[1256] Step 5:
[1257] The server references the user's location information, current season, weather information, and past suggestion history, and uses generative AI to generate multiple action suggestions.
[1258] Step 6:
[1259] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[1260] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[1261] Step 7:
[1262] The server sends the selected proposal as text data to the terminal.
[1263] Step 8:
[1264] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[1265] The device generates audio data which is then presented to the user through the speaker (e.g., "How about a drive along the Shonan coast?").
[1266] Step 9:
[1267] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[1268] The device restarts its speech recognition module and converts the user's feedback into text data.
[1269] Step 10:
[1270] The device sends the converted feedback back to the server.
[1271] Step 11:
[1272] The server analyzes the feedback and generates multiple action suggestions again based on the new conditions.
[1273] Select the best option from the revised proposals (e.g., choose "Hayama Beach").
[1274] Step 12:
[1275] The server sends the re-selected proposals to the terminal as text data.
[1276] Step 13:
[1277] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[1278] Through these steps, users can quickly and intuitively decide on their next course of action.
[1279] (Example 1)
[1280] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1281] Conventional action planning systems have the problem of requiring users to search for, evaluate, and judge information themselves, which is time-consuming. Furthermore, the quality of suggestions is often inconsistent, making it difficult to accurately reflect the user's intentions. In addition, there is a lack of voice-based interfaces, hindering the intuitive and rapid development of action plans.
[1282] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1283] In this invention, the server includes means for analyzing text data and understanding the user's intent, means for generating multiple action suggestions using generative artificial intelligence based on the analyzed data, and means for selecting the optimal action suggestion and transmitting it along with detailed information. This makes it possible to accurately reflect the user's intent and provide high-quality action suggestions quickly and intuitively.
[1284] "Means for receiving voice input" refers to a combination of hardware and software for capturing the user's voice as a digital signal.
[1285] "Means for converting voice input into text data" refers to speech recognition technology that analyzes captured voice signals and converts them into corresponding text data.
[1286] "Means for transmitting location information to a server" refers to a communication module that transmits location information, which identifies the user's current location, to a server.
[1287] "Methods for analyzing text data and understanding user intent" refers to the process of analyzing text data using natural language processing technology to grasp the intent behind user requests and questions.
[1288] "Means for generating multiple action suggestions using generative artificial intelligence" refers to an algorithm that uses artificial intelligence technology to create multiple action options based on user requests.
[1289] The "means for selecting the most suitable action proposal and sending it along with detailed information" refers to a function that selects the most appropriate action proposal from those generated and sends its detailed information (e.g., access methods and local conditions) to the terminal.
[1290] "Means for converting generated action suggestions into speech and presenting them to the user" refers to speech synthesis technology and output devices that convert text data into speech data and present suggestions to the user visually and aurally.
[1291] "A means of receiving user feedback, converting audio into text data, and sending it to a server" refers to a process that receives user feedback as voice input, converts it into text, and transmits it to a server.
[1292] "A means of generating new suggestions based on received feedback and presenting them in audio format" refers to a process of generating new action suggestions based on user feedback, converting them back into audio data, and presenting them to the user.
[1293] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[1294] System Configuration
[1295] This system consists of a terminal that accepts voice input and a server that analyzes the data. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI.
[1296] Hardware and software to be used
[1297] The device includes a microphone for receiving voice input, a speech recognition module (e.g., Google Cloud Speech-to-Text API, Amazon Transcribe, etc.) for converting speech to text data, and a GPS module for obtaining the user's location information. A communication module is used to send the converted text data and location information to the server. On the server side, natural language processing technology (e.g., BERT, GPT-3, etc.) is used as a data analysis module to understand the user's intent. Furthermore, generative AI (e.g., OpenAI's GPT-3, etc.) is used to generate multiple action suggestions. After the server selects the optimal action suggestion, it sends it to the device along with detailed information (such as how to access it and the local weather). The device uses a speech synthesis module (e.g., Google Text-to-Speech API, Amazon Polly, etc.) to convert the suggested content into speech and present it to the user.
[1298] Specific example
[1299] As a concrete example, suppose a user asks their device, "Where should we go today?" The device converts the voice input into text data and sends it to the server along with the user's location information. The server analyzes the data and generates several candidate locations. From these, it selects "Shonan Coast" as the best suggestion and sends it to the device along with detailed information (how to get there, local weather, etc.). The device then suggests to the user via voice, "How about a drive to Shonan Coast?" If the user provides feedback, "Isn't there somewhere a little closer?", the server suggests "Hayama Coast" as a new option, and the device communicates this to the user. Through this series of operations, the user can decide on their next course of action appropriately and quickly.
[1300] As described above, this system is designed to allow users to make decisions about their next actions in an enjoyable way through its user-friendly language and intuitive interface. Furthermore, by providing location information and detailed information, users can plan their actions with peace of mind.
[1301] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1302] Step 1:
[1303] User input reception
[1304] The user speaks into the device to ask a question or make a request (e.g., "Where should we go today?"). The device's microphone collects the voice data and uses it as input for a speech recognition module (such as Google Cloud Speech-to-Text API or Amazon Transcribe). The speech recognition module converts the voice data into text data, generating a text-based input. This text data includes the user's intent or inquiry.
[1305] Step 2:
[1306] Acquiring and transmitting location information
[1307] The device uses a GPS module to determine the user's current location. The acquired location information is sent to the server along with text data. Data communication is conducted using the HTTPS protocol during transmission, ensuring data security.
[1308] Step 3:
[1309] Data Analysis
[1310] The server analyzes the received text data and location information. It uses natural language processing models (e.g., BERT or GPT-3) to analyze the text data and understand the user's intent. This process extracts keywords and contextual information from the text data. Based on the analysis results, the server generates data that reflects the user's wishes and requests.
[1311] Step 4:
[1312] Generating action proposals
[1313] The server generates multiple action suggestions based on the analysis results using a generative AI (e.g., OpenAI's GPT-3). The server also takes into account external data such as the user's location, season, and weather information (e.g., using the OpenWeatherMap API). This improves the accuracy and appropriateness of the suggestions. The generative AI generates multiple candidate suggestions and adds detailed information to each suggestion.
[1314] Step 5:
[1315] Selecting and sending the most suitable proposal
[1316] The server uses evaluation criteria (e.g., distance, weather conditions, event information, etc.) to select the most suitable action from the generated suggestions. It then sends the best suggestion and its detailed information (such as access methods) to the terminal. The server uses asynchronous communication to transmit data quickly and efficiently.
[1317] Step 6:
[1318] Presentation to the user
[1319] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Google Text-to-Speech API or Amazon Polly). The generated speech data is then presented to the user through the speaker (e.g., "How about a drive to Shonan Beach?"). Simultaneously, displaying the suggestions in text format on the screen is also being considered.
[1320] Step 7:
[1321] Accepting feedback
[1322] The user provides voice feedback on the suggested location (e.g., "Is there a place a little closer?"). The device uses a voice recognition module to convert the voice feedback into text data. This text data is then sent back to the server for further analysis.
[1323] Step 8:
[1324] Generation and presentation of revised proposals
[1325] The server analyzes the feedback and generates new suggestions based on the new conditions. Generative AI is used again in this re-suggestion to create new, appropriate candidates. The server sends detailed information about the re-suggestion to the terminal, which then presents it to the user again (e.g., "How about Hayama Beach?"). This allows the system to repeat suggestions until the user is satisfied.
[1326] Through this series of processes, users can intuitively and efficiently plan their actions.
[1327] (Application Example 1)
[1328] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1329] In modern society, efficiently and intuitively planning time with family and friends is crucial for users to make the most of their limited time. However, users often have to go through the trouble of translating their desires into concrete and appropriate action suggestions, and obtaining detailed information to judge the appropriateness of those suggestions. Furthermore, the lack of means to properly make reservations or orders for the suggested actions can hinder the smooth execution of those suggestions. There is a need for a system that can solve this problem.
[1330] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1331] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data and understanding the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, and means for making reservations or orders related to the suggested actions. This allows the user to easily receive action suggestions through voice input, provide feedback to obtain the most suitable suggestion, and smoothly transition to actual actions based on the suggestions.
[1332] "Means for receiving voice input" refers to hardware or software that receives voice data and converts it into a format that can be processed within the system.
[1333] "Means for converting received voice input into text data" refers to software or hardware that uses speech recognition technology to convert a user's voice input into corresponding text data.
[1334] "Means for analyzing converted text data and understanding user intent" refers to software or hardware that uses natural language processing techniques to analyze the content of text data and understand what the user is asking for.
[1335] "Means for generating multiple action suggestions based on analyzed data" refers to an algorithm for generating multiple action options for the user based on the analysis results, and the hardware or software that executes it.
[1336] "Means for presenting generated action suggestions" refers to devices and software that present generated action options to the user, either visually or audibly.
[1337] "Means for receiving user feedback" refers to the hardware and software used to receive and incorporate user responses to suggestions into the system.
[1338] "Means for generating and presenting new suggestions based on received feedback" refers to algorithms and devices that generate and present new action suggestions based on user feedback.
[1339] "Means of making reservations or orders related to suggested actions" refers to online services for making restaurant reservations or ordering goods based on actions suggested within the system, as well as the software and hardware for executing them.
[1340] "Means of transmitting location information" refers to a GPS module and communication module used to determine the user's current location and transmit it to a server.
[1341] This invention provides a system for users to plan meals efficiently and intuitively. The system consists of a terminal that accepts voice input and a server that analyzes the data. Specific embodiments of this system are described in detail below.
[1342] System Configuration
[1343] This system includes means for receiving voice input, means for converting voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for presenting the generated action suggestions, means for receiving feedback from the user, means for generating and presenting suggestions again based on the received feedback, means for making reservations or orders related to the suggested actions, and means for transmitting location information.
[1344] Processing flow
[1345] 1. Accepting voice input
[1346] Users input their meal requests by voice into a device such as a smartphone. For example, they might say, "I'd like to eat some delicious pizza nearby."
[1347] The device is equipped with a microphone and accepts voice input.
[1348] 2. Text conversion of audio data
[1349] The device uses a speech recognition module (e.g., Google Speech Recognition) to convert the audio data into text data.
[1350] 3. Acquisition of location information
[1351] The device uses a GPS module to obtain the user's current location.
[1352] 4. Sending data
[1353] The device sends the acquired text data and location information to the server. A communication module is used.
[1354] 5. Data Analysis and Proposal Generation
[1355] The server analyzes text data and uses natural language processing techniques to understand the user's intent.
[1356] The server uses generative AI (e.g., GPT-3) to generate multiple action suggestions (e.g., restaurant or cafe selection) considering the user's intent, location, season, weather information, etc.
[1357] 6. Presentation of Action Proposals
[1358] The server sends the generated action suggestion to the device. The device then presents the suggestion to the user via text or voice (e.g., "I'd like to recommend John's Pizza nearby. Would you like to make a reservation?").
[1359] 7. Accepting Feedback
[1360] Users provide feedback on the suggested locations, such as "Is there a closer location?"
[1361] The device converts the feedback into text data and sends it to the server.
[1362] 8. Reanalysis and Re-proposal
[1363] The server analyzes the feedback and generates new suggestions based on the new conditions.
[1364] The revised suggestion is sent to the device, and the device presents the suggestion to the user again (e.g., "How about 'Erika's Pizza,' which is a 5-minute walk from there?").
[1365] 9. Execute reservation / order
[1366] If the user agrees to the suggestions, they can make restaurant reservations or order delivery from their device. The server then sends the reservation and order information to the online service to complete the process.
[1367] Specific example
[1368] When a user speaks to a smartphone app saying, "I want to eat some delicious pizza nearby," the app converts the speech to text, retrieves the user's location information, and sends it to a server. The server uses generative AI (e.g., GPT-3) to suggest the best pizza restaurant based on the user's location and preferences. One option presented is "John's Pizza - 5 minutes away," and the app responds with a voice message saying, "I'd like to introduce you to a nearby 'John's Pizza.' Would you like to make a reservation?"
[1369] Example of a prompt
[1370] User: "I want to eat some delicious pizza nearby."
[1371] system:
[1372] Speech recognition: Converts speech to text → "I want to eat some delicious pizza nearby."
[1373] Location information acquisition: Obtain current location from GPS.
[1374] Server: Receives text and location information, and generates suggestions using generative AI → "John's Pizza"
[1375] Response: Voice response from the speech generation module → "I'd like to recommend a nearby 'John's Pizza'. Would you like to make a reservation?"
[1376] Thus, by using the system of the present invention, users can not only easily perform voice input and receive appropriate suggestions, but also smoothly carry out actions based on those suggestions.
[1377] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1378] Step 1:
[1379] The user inputs their dining requests via voice. Using their smartphone's microphone, the user might say, "I want to eat some delicious pizza nearby." This voice data becomes the input.
[1380] Step 2:
[1381] The device converts the voice data into text data. A speech recognition module (e.g., Google Speech Recognition) analyzes the voice and outputs the text data "I want to eat some delicious pizza nearby." This text data becomes the input for the next processing step.
[1382] Step 3:
[1383] The device acquires its current location information. Using a GPS module, it measures the user's current location and obtains latitude and longitude information. This location information becomes the input for the next processing step.
[1384] Step 4:
[1385] The terminal sends text data and location information to the server. The text data and location information are passed to the server via a communication module. This text data and location information becomes the input for the next processing step.
[1386] Step 5:
[1387] The server analyzes the text data to understand the user's intent. Using natural language processing techniques, it analyzes the received text data and outputs the user's intent, "I want recommendations for nearby pizza restaurants." This intent becomes the input for the next processing step.
[1388] Step 6:
[1389] The server generates multiple action suggestions based on the user's intent and location information. Using a generative AI (e.g., GPT-3), it generates several pizza restaurant candidates, taking into account the user's intent, location information, and additional information such as weather and time of day. These action suggestions become the input for the next processing step.
[1390] Step 7:
[1391] The server selects the most suitable action suggestion and sends it to the terminal along with detailed information. The terminal then selects the most appropriate restaurant from the multiple suggestions and sends it along with its details (address, opening hours, menu, etc.). This becomes the input for the next processing step.
[1392] Step 8:
[1393] The device presents the user with an action suggestion. Using a speech synthesis module that converts text data into speech, it suggests to the user verbally, "I'd like to recommend a nearby 'John's Pizza.' Would you like to make a reservation?" This suggestion becomes the input for the next processing step.
[1394] Step 9:
[1395] The user provides feedback. They give voice feedback on the suggested location, such as "Is there a closer location?" This voice data becomes the input for the next processing step.
[1396] Step 10:
[1397] The device converts the user's voice feedback into text data. The voice recognition module is used again to convert the feedback into text data. This feedback becomes the input for the next processing step.
[1398] Step 11:
[1399] The terminal sends feedback text data to the server. The feedback text data is passed to the server via the communication module. This feedback becomes the input for the next processing step.
[1400] Step 12:
[1401] The server analyzes the feedback and generates new action suggestions. The generative AI is used again to generate new candidates based on the feedback and new conditions. These new action suggestions become the input for the next processing step.
[1402] Step 13:
[1403] The server sends a revised suggestion to the terminal, which then presents the suggestion to the user again. It then makes another voice suggestion, such as, "How about 'Erika's Pizza,' which is a 5-minute walk from there?" This process is repeated until the user is satisfied.
[1404] Step 14:
[1405] If the user agrees to the final action suggestion, they can make a reservation or order delivery from the relevant restaurant using their device. The device uses a communication module to send the reservation or order information to the server, which then transmits this information to the online service to complete the process. This reservation or order becomes the final output.
[1406] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1407] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, the invention aims to provide users with more personalized suggestions by incorporating an emotion engine that recognizes user emotions. Specific embodiments of this system are described in detail below.
[1408] System Configuration
[1409] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[1410] Program Processing Description
[1411] 1. User input reception:
[1412] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1413] The terminal accepts voice input and activates the voice recognition module.
[1414] 2. Emotional analysis of voice data:
[1415] The device sends voice input to the emotion engine.
[1416] The emotion engine analyzes the audio data and extracts emotional states (e.g., joy, excitement, fatigue, etc.).
[1417] The extracted emotion data is sent back to the device.
[1418] 3. Convert to text data:
[1419] The device converts voice input into text data.
[1420] The converted text data is reviewed to determine if it is a valid request.
[1421] 4. Sending data to the server:
[1422] The device sends the converted text data, user location information, and sentiment data to the server.
[1423] 5. Data analysis and proposal generation:
[1424] The server analyzes the text data it receives to understand the user's intent.
[1425] The server uses generative AI to generate multiple action suggestions, taking into account location information, current season, weather information, sentiment data, and past suggestion history.
[1426] The server selects the most suitable action from the generated suggestions and sends it to the terminal along with detailed information.
[1427] 6. Presentation to the user:
[1428] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[1429] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[1430] 7. Accepting feedback:
[1431] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[1432] The device then uses the speech recognition module again to convert the voice feedback into text data.
[1433] 8. Data retransmission and reanalysis:
[1434] The device then sends the converted feedback and sentiment data back to the server.
[1435] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[1436] The server selects the most suitable proposal from the revised proposals and sends the selected proposal to the terminal.
[1437] 9. Final presentation:
[1438] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[1439] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[1440] The following describes the processing flow.
[1441] System Configuration
[1442] This system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[1443] Program processing steps
[1444] Step 1:
[1445] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1446] The terminal accepts voice input and activates the voice recognition module.
[1447] Step 2:
[1448] The device sends voice data to the emotion engine.
[1449] The emotion engine analyzes the audio data and extracts the user's emotional state (e.g., joy, excitement, fatigue, etc.).
[1450] The extracted emotion data is sent back to the device.
[1451] Step 3:
[1452] The device converts voice input into text data.
[1453] Review the converted text data to determine if it is a valid request.
[1454] Step 4:
[1455] The device transmits the converted text data, user location information, and sentiment data to the server via a communication module.
[1456] Step 5:
[1457] The server analyzes the text data it receives to understand the user's intent.
[1458] The server uses generative AI to generate multiple action suggestions based on location information, current season, weather information, sentiment data, and past suggestion history.
[1459] Step 6:
[1460] The server selects the most suitable action from the generated suggestions (e.g., "Shonan Coast").
[1461] Summarize the selected proposals and their detailed information (access methods, points of interest, weather forecast, etc.).
[1462] Step 7:
[1463] The server sends the selected proposal as text data to the terminal.
[1464] Step 8:
[1465] The terminal receives a suggestion and converts it into speech data using a speech synthesis module.
[1466] The device presents the generated audio data to the user through its speaker (e.g., "How about a drive along the Shonan coast?").
[1467] Step 9:
[1468] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?").
[1469] The device then uses the speech recognition module again to convert the voice feedback into text data.
[1470] Step 10:
[1471] The device then sends the converted feedback and sentiment data back to the server.
[1472] Step 11:
[1473] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions.
[1474] The server selects the best option from the alternative suggestions (e.g., "Hayama Beach").
[1475] Step 12:
[1476] The server sends the re-selected proposals to the terminal as text data.
[1477] Step 13:
[1478] The terminal then uses a speech synthesis module to convert the received suggestion into voice data and presents it to the user (e.g., "How about Hayama Beach?").
[1479] Through these steps, users can quickly and intuitively decide on their next course of action. By using an emotion engine, personalized suggestions that align with the user's feelings become possible, further improving user satisfaction.
[1480] (Example 2)
[1481] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1482] Conventional behavior suggestion systems often fail to adequately consider users' emotional states and offer only uniform suggestions, resulting in low user satisfaction. Therefore, there is a need to develop a system that provides individually optimized suggestions based on the user's emotional state, thereby increasing user satisfaction. Furthermore, the ability to quickly respond to received feedback and offer revised suggestions is also required.
[1483] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for performing sentiment analysis on received voice data, means for extracting the user's emotional state based on the sentiment analysis, and means for analyzing the converted text data and sentiment data to understand the user's intent. This enables highly personalized suggestions based on sentiment analysis.
[1484] "Voice input" is a method of input that allows users to communicate questions or requests to a system using their voice.
[1485] "Text data" refers to data obtained by analyzing voice input and converting it into text information.
[1486] "Emotional analysis" is the process of extracting a user's emotional state (e.g., joy, excitement, fatigue, etc.) from audio data.
[1487] "Emotional data" refers to data that indicates the user's emotional state, obtained as a result of emotion analysis.
[1488] "Understanding user intent" means analyzing converted text and sentiment data to grasp what the user wants.
[1489] A "generative AI model" is an artificial intelligence algorithm that generates appropriate action suggestions based on given data.
[1490] "Action suggestions" refer to ideas and plans for proposing actions and activities that are appropriate for the user.
[1491] "Feedback" refers to users providing feedback, opinions, and requests regarding suggestions to the system.
[1492] This invention is a system for users to efficiently and intuitively plan enjoyable time spent with family and friends. This system includes a dedicated terminal and is intended to be installed in places where people gather. Furthermore, by combining this system with an emotion engine that recognizes the user's emotions, the invention aims to provide users with more personalized suggestions.
[1493] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module and a communication module, while the server is equipped with a data analysis module and a generative AI model. The emotion engine extracts emotions from the user's voice data and adjusts action suggestions based on the analysis results.
[1494] First, the system starts working when the user inputs a question or request via voice into the device. For example, a voice input might be, "Where should we go today?" The device receives the voice input, activates a speech recognition module (such as the Google Cloud Speech-to-Text API), and converts the voice into text data. This conversion may take several seconds.
[1495] Next, the terminal uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the voice data and extract the user's emotional state. The extracted emotion data is sent back to the terminal and then transmitted to the server along with the converted text data. The server receives this data and uses natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to analyze the user's intent.
[1496] The server further analyzes location information, current season, weather information, and past behavioral history, and uses a generative AI model (e.g., OpenAI GPT-3) to generate multiple action suggestions. It selects the most suitable suggestion from the generated suggestions and sends it back to the terminal. The terminal converts the received suggestion into audio data using a speech synthesis module (e.g., Amazon Polly) and presents it to the user through its speaker.
[1497] The user provides feedback on the presented suggestion (e.g., "Is there a place a little closer?"). The device uses the speech recognition module again to convert the voice feedback into text data and sends it back to the server. The server analyzes the feedback and sentiment data, generates multiple action suggestions again based on the new conditions, and sends the best suggestion back to the device. Finally, the device uses the speech synthesis module to convert the received suggestion into voice data and presents it to the user.
[1498] As a concrete example, the following prompt sentences could be input into the generation AI model:
[1499] Prompt: "Please suggest some family-friendly outing destinations. Users are happy and prefer nearby locations."
[1500] This system can provide more personalized action suggestions by taking into account the user's emotional state. Furthermore, by providing location information and detailed information, users can plan their actions with confidence. The inclusion of an emotion engine is expected to further improve user satisfaction.
[1501] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1502] Step 1:
[1503] The user inputs questions or requests into the device using voice (e.g., "Where should we go for fun today?").
[1504] The device accepts voice input and activates a speech recognition module (e.g., Google Cloud Speech-to-Text API). The voice data is converted into text data. The input is the user's voice data, and the output is text data.
[1505] Step 2:
[1506] The terminal sends the converted text data to an emotion engine (e.g., IBM Watson Tone Analyzer). The emotion engine analyzes the audio data and extracts the user's emotional state. The input is audio data, and the output is emotion data.
[1507] Step 3:
[1508] The device verifies the sentiment data and text data returned from the sentiment engine. It performs text verification to determine if the request is appropriate. The input is sentiment data and text data, and the output is the verified text data and sentiment data.
[1509] Step 4:
[1510] The device sends verified text data, user location information (e.g., GPS data), and sentiment data to the server. This communication typically uses the HTTP or HTTPS protocol. Inputs are text data, location information, and sentiment data, while output is a notification to the server that the transmission is complete.
[1511] Step 5:
[1512] The server analyzes the received text data using natural language processing algorithms (e.g., SpaCy or NLTK) to understand the user's intent. The input is text data, and the output is user intent data.
[1513] Step 6:
[1514] The server considers the user's location, current season, weather information, and sentiment data, and also accesses a database of past behavioral history. It generates multiple action suggestions using a generative AI model (e.g., OpenAI GPT-3). The input is location information, season, weather information, sentiment data, and past behavioral history, and the output is a list of action suggestions.
[1515] Step 7:
[1516] The server uses a machine learning algorithm to select the most suitable action from the generated suggestions. The selected action suggestion, along with detailed information, is then sent to the terminal. The input is a list of action suggestions, and the output is the optimal action suggestion and its details.
[1517] Step 8:
[1518] The terminal receives suggestions from the server and converts them into speech data using a speech synthesis module (e.g., Amazon Polly). This data is then presented to the user through the speaker. The input consists of optimal action suggestions and detailed information, while the output is speech data.
[1519] Step 9:
[1520] The user provides feedback on the suggested location (e.g., "Is there a place a little closer?"). The device then uses its speech recognition module again to convert the voice feedback into text data. The input is the user's voice data, and the output is the text data of the feedback.
[1521] Step 10:
[1522] The device then resends the converted feedback and sentiment data to the server. The inputs are the feedback text data and sentiment data, and the output is a notification that the transmission to the server is complete.
[1523] Step 11:
[1524] The server analyzes feedback and sentiment data and generates multiple action suggestions again based on the new conditions. The input is feedback text data and sentiment data, and the output is a list of new action suggestions.
[1525] Step 12:
[1526] The server selects the most suitable action from the regenerated suggestions and sends the selected suggestion to the terminal. The input is a list of new action suggestions, and the output is the re-selected action suggestion.
[1527] Step 13:
[1528] The terminal receives the suggested content again, converts it into speech data using a speech synthesis module, and presents it to the user. The input is the re-selected action suggestion, and the output is the speech data.
[1529] (Application Example 2)
[1530] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1531] Conventional behavior suggestion systems have difficulty taking into account the user's emotional state and detailed location information, resulting in suggestions that are not sufficiently personalized. Furthermore, the process of regenerating suggestions based on user voice feedback is inefficient, making it difficult to increase user satisfaction.
[1532] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text data, means for analyzing the converted text data to understand the user's intent, means for generating multiple action suggestions based on the analyzed data, means for extracting the emotional state and adjusting the action suggestions based on the extracted emotional data, means for optimizing the action suggestions generated based on the user's current location information, emotional data, and past suggestion history, means for presenting the generated action suggestions, means for receiving feedback from the user, and means for generating and presenting suggestions again based on the received feedback. This makes it possible to provide personalized action suggestions that take into account the user's emotional state and location information, thereby improving user satisfaction.
[1533] A "means for receiving voice input" refers to a module for acquiring voice input from a user as digital data.
[1534] "Means for converting voice input into text data" refers to a program that converts acquired voice data into text information.
[1535] "Means of analyzing text data and understanding user intent" refers to algorithms that analyze converted text data to understand the content of user utterances and requests.
[1536] "A means of generating multiple action suggestions based on analyzed data" refers to a program that generates multiple actions or choices that a user should take, based on the results of analyzing text data.
[1537] "Means for extracting emotional states" refers to an analysis engine that automatically extracts a user's emotions from audio and video data.
[1538] "Means for adjusting action suggestions based on extracted emotional data" refers to a program that modifies generated action suggestions to a more appropriate form according to the user's emotional state.
[1539] "Means for transmitting location information" refers to a module that acquires data on the user's current location and sends it to an analysis server.
[1540] "Means for optimizing action suggestions generated based on the user's current location information, emotional data, and past suggestion history" refers to a program that generates optimal action suggestions by considering multiple factors such as the user's location information, emotional state, and past suggestion history.
[1541] "Means for presenting generated action suggestions" refers to output devices such as displays and speech synthesis modules, as well as programs, for providing generated action suggestions to the user.
[1542] "Means for receiving user feedback" refers to a module that receives responses and opinions from users regarding suggestions they have made, in the form of audio or text.
[1543] "A means of generating and presenting new suggestions based on received feedback" refers to a program that analyzes feedback received from users, generates new action suggestions based on that analysis, and presents them to the user again.
[1544] "Means of providing detailed information" refers to modules that provide users with background information and additional explanations for the generated action suggestions.
[1545] This invention provides a system that allows users to efficiently receive action suggestions in physical stores using voice input. The specific configuration and operation of this system are described in detail below.
[1546] The system consists of a terminal that accepts voice input, a server that analyzes the data, and an emotion engine that recognizes the user's emotions. The terminal is equipped with a voice recognition module, a communication module, and an emotion recognition camera. The server includes a data analysis module, a generative AI model, and a database.
[1547] Specific hardware and software examples include the following:
[1548] Speech recognition module: Google Speech-to-Text API
[1549] Emotion analysis engine: Microsoft Azure Emotion API
[1550] Generative AI: OpenAI GPT Model
[1551] Communication module: HTTP / HTTPS
[1552] Data Analysis Module: Python Scripts and SQL Databases
[1553] System operation
[1554] 1. Accepting voice input:
[1555] The user initiates voice input by speaking into a smart terminal installed in the store. The user's voice is captured through the microphone and transmitted to the system as digital data.
[1556] 2. Text conversion and analysis of audio data:
[1557] The device converts the acquired audio data into text data using the Google Speech-to-Text API. The converted text data is sent to the server, where a data analysis module analyzes the user's intent.
[1558] 3. Extraction of emotional states:
[1559] The device's built-in emotion recognition camera and audio data are used to extract the user's emotional state via the Microsoft Azure Emotion API. This emotional data is then sent to a server.
[1560] 4. Generating action proposals:
[1561] The server uses the OpenAI GPT model to generate multiple action suggestions based on analyzed text data, sentiment data, user location information, and past suggestion history. These action suggestions are adjusted according to the user's emotional state.
[1562] 5. Presenting action proposals:
[1563] The generated action suggestions are sent from the server to the terminal and presented to the user using the terminal's display or speech synthesis module. A concrete example is, "There's a new autumn book fair going on at the bookstore on the third floor."
[1564] 6. Receiving user feedback and making revisions:
[1565] When a user provides feedback on a suggested solution, the device accepts voice input again and performs re-analysis. Based on the feedback, the server regenerates multiple action suggestions under new conditions and presents them to the user again. An example of a specific prompt is: "The user is asking, 'Are there any stores closer?' Their emotional state is excited, and their current location is the first floor of a shopping mall. Please tell me about recommended stores and events."
[1566] This allows the system to provide personalized action suggestions that take into account the user's emotional state and location, thereby improving user satisfaction.
[1567] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1568] Step 1:
[1569] The user initiates voice input. The user speaks into the smart device, and the voice input is acquired through the device's microphone. The input is the user's voice data, and the output is digital voice data.
[1570] Step 2:
[1571] The device converts the audio data into text data. The acquired audio data is converted into text data using the Google Speech-to-Text API. The input is audio data, and the output is text data.
[1572] Step 3:
[1573] The terminal sends text and audio data to the server. The terminal sends the converted text data and the acquired audio data to the server. The input is text and audio data, and the output is the data sent to the server.
[1574] Step 4:
[1575] The server analyzes text data to understand the user's intent. The data analysis module analyzes text data to understand the user's requests and intent. The input is text data, and the output is the analysis result.
[1576] Step 5:
[1577] The server extracts emotional data. The audio data is analyzed for emotional state using the Microsoft Azure Emotion API, and emotional data is extracted. The input is audio data, and the output is emotional data.
[1578] Step 6:
[1579] The server generates action suggestions. The server uses the OpenAI GPT model to generate action suggestions, taking into account analysis results, sentiment data, user location information, and past suggestion history. The inputs are analysis results, sentiment data, location information, and past suggestion history, and the output is multiple action suggestions.
[1580] Step 7:
[1581] The server optimizes the action suggestions. Using generative AI and a database, it optimizes the generated action suggestions according to the user's emotional state. The input is the action suggestion, and the output is the optimized action suggestion.
[1582] Step 8:
[1583] The server sends optimized action suggestions to the terminal, which then presents them to the user. The input is the optimized action suggestions, and the output is the data sent to the terminal.
[1584] Step 9:
[1585] The device presents action suggestions. Using the device's display and speech synthesis module, the generated action suggestions are presented to the user visually and audibly. The input is optimized action suggestions, and the output is the presentation of suggestions to the user.
[1586] Step 10:
[1587] The user provides feedback on the proposal. The user inputs the feedback by voice, which is captured through the device's microphone. The input is feedback voice data, and the output is digital voice data.
[1588] Step 11:
[1589] The device converts the feedback audio data into text. The acquired feedback audio data is then converted back into text data using the Google Speech-to-Text API. The input is the feedback audio data, and the output is text data.
[1590] Step 12:
[1591] The server re-parses the feedback text data and sentiment data. The terminal sends the converted feedback text data and sentiment data to the server, which then re-parses it. The input is the feedback text data and sentiment data, and the output is the re-parsed result.
[1592] Step 13:
[1593] The server generates action suggestions again. Based on the re-analysis results, it generates action suggestions again using the OpenAI GPT model according to the new conditions. The input is the re-analysis results, and the output is the regenerated action suggestions.
[1594] Step 14:
[1595] The server optimizes the regenerated action suggestions and sends them to the terminal. The terminal then presents them to the user. The input is the regenerated action suggestions, and the output is the data sent to the terminal.
[1596] Step 15:
[1597] The terminal presents the action suggestion again. The action suggestion is presented to the user again using a display or speech synthesis module. The input is the regenerated action suggestion, and the output is the presentation of the revised suggestion to the user.
[1598] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1599] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1600] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1601] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1602] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1603] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1604] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1605] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1606] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1607] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1608] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1609] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1610] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1611] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1612] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1613] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1614] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1615] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1616] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1617] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1618] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1619] The following is further disclosed regarding the embodiments described above.
[1620] (Claim 1)
[1621] A means of accepting voice input,
[1622] A means of converting received voice input into text data,
[1623] A means of analyzing the converted text data and understanding the user's intent,
[1624] A means for generating multiple action suggestions based on analyzed data,
[1625] A means of presenting the generated action proposals,
[1626] A means of receiving user feedback,
[1627] A means of generating and presenting proposals again based on the feedback received,
[1628] A system that includes this.
[1629] (Claim 2)
[1630] The system according to claim 1, comprising means for transmitting location information along with the received voice input.
[1631] (Claim 3)
[1632] The system according to claim 1, comprising means for providing the user with detailed information about the suggested action.
[1633] "Example 1"
[1634] (Claim 1)
[1635] A means of accepting voice input,
[1636] A means of converting received voice input into text data,
[1637] A means of sending location information to a server along with the converted text data,
[1638] A means of analyzing text data to understand user intent,
[1639] A means for generating multiple action suggestions using generative artificial intelligence based on analyzed data,
[1640] A means of selecting the most appropriate action proposal and sending it along with detailed information,
[1641] A means of converting the generated action suggestions into audio and presenting them to the user,
[1642] A means of receiving feedback from users, converting audio into text data, and sending it to a server,
[1643] A means of generating a new proposal based on the received feedback, converting it into audio, and presenting it,
[1644] A system that includes this.
[1645] (Claim 2)
[1646] The system according to claim 1, characterized in that it takes external data such as weather information and seasonal information into consideration during the analysis process.
[1647] (Claim 3)
[1648] The system according to claim 1, characterized in that it includes means of providing access methods and detailed local information when presenting a proposal.
[1649] "Application Example 1"
[1650] (Claim 1)
[1651] A means of accepting voice input,
[1652] A means of converting received voice input into text data,
[1653] A means of analyzing the converted text data and understanding the user's intent,
[1654] A means for generating multiple action suggestions based on analyzed data,
[1655] A means of presenting the generated action proposals,
[1656] A means of receiving user feedback,
[1657] A means of generating and presenting proposals again based on the feedback received,
[1658] Means of making a reservation or order related to the proposed action,
[1659] A system that includes this.
[1660] (Claim 2)
[1661] The system according to claim 1, comprising means for transmitting location information along with the received voice input.
[1662] (Claim 3)
[1663] The system according to claim 1, comprising means for providing the user with detailed information about the suggested action.
[1664] "Example 2 of combining an emotion engine"
[1665] (Claim 1)
[1666] A means of accepting voice input,
[1667] A means of converting received voice input into text data,
[1668] A means of performing emotion analysis on converted audio data,
[1669] A means for extracting a user's emotional state based on emotion analysis,
[1670] A means of analyzing converted text data and sentiment data to understand user intent,
[1671] A means of utilizing a generative AI model that generates multiple action suggestions based on analyzed data,
[1672] A means of selecting and presenting the most suitable proposal from the generated action proposals,
[1673] A means of receiving user feedback and generating new suggestions,
[1674] A system that includes this.
[1675] (Claim 2)
[1676] The system according to claim 1, further comprising means for transmitting the user's location information in addition to the received voice input.
[1677] (Claim 3)
[1678] The system according to claim 1, comprising means for providing the user with detailed information about the suggested action.
[1679] "Application example 2 of combining emotional engines"
[1680] (Claim 1)
[1681] A means of accepting voice input,
[1682] A means of converting received voice input into text data,
[1683] A means of analyzing the converted text data and understanding the user's intent,
[1684] A means for generating multiple action suggestions based on analyzed data,
[1685] A means of presenting the generated action proposals,
[1686] A means of receiving user feedback,
[1687] A means of generating and presenting proposals again based on the feedback received,
[1688] A means for extracting emotional states and adjusting action suggestions based on the extracted emotional data,
[1689] A means for optimizing action suggestions generated based on the user's current location information, sentiment data, and past suggestion history,
[1690] A system that includes this.
[1691] (Claim 2)
[1692] The system according to claim 1, comprising means for transmitting location information along with the received voice input.
[1693] (Claim 3)
[1694] The system according to claim 1, comprising means for providing the user with detailed information about the suggested action. [Explanation of symbols]
[1695] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of accepting voice input, A means of converting received voice input into text data, A means of analyzing the converted text data and understanding the user's intent, A means for generating multiple action suggestions based on analyzed data, A means of presenting the generated action proposals, A means of receiving user feedback, A means of generating and presenting proposals again based on the feedback received, A system that includes this.
2. The system according to claim 1, further comprising means for transmitting location information along with the received voice input.
3. The system according to claim 1, comprising means for providing the user with detailed information about suggested actions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A