System
The system addresses the challenge of slow data acquisition by converting voice input to text, analyzing intent, and visually displaying market data, allowing for rapid and effective investment decision-making.
Patent Information
- Application Number
- JP2024120536
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Traditional systems lack efficient means for accessing market data through voice input and visually understanding data, impeding the speed of investment decision-making.
A system that recognizes voice input, converts it into text, analyzes intent and entities, acquires market data, and visually displays the data with voice responses, utilizing speech recognition, natural language processing, and data visualization.
Enables users to quickly and efficiently acquire and understand necessary investment information through voice commands, providing real-time market analysis and investment advice.
Smart Images

Figure 2026019127000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Today's individual investors and financial professionals demand fast and accurate market data acquisition and analysis. However, traditional systems often lack sufficient means to access information through voice input or visually understand data. This creates a challenge as the time and effort required to acquire and interpret data impedes the speed of investment decision-making. [Means for solving the problem]
[0005] In order to solve the above problems, the present invention provides the following means: a means for recognizing voice input from a user and converting it into text; a means for analyzing intent and entities from the text and acquiring market data based on the analysis results; and a means for visually displaying the acquired market data and outputting a response generated based on the data by voice, thereby providing a system that enables a user to quickly and efficiently acquire and understand necessary investment information.
[0006] "Voice input" is a means by which a user provides instructions or information to a system through speech.
[0007] "Speech recognition" is a technology that captures voice input as a digital signal and converts it into text.
[0008] "Text conversion" is the process of converting speech data that has been recognized into text data.
[0009] "Natural language processing (NLP)" is a technology that analyzes user intent from text data and extracts information for specific purposes.
[0010] An "intent" is the goal or request that a user is trying to achieve through voice input.
[0011] An "entity" is an information unit that refers to a specific object or attribute in text data.
[0012] "Market data" refers to data related to financial markets, such as stock prices, trading volume, and company information.
[0013] "Visual display" refers to the technique of presenting acquired market data to a user in the form of graphs, charts, etc.
[0014] "Voice response" is a means of conveying system-generated responses to the user by voice.
[0015] "Multimodal input" is a technology that processes input data in multiple formats, including not only audio but also images and videos. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system works by combining speech recognition, natural language processing, data acquisition, and visual display.
[0038] ---
[0039] System Configuration
[0040] 1. Speech Recognition and Natural Language Processing
[0041] 1. Users
[0042] The user speaks a voice command, such as "What is the current Apple stock price?"
[0043] 2. Terminal
[0044] The device captures the user's voice input with a microphone and transmits it to the server.
[0045] 3. Server
[0046] The server uses a speech recognition system to convert the voice input into text, then uses a natural language processing system to parse the text and identify intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[0047] 2. Obtaining market data
[0048] 1. Server
[0049] Based on the intent and entity, send a request to a market data service to retrieve specific stock information (e.g., "Apple" stock price).
[0050] 3. Visual display and audio response
[0051] 1. Server
[0052] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[0053] 2. Terminal
[0054] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[0055] 3. Server
[0056] The server generates a response message to provide the user with the stock price information by voice, and creates a voice response using a voice synthesis system.
[0057] 4. Terminal
[0058] The terminal plays the voice response sent from the server and provides the information to the user.
[0059] ---
[0060] Specific examples
[0061] 1. Obtaining stock price information
[0062] 1. Users
[0063] User: "What's Apple's stock price right now?"
[0064] 2. Terminal
[0065] The terminal captures the user's voice and transmits it to the server.
[0066] 3. Server
[0067] The server converts the speech into text using a speech recognition system.
[0068] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[0069] Based on the intent, retrieve Apple's current stock price from a market data service.
[0070] 4. Terminal
[0071] Receive market data sent from the server and display it visually in the data visualizer.
[0072] 5. Server
[0073] Generate a voice response message saying "Apple's current stock price is $145" and create the voice response using a speech synthesis system.
[0074] 6. Terminal
[0075] A voice response is played to provide information to the user.
[0076] 2. Providing investment advice
[0077] 1. Users
[0078] User: "What investment advice would you give me?"
[0079] 2. Terminal
[0080] The device captures the audio and sends it to the server.
[0081] 3. Server
[0082] The server converts the speech to text and uses a natural language processing system to identify the intent "provide investment advice."
[0083] Analyze market data and generate investment advice based on current market trends.
[0084] Generate a voice response message that reads, "Based on current market trends, we recommend investing in technology stocks," and use a speech synthesis system to create the voice response.
[0085] 4. Terminal
[0086] A voice response is played to provide investment advice to the user.
[0087] As described above, the present invention provides a system that combines speech recognition, natural language processing, market data acquisition and visual display to enable users to efficiently acquire and understand investment information.
[0088] The processing flow will be explained below.
[0089] When obtaining stock price information
[0090] Step 1:
[0091] The user issues a voice command such as, "What is the current Apple stock price?"
[0092] Step 2:
[0093] The device captures the user's voice with a microphone and converts it into a digital signal.
[0094] Step 3:
[0095] The device transmits the captured audio data to the server.
[0096] Step 4:
[0097] The server analyzes the received voice data using a voice recognition system and converts it into text.
[0098] python
[0099] text_query = self.voice_recognition_system.transcribe(audio_input)
[0100] Step 5:
[0101] The server analyzes the text using a natural language processing system to identify intent and entities.
[0102] python
[0103] intent, entities = self.natural_language_processing.analyze_query(text_query)
[0104] Step 6:
[0105] Based on the analysis results, the server determines that the user's intent is to "get stock price information." At the same time, it extracts the entity "Apple."
[0106] Step 7:
[0107] The server sends a request to a market data service to get current stock price information for Apple.
[0108] python
[0109] market_data = self.market_data_service.get_stock_data(stock_ticker)
[0110] Step 8:
[0111] The server analyzes the acquired market data and sends it to the terminal.
[0112] Step 9:
[0113] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[0114] python
[0115] self.data_visualizer.visualize_data(market_data)
[0116] Step 10:
[0117] The server generates a response message to provide to the user: "Apple's current stock price is $145."
[0118] python
[0119] self.speak_response(f"Here is the market data for {stock_ticker}")
[0120] Step 11:
[0121] The server converts the generated message into a voice response using a speech synthesis system.
[0122] Step 12:
[0123] The terminal plays the voice response sent from the server and provides the information to the user.
[0124] python
[0125] play_audio_response(message)
[0126] ---
[0127] When providing investment advice
[0128] Step 1:
[0129] The user issues the voice command "Give me some investment advice."
[0130] Step 2:
[0131] The device captures the user's voice with a microphone and converts it into a digital signal.
[0132] Step 3:
[0133] The device transmits the captured audio data to the server.
[0134] Step 4:
[0135] The server analyzes the received voice data using a voice recognition system and converts it into text.
[0136] python
[0137] text_query = self.voice_recognition_system.transcribe(audio_input)
[0138] Step 5:
[0139] The server analyzes the text using a natural language processing system to identify the intent.
[0140] python
[0141] intent, entities = self.natural_language_processing.analyze_query(text_query)
[0142] Step 6:
[0143] Based on the analysis results, the server determines that the user's intent is to "provide investment advice."
[0144] Step 7:
[0145] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[0146] python
[0147] analysis = self.market_data_service.analyze_market()
[0148] Step 8:
[0149] The server generates a response message to provide to the user, which might say, "Based on current market trends, we recommend investing in technology stocks."
[0150] python
[0151] self.speak_response(f"My investment advice based on current market trends is: {analysis}")
[0152] Step 9:
[0153] The server converts the generated message into a voice response using a speech synthesis system.
[0154] Step 10:
[0155] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[0156] python
[0157] play_audio_response(message)
[0158] Example 1
[0159] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0160] Conventional investment information acquisition systems make it difficult for users to accurately and quickly acquire and understand market data in real time. Furthermore, since an intuitive method for acquiring information through voice commands was uncommon, improving the user experience was a challenge. Furthermore, the functionality for visually representing acquired data and providing voice responses was insufficient, making it difficult to meet the diverse needs of users.
[0161] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0162] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for outputting a response generated based on the market data by voice, means for using an existing voice recognition system and a natural language processing model for performing voice recognition and natural language processing, means for sending a request to an external data providing service to acquire market data and deserializing the result, means for using a data visualization tool to visually display the acquired market data in the form of graphs and charts, and means for using a voice synthesis system to generate a voice response, thereby enabling a user to intuitively acquire real-time market data through voice commands and quickly understand information visually and audibly.
[0163] "Means for recognizing voice input from a user" refers to a function that uses a voice input device such as a microphone to capture voice commands spoken by a user and obtain them in digital form.
[0164] "Means for converting said voice input to text" refers to a function that uses a voice recognition system to convert captured voice data into text form.
[0165] "Means for analyzing intent and entities from the text" refers to a function that uses a natural language processing model to identify objects (entities) related to a user's intent from text data.
[0166] "Means for obtaining market data based on the analyzed intent" refers to the function of sending a request derived from the analysis result to a market data providing service and obtaining the required data.
[0167] "Means for visually displaying the acquired market data" refers to the ability to use data visualization tools to display the acquired market data to a user in a visual format, such as a graph or chart.
[0168] "Means for outputting a response generated based on the market data in voice" refers to a function for generating a response message based on the acquired market data and providing it to the user in voice format using a voice synthesis system.
[0169] "Means of using existing speech recognition systems and natural language processing models for speech recognition and natural language processing" refers to the ability to apply existing speech recognition and natural language processing technologies to efficiently convert speech data into text and analyze the text.
[0170] "Means for sending requests to external data providers to retrieve market data and deserializing the results" means the functionality for sending requests to external market data providers to retrieve data and parsing the retrieved data into an appropriate format.
[0171] "Means of using data visualization tools to visually display acquired market data in the form of graphs and charts" refers to the ability to display data in a visually easy-to-understand format using dedicated visualization software or libraries.
[0172] "Means for using a speech synthesis system to generate a voice response" refers to the ability to use speech synthesis technology to convert a text message into speech form and provide it to a user.
[0173] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system combines speech recognition, natural language processing, data acquisition, and visual display. Specifically, the system integrates a speech recognition system, a natural language processing model, a market data provider, a data visualization tool, and a speech synthesis system.
[0174] System Configuration
[0175] 1. Speech Recognition and Natural Language Processing
[0176] 1. Users
[0177] The user speaks a voice command such as "What is the current Apple stock price?"
[0178] 2. Terminal
[0179] The device captures the user's voice input using the device's built-in microphone, which may include a smartphone or smart speaker.
[0180] The device sends the captured audio data to the server using an HTTP request.
[0181] 3. Server
[0182] The server uses a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text.
[0183] The server uses a natural language processing model (e.g., OpenAI GPT-3) to identify intents (user intentions) and entities (specific targets) from the text data.
[0184] 2. Obtaining market data
[0185] 1. Server
[0186] The server sends a request based on the intent and entity to an external market data provider (e.g., Yahoo Finance API) using the HTTP GET method.
[0187] The server receives the data returned by the API and deserializes it in JSON format, which contains the latest market information for a specific financial instrument.
[0188] 3. Visual display and audio response
[0189] 1. Server
[0190] The server analyzes the acquired market data, formats it into a format that is easy for users to understand, and sends the formatted data to the terminal.
[0191] 2. Terminal
[0192] The terminal receives the data sent from the server and uses data visualization tools (e.g., Plotly or Matplotlib) to generate and visually display graphs and charts, including line graphs and bar graphs.
[0193] 3. Server
[0194] The server generates a response message such as "Apple's current stock price is $145" and creates the audio data using a speech synthesis system (e.g., Amazon Polly).
[0195] 4. Terminal
[0196] The terminal plays the audio data sent from the server and provides the information to the user, without the user having to perform any specific operation.
[0197] Specific examples
[0198] Get stock quotes
[0199] 1. Users
[0200] User: "What's Apple's stock price right now?"
[0201] 2. Terminal
[0202] The terminal captures the user's voice and transmits the voice data to the server.
[0203] 3. Server
[0204] The server converts the speech to text using the Google Cloud Speech-to-Text API, then uses OpenAI's GPT-3 to identify the intent "get stock quotes" and the entity "Apple" from the text.
[0205] The server sends a request to the Yahoo Finance API to get Apple's current stock price.
[0206] 4. Terminal
[0207] The terminal receives market data sent from the server and displays it as a line graph using Plotly.
[0208] 5. Server
[0209] The server generates a text message saying, "Apple's current stock price is $145," and creates voice data using Amazon Polly.
[0210] 6. Terminal
[0211] The terminal plays back the audio data and provides information to the user.
[0212] This invention allows users to intuitively obtain real-time market data through voice commands, providing information quickly and easily understood both visually and audibly. As described above, the present invention combines voice recognition, natural language processing, market data acquisition and visual display to provide a system that allows users to efficiently obtain and understand investment information.
[0213] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0214] Step 1: Capture and send audio input
[0215] 1. Users
[0216] The user speaks a voice command such as "What is the current Apple stock price?"
[0217] Input: User's voice command
[0218] Output: Audio data
[0219] 2. Terminal
[0220] The terminal uses the device's built-in microphone to capture the user's voice.
[0221] The captured audio is converted into a digital format and data compressed.
[0222] The terminal transmits the compressed audio data to the server as an HTTP request.
[0223] Input: User's voice data
[0224] Output: HTTP request to the server
[0225] Step 2: Convert audio data to text
[0226] 1. Server
[0227] The server sends the received audio data to the Google Cloud Speech-to-Text API.
[0228] The Google Cloud Speech-to-Text API converts the audio data into text format and returns it to the server.
[0229] Input: Compressed audio data
[0230] Output: Text data
[0231] Step 3: Identifying intents and entities through natural language processing
[0232] 1. Server
[0233] The server inputs the acquired text data into the OpenAI GPT-3 model to analyze intent and entities.
[0234] GPT-3 parses text and extracts intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[0235] Input: Text data
[0236] Output: Intents and entities
[0237] Step 4: Obtain market data
[0238] 1. Server
[0239] The server sends an HTTP GET request to the Yahoo Finance API to retrieve market data for a particular financial instrument.
[0240] The Yahoo Finance API returns relevant market data in JSON format upon request.
[0241] The server deserializes the received JSON data and extracts the necessary information.
[0242] Input: Intents and Entities
[0243] Output: Market data
[0244] Step 5: Format and transmit market data
[0245] 1. Server
[0246] The server formats the acquired market data into a format that is easy for users to understand. The formatted data is organized, for example, as time-series data of stock prices.
[0247] The server transmits the formatted market data to the terminal.
[0248] Input: Market Data
[0249] Output: Formatted market data
[0250] Step 6: Generate a visual representation
[0251] 1. Terminal
[0252] The terminal uses the received market data to generate graphs and charts using data visualization tools such as Plotly and Matplotlib.
[0253] The generated visual representation is displayed on the screen of the terminal.
[0254] Input: Formatted market data
[0255] Output: Graphs and charts
[0256] Step 7: Generate and play a voice response
[0257] 1. Server
[0258] The server generates a response message such as "Apple's current stock price is $145."
[0259] The response message is sent to Amazon Polly and converted into voice data.
[0260] The generated voice data is transmitted to the terminal.
[0261] Input: Response message
[0262] Output: Audio data
[0263] 2. Terminal
[0264] The terminal plays back the received audio data and provides the information to the user.
[0265] Input: Audio data
[0266] Output: Audio information
[0267] This specific processing step allows users to obtain real-time market analysis and stock information through voice commands, and to understand it visually and audibly. The system provides an intuitive and easy-to-use interface for users.
[0268] (Application example 1)
[0269] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0270] In modern society, there is a demand for quick and efficient understanding of security risks and market trends. It is also important to enable users to easily access the information they need in real time by displaying that information visually and outputting it audibly. The present invention aims to provide a system that supports the acquisition and understanding of such information, thereby enabling users to quickly grasp risk management and market trends.
[0271] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0272] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for audio outputting a response generated based on the market data, means for acquiring data for analyzing crisis management and market trends in real time, means for visually displaying crisis management information from the data, and means for audio outputting a response generated based on the crisis management information, thereby enabling a user to acquire market analysis and security-related information in real time through voice commands and respond quickly.
[0273] "Means for recognizing voice input from a user" refers to technology that allows a device to receive speech uttered by a user and process it as a digital signal.
[0274] The "means for converting said voice input into text" refers to a technology for converting a voice signal into a string of characters, typically using a voice recognition algorithm.
[0275] The "means for analyzing intent and entities from the text" is a natural language processing technology for extracting a user's intent (intent) and specific target elements (entities) from text data.
[0276] The "means for acquiring market data based on the analyzed intent" is a technology for acquiring market data corresponding to the intent from an external database or API.
[0277] The "means for visually displaying acquired market data" refers to technology for displaying market data on a screen in the form of graphs, charts, etc.
[0278] The "means for outputting a response generated based on the market data as voice" is a technology that uses voice synthesis technology to provide a text response created based on market data to a user.
[0279] "Data acquisition means for analyzing crisis management and market trends in real time" refers to technology for collecting and analyzing current crisis management information and market trends in real time.
[0280] The "means for visually displaying crisis management information from the data" refers to a technology that visually displays collected and analyzed crisis management information in an easy-to-understand manner for users.
[0281] The "means for outputting a response generated based on the crisis management information as voice" is a technology that uses a voice synthesis technology to provide a text response generated based on the crisis management information to a user.
[0282] System Configuration
[0283] 1. Speech Recognition and Natural Language Processing
[0284] 1. Users
[0285] The user uses their smartphone or smart glasses to issue a voice command such as "Tell me the current crisis management information."
[0286] 2. Terminal
[0287] The device (smartphone or smart glasses) captures the user's voice input with a microphone and sends it to a server.
[0288] 3. Server
[0289] The server converts the voice input into text using the Google Speech-to-Text API, then uses a natural language processing system such as spaCy to parse the text and identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends").
[0290] 2. Obtaining market data and crisis management information
[0291] 1. Server
[0292] Use the News API and custom web scrapers to gather up-to-date crisis and market data based on intent and entities.
[0293] 3. Visual display and audio response
[0294] 1. Server
[0295] The acquired information is analyzed and presented in a visually easy-to-understand format using Plotly, and a voice response is generated using the Google Text-to-Speech API, stating, "The current trend in crisis management is an increase in phishing attacks."
[0296] 2. Terminal
[0297] The terminal displays the visual data sent from the server on its screen and plays back audio responses to provide information to the user.
[0298] Specific examples
[0299] 1. Speech Recognition and Natural Language Processing
[0300] A user utters, "What are the current trends in crisis management?"
[0301] The terminal captures the user's voice and sends it to the server.
[0302] The server converts the speech into text, and a natural language processing system identifies "obtaining crisis management information" and "current trends."
[0303] 2. Data Acquisition
[0304] The server collects crisis management information using the News API and custom web scrapers.
[0305] 3. Visual display and audio response
[0306] The server parses the information, formats it for visual display using Plotly, and generates an audio response using the Google Text-to-Speech API.
[0307] The terminal displays the displayed data on a screen and plays audio responses to provide information to the user.
[0308] Prompt Sentence Examples
[0309] Below are some example prompts to be fed into the generative AI model:
[0310] When a user says, "What are the current crisis management trends?" the application performs the following steps:
[0311] 1. Use the Google Speech-to-Text API to convert voice commands into text.
[0312] 2. Using spaCy, we analyzed the intent “Get crisis management information” and the entity “Current trends” from the text.
[0313] 3. Use the News API or a custom web scraper to gather the latest crisis management information.
[0314] 4. Use Plotly to generate visually easy-to-understand graphs and charts and display them on the user screen.
[0315] 5. Using the Google Text-to-Speech API, a voice response was generated that said, "The current trend in crisis management is an increase in phishing attacks," and played on a smartphone or smart glasses.
[0316] This allows users to easily obtain and understand real-time security information through voice commands.
[0317] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0318] Step 1:
[0319] The user says, "Tell me about the current trends in crisis management." The input is the user's voice command, and the output is voice data. This allows the system to recognize the user's request and begin processing.
[0320] Step 2:
[0321] The device captures the user's voice input with a microphone and sends it to the server. The input is voice data, and the output is a digital signal. The voice data is transmitted to the server via the Internet.
[0322] Step 3:
[0323] The server converts the voice input into text using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This process recognizes the voice as text.
[0324] Step 4:
[0325] The server uses a natural language processing system such as spaCy to identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends") from the text data. The input is text data, and the output is the parsed intent and entities. This clarifies the user's intent and specific request.
[0326] Step 5:
[0327] The server collects the latest crisis management information using the News API or a custom web scraper based on the intent and entity. The input is the intent and entity, and the output is the crisis management information data. This allows the appropriate information to be retrieved.
[0328] Step 6:
[0329] The server analyzes the acquired information and uses Plotly to generate graphs and charts in a visually easy-to-understand format. The input is crisis management information data, and the output is visual display data. This converts the information into a format that is easy for users to understand.
[0330] Step 7:
[0331] The server uses the Google Text-to-Speech API to generate a voice response stating, "Current crisis management trends include an increase in phishing attacks." The input is the text response message, and the output is the audio data. This provides audio information along with visual information.
[0332] Step 8:
[0333] The terminal displays the visual data sent from the server on the screen and plays back audio responses to provide information to the user. The input is visual display data and audio data, and the output is the displayed data and played audio. This allows the user to obtain and understand the information they need in real time.
[0334] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0335] The present invention provides more personalized market data and investment advice by combining a system that recognizes user voice input, converts text, and processes natural language with an emotion engine. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[0336] ---
[0337] System Configuration
[0338] 1. Speech Recognition and Natural Language Processing
[0339] 1. Users
[0340] Users can speak voice commands such as "What is the current Apple stock price?" or "I'd like some investment advice."
[0341] 2. Terminal
[0342] The device captures the user's voice input with a microphone and transmits it to the server.
[0343] 3. Server
[0344] The server uses a speech recognition system to convert the voice input into text, and then uses a natural language processing system to parse the text and identify intents and entities.
[0345] 2. Emotion recognition
[0346] 1. Server
[0347] The server recognizes the user's emotions using an emotion engine based on the voice recognition data and text data.
[0348] 3. Obtaining market data
[0349] 1. Server
[0350] Based on the intent and entity, a request is sent to the market data service to retrieve specific stock information or market data.
[0351] 4. Regulating responses based on emotions
[0352] 1. Server
[0353] The emotion engine adjusts the tone and content of responses based on the user's emotions. For example, if it detects that the user is feeling stressed, it will generate a more reassuring response.
[0354] 5. Visual display and audio response
[0355] 1. Server
[0356] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[0357] 2. Terminal
[0358] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[0359] 3. Server
[0360] The server generates a response message to be presented to the user and uses a speech synthesis system to create a voice response.
[0361] 4. Terminal
[0362] The terminal plays the voice response sent from the server and provides the information to the user.
[0363] ---
[0364] Specific examples
[0365] 1. Obtaining stock price information
[0366] 1. Users
[0367] User: "What's Apple's stock price right now?"
[0368] 2. Terminal
[0369] The terminal captures the user's voice and transmits it to the server.
[0370] 3. Server
[0371] The server converts the speech into text using a speech recognition system.
[0372] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[0373] Recognize user emotions with an emotion engine.
[0374] Based on the intent, retrieve Apple's current stock price from a market data service.
[0375] 4. Terminal
[0376] Receive market data sent from the server and display it visually in the data visualizer.
[0377] 5. Server
[0378] Generate a voice response message saying "Apple's current stock price is $145" and adjust the tone depending on the user's emotion.
[0379] Generate a reassuring response like, "Apple stock is currently at $145, don't worry, there are good investment opportunities waiting for you."
[0380] 6. Terminal
[0381] A voice response is played to provide information to the user.
[0382] 2. Providing investment advice
[0383] 1. Users
[0384] User: "What investment advice would you give me?"
[0385] 2. Terminal
[0386] The device captures the audio and sends it to the server.
[0387] 3. Server
[0388] The server converts the speech to text and uses a natural language processing system to identify the intent: "Provide investment advice."
[0389] Recognize user emotions with an emotion engine.
[0390] Comprehensively analyze market data and generate investment advice based on current market trends.
[0391] 4. Server
[0392] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions.
[0393] "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and a good long-term view," he suggests.
[0394] 5. Terminal
[0395] A voice response is played to provide investment advice to the user.
[0396] As described above, the present invention provides a system that enables users to efficiently acquire and understand personalized investment information by integrating and operating speech recognition, natural language processing, market data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[0397] The processing flow will be explained below.
[0398] When obtaining stock price information
[0399] Step 1:
[0400] The user issues a voice command such as, "What is the current Apple stock price?"
[0401] Step 2:
[0402] The device captures the user's voice with a microphone and converts it into a digital signal.
[0403] Step 3:
[0404] The device transmits the captured audio data to the server.
[0405] Step 4:
[0406] The server analyzes the received voice data using a voice recognition system and converts it into text.
[0407] python
[0408] text_query = self.voice_recognition_system.transcribe(audio_input)
[0409] Step 5:
[0410] The server analyzes the text using a natural language processing system to identify intent and entities.
[0411] python
[0412] intent, entities = self.natural_language_processing.analyze_query(text_query)
[0413] Step 6:
[0414] The server verifies that the intent parsed from the text is "Get stock quotes" and identifies "Apple" as the entity.
[0415] Step 7:
[0416] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[0417] python
[0418] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[0419] Step 8:
[0420] The server takes into account the perceived sentiment and sends a request to a market data service to get current stock price information for Apple.
[0421] python
[0422] market_data = self.market_data_service.get_stock_data(stock_ticker)
[0423] Step 9:
[0424] The server analyzes the acquired market data and sends it to the terminal.
[0425] Step 10:
[0426] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[0427] python
[0428] self.data_visualizer.visualize_data(market_data)
[0429] Step 11:
[0430] The server generates a response message to be delivered to the user, such as "Apple's current stock price is $145," and adjusts the tone according to the user's emotion as determined by the emotion engine.
[0431] python
[0432] response_message = f"Apple's current stock price is ${market_data['price']}."
[0433] response_message = self.emotion_engine.adjust_tone(response_message, user_emotion)
[0434] Step 12:
[0435] The server converts the generated message into a voice response using a speech synthesis system.
[0436] Step 13:
[0437] The terminal plays the voice response sent from the server and provides the information to the user.
[0438] python
[0439] play_audio_response(response_message)
[0440] ---
[0441] When providing investment advice
[0442] Step 1:
[0443] The user issues the voice command "Give me some investment advice."
[0444] Step 2:
[0445] The device captures the user's voice with a microphone and converts it into a digital signal.
[0446] Step 3:
[0447] The device transmits the captured audio data to the server.
[0448] Step 4:
[0449] The server analyzes the received voice data using a voice recognition system and converts it into text.
[0450] python
[0451] text_query = self.voice_recognition_system.transcribe(audio_input)
[0452] Step 5:
[0453] The server analyzes the text using a natural language processing system to identify the intent.
[0454] python
[0455] intent, entities = self.natural_language_processing.analyze_query(text_query)
[0456] Step 6:
[0457] The server determines that the intent parsed from the text is "provide investment advice."
[0458] Step 7:
[0459] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[0460] python
[0461] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[0462] Step 8:
[0463] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[0464] python
[0465] analysis = self.market_data_service.analyze_market()
[0466] investment_advice = f"My investment advice based on current market trends is: {analysis}"
[0467] Step 9:
[0468] The server adjusts the content and tone of the investment advice based on the perceived user sentiment.
[0469] python
[0470] investment_advice = self.emotion_engine.adjust_tone(investment_advice, user_emotion)
[0471] Step 10:
[0472] The server converts the generated investment advice message into a voice response using a voice synthesis system.
[0473] Step 11:
[0474] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[0475] python
[0476] play_audio_response(investment_advice)
[0477] Through the above steps, the present invention combines voice input, natural language processing, emotion recognition, data acquisition, and visual display to realize a system that provides users with personalized investment information, allowing them to acquire and understand investment information more quickly and efficiently.
[0478] Example 2
[0479] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0480] Conventional systems simply convert voice input from users into text and provide market data based on specific intents and entities, but are unable to generate personalized responses that take into account the user's emotions. This has led to the challenge of not being able to provide responses with the appropriate tone and content according to the user's emotions.
[0481] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0482] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for recognizing user emotions from the voice input and text data, means for acquiring market data based on the analyzed intent and entities, means for visually displaying the acquired market data, means for adjusting the tone and content of a response generated based on the user emotions, and means for outputting the market data and the adjusted response by voice, thereby enabling a user to efficiently acquire and understand personalized investment information.
[0483] A "user" is a human user who requests information by speaking voice commands to the system.
[0484] "Voice input" is voice data uttered by the user.
[0485] "Speech recognition" is a technology that analyzes a user's voice input and converts the content into text data.
[0486] "Text conversion" is the process of converting audio data into text data.
[0487] "Intent" refers to identifying the user's intention or purpose for text input.
[0488] An "entity" is an element that extracts and identifies specific keywords or items within a text.
[0489] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state from voice input or text data.
[0490] "Market data" refers to financial-related data such as specific stock information and economic indicators.
[0491] "Visual display" refers to converting acquired market data into a visually understandable format, such as a graph or chart, and displaying it.
[0492] "Response generation" is the process of creating a response message to be provided to a user based on acquired market data and user sentiment.
[0493] "Audio output" is the process of playing back the generated response message as audio.
[0494] This invention relates to a system that recognizes user voice input, performs text conversion and natural language processing, and combines it with an emotion engine to provide personalized market data and investment advice. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[0495] System configuration
[0496] The system includes the following main components:
[0497] 1. Acquiring voice input
[0498] User
[0499] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[0500] Terminal
[0501] The device will use a built-in microphone to capture the user's voice input, and at this stage it is advisable to use a high-quality microphone and noise-cancelling technology.
[0502] 2. Sending audio data to the server
[0503] Terminal
[0504] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[0505] 3. Speech to text conversion
[0506] server
[0507] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[0508] 4. Analysis using natural language processing
[0509] server
[0510] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[0511] 5. Emotion recognition
[0512] server
[0513] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[0514] 6. Obtaining Market Data
[0515] server
[0516] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[0517] 7. Response generation and coordination
[0518] server
[0519] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[0520] 8. Visual Indications
[0521] server
[0522] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[0523] Terminal
[0524] The terminal displays the visual data sent from the server on a browser or dedicated application.
[0525] 9. Providing voice responses
[0526] Terminal
[0527] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[0528] Specific examples
[0529] Get stock quotes
[0530] 1. Users
[0531] User: "What's Apple's stock price right now?"
[0532] 2. Terminal
[0533] The terminal captures the user's voice and transmits it to the server.
[0534] 3. Server
[0535] The server converts the speech to text using a speech recognition system. The natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple." The emotion engine recognizes the user's emotions. Based on the intent, the server retrieves Apple's current stock price from a market data service.
[0536] 4. Terminal
[0537] Receive market data sent from the server and display it visually using D3.js, Chart.js, etc.
[0538] 5. Server
[0539] Generate a voice response message such as "Apple's current stock price is $145" and adjust the tone depending on the user's emotions. Generate a reassuring response such as "Apple's current stock price is $145, don't worry, there are good investment opportunities waiting for you."
[0540] 6. Terminal
[0541] A voice response is played to provide information to the user.
[0542] Providing investment advice
[0543] 1. Users
[0544] User: "What investment advice would you give me?"
[0545] 2. Terminal
[0546] The device captures the audio and sends it to the server.
[0547] 3. Server
[0548] The server converts the speech into text and uses a natural language processing system to identify the intent "provide investment advice." The emotion engine recognizes the user's emotions. The server then comprehensively analyzes market data and generates investment advice based on current market trends.
[0549] 4. Server
[0550] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions. It also suggests, "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and are good for the long term."
[0551] 5. Terminal
[0552] A voice response is played to provide investment advice to the user.
[0553] Prompt Sentence Examples
[0554] User says: "What is Apple's stock price right now?"
[0555] Prompt: "What is the process flow if the user says, 'What is Apple's stock price now?'"
[0556] In this way, the system integrates multiple technologies, including user emotion recognition, to provide a way for users to efficiently obtain and understand personalized investment information.
[0557] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0558] Step 1:
[0559] Acquiring voice input
[0560] User
[0561] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[0562] Terminal
[0563] The device uses a built-in microphone to capture the user's voice input, utilizing high-quality microphone and noise-canceling technology to collect voice data in real time.
[0564] Input: User's voice command
[0565] Output: Captured audio data
[0566] Step 2:
[0567] Sending audio data to the server
[0568] Terminal
[0569] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[0570] Input: Captured audio data
[0571] Output: Audio data sent to the server
[0572] Step 3:
[0573] Speech-to-text conversion
[0574] server
[0575] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[0576] Input: Transmitted audio data
[0577] Output: Converted text data
[0578] Step 4:
[0579] Analysis using natural language processing
[0580] server
[0581] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[0582] Input: Text data
[0583] Output: Identified intents and entities
[0584] Step 5:
[0585] emotion recognition
[0586] server
[0587] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[0588] Input: Audio and text data
[0589] Output: Recognized user emotion
[0590] Step 6:
[0591] Obtaining Market Data
[0592] server
[0593] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[0594] Input: Identified intents and entities
[0595] Output: Market data
[0596] Step 7:
[0597] Response generation and coordination
[0598] server
[0599] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[0600] Input: Market data and user sentiment
[0601] Output: Adjusted response message
[0602] Step 8:
[0603] Visual Indication
[0604] server
[0605] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[0606] Input: Market Data
[0607] Output: Visually displayable data
[0608] Terminal
[0609] The terminal displays the visual data sent from the server on a browser or dedicated application.
[0610] Input: Visually displayable data
[0611] Output: Graphs and charts that are displayed to the user
[0612] Step 9:
[0613] Providing voice responses
[0614] Terminal
[0615] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[0616] Input: Tailored response message
[0617] Output: The audio response the user hears
[0618] In this way, each processing step works in conjunction with the others, allowing the user to efficiently obtain and understand personalized investment information.
[0619] (Application example 2)
[0620] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0621] Conventional food delivery services have the problem that they do not suggest personalized menus based on the user's emotions, making it difficult to select the optimal menu based on the user's psychological state. Also, because they do not suggest optimal dishes based on the user's emotions, the improvement in satisfaction is limited.
[0622] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring menu information based on the analyzed intent, means for analyzing the user's emotion based on the intent and entities, means for adjusting the recommended menu based on the emotion information, means for visually displaying the acquired menu information, and means for outputting a response generated based on the menu information by voice. This makes it possible to propose an optimal menu based on the user's emotion, thereby significantly improving user satisfaction.
[0623] A "user" is someone who uses the system.
[0624] "Voice input" is data of the voice that the user utters to the system.
[0625] "Text" refers to speech input converted into character data.
[0626] An "intent" is the purpose or request that a user intends to make to a system.
[0627] "Entity" means a concrete object or subject that is relevant to an intent.
[0628] "Menu information" is data related to the provision of food and drink, and is a list of dishes and drinks suggested to the user.
[0629] "Emotion" refers to the user's mental or emotional state.
[0630] A "recommended menu" is a selection of food and drink options suggested based on the user's intent and emotions.
[0631] The "visual display means" is a method for presenting the acquired menu information and recommended menus to the user in a graphical format.
[0632] "Audio output means" refers to a method for transmitting acquired information or generated responses to the user in audio form.
[0633] The term "system" refers to a series of functional blocks that operate in combination with the above means.
[0634] This invention relates to a food delivery system that recognizes and analyzes voice input from a user to suggest menu items based on the user's emotions. Below, we will explain the program and its detailed processing for specifically implementing the invention, the hardware and software used, and specific examples of use.
[0635] Hardware and Software Used
[0636] Hardware: smartphone, microphone, server
[0637] software:
[0638] speech_recognition library: for speech recognition
[0639] textblob library: for natural language processing
[0640] sentiment_analysis module: for emotion recognition
[0641] delivery_service_api: To obtain and recommend menu information
[0642] Explanation of the program processing flow
[0643] 1. Voice Recognition
[0644] When a user speaks, the microphone connected to the smartphone captures the voice data. The speech_recognition library is used to convert this voice data into text. An example of voice input might be, "I'm tired today, so I want some food to relax me."
[0645] 2. Natural Language Processing
[0646] The converted text data is then analyzed in detail using the textblob library, where the user's intent and the desired object (entity) are extracted. For example, from this utterance, the intent "I want to relax" is identified, and the entity "food" is identified.
[0647] 3. Emotion recognition
[0648] The analyzed text data is then subjected to emotion recognition by the sentiment_analysis module, which determines the user's emotional state as "fatigue" and selects the most appropriate menu based on this.
[0649] 4. Menu information acquisition and recommendation
[0650] Based on the user's sentiment and intent, the delivery_service_api is used to obtain the most suitable menu information. For example, it may recommend herbal tea or Japanese soup as a relaxing food.
[0651] 5. Visual and audio output
[0652] The acquired menu information is sent from the server to the smartphone and displayed visually on the smartphone screen. This information is also output as a voice response. For example, a voice response such as "Today, we recommend relaxing herbal tea or Japanese-style soup."
[0653] Specific examples
[0654] Usage example 1
[0655] User: "I'm tired today and I want some food to help me relax."
[0656] Server: Converts speech to text, identifies intent "relaxation", and determines emotional state as "fatigue".
[0657] System: Retrieves recommended menu items such as herbal tea and Japanese-style soup, and presents them to the user visually and audibly.
[0658] Prompt Sentence Examples
[0659] "User: "I'm tired today and I want some food to help me relax."
[0660] Server: "Thank you for your hard work. To help you relax, we recommend the following: herbal tea, Japanese soup. Please let us know if there's anything else you'd like."
[0661] In this way, the present invention realizes a system that is sensitive to the user's emotions and provides optimal menu suggestions and food delivery services.
[0662] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0663] Step 1:
[0664] The user opens the smartphone application and speaks into the microphone, for example, saying, "I'm tired today, so I want some food to help me relax." The voice recognition module in the smartphone captures this voice data.
[0665] Input: User voice input
[0666] Output: Captured audio data
[0667] Step 2:
[0668] The device converts the captured voice data into text data using the speech_recognition library. The user's statement, "I'm tired today, so I want some food to relax me," is converted into text.
[0669] Input: Captured audio data
[0670] Output: Converted text data
[0671] Step 3:
[0672] The server receives the converted text data and parses it using the textblob library to extract the intent "I want to relax" and the entity "food."
[0673] Input: Converted text data
[0674] Output: Intents and entities
[0675] Step 4:
[0676] The server analyzes the user's emotion using the sentiment_analysis module based on the intent and entities. In this case, the user's emotion is determined to be "fatigue."
[0677] Input: Intents and Entities
[0678] Output: User's emotional information
[0679] Step 5:
[0680] The server uses the preprocessed intent and emotion information to retrieve appropriate menu information via the delivery_service_api. For example, herbal tea or Japanese soup is selected as relaxing food.
[0681] Input: Intent and emotion information
[0682] Output: Recommended menu information
[0683] Step 6:
[0684] The server converts the recommended menu information into a visually displayable format and sends it to the terminal, which then uses its visualization engine to display the received data on the screen.
[0685] Input: Recommended menu information
[0686] Output: Visualized menu information
[0687] Step 7:
[0688] The server generates a voice response message to convey to the user and converts it into voice data using a speech synthesis system. For example, a message such as "Today, we recommend a relaxing herbal tea or Japanese-style soup."
[0689] Input: Recommended menu information
[0690] Output: Voice response message
[0691] Step 8:
[0692] The terminal plays the voice response message sent from the server to provide information to the user, who can then check the details of the relaxing menu through the visual menu display and voice response.
[0693] Input: Voice response message
[0694] Output: Providing audio and visual information to the user
[0695] As described above, a series of processes are carried out, from recognizing the voice input from the user, analyzing emotions based on that, and recommending and providing the most suitable menu.
[0696] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0697] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0698] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0699] [Second embodiment]
[0700] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0701] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0702] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0703] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0704] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0705] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0706] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0707] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0708] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0709] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0710] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0711] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0712] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system works by combining speech recognition, natural language processing, data acquisition, and visual display.
[0713] ---
[0714] System Configuration
[0715] 1. Speech Recognition and Natural Language Processing
[0716] 1. Users
[0717] The user speaks a voice command, such as "What is the current Apple stock price?"
[0718] 2. Terminal
[0719] The device captures the user's voice input with a microphone and transmits it to the server.
[0720] 3. Server
[0721] The server uses a speech recognition system to convert the voice input into text, then uses a natural language processing system to parse the text and identify intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[0722] 2. Obtaining market data
[0723] 1. Server
[0724] Based on the intent and entity, send a request to a market data service to retrieve specific stock information (e.g., "Apple" stock price).
[0725] 3. Visual display and audio response
[0726] 1. Server
[0727] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[0728] 2. Terminal
[0729] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[0730] 3. Server
[0731] The server generates a response message to provide the user with the stock price information by voice, and creates a voice response using a voice synthesis system.
[0732] 4. Terminal
[0733] The terminal plays the voice response sent from the server and provides the information to the user.
[0734] ---
[0735] Specific examples
[0736] 1. Obtaining stock price information
[0737] 1. Users
[0738] User: "What's Apple's stock price right now?"
[0739] 2. Terminal
[0740] The terminal captures the user's voice and transmits it to the server.
[0741] 3. Server
[0742] The server converts the speech into text using a speech recognition system.
[0743] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[0744] Based on the intent, retrieve Apple's current stock price from a market data service.
[0745] 4. Terminal
[0746] Receive market data sent from the server and display it visually in the data visualizer.
[0747] 5. Server
[0748] Generate a voice response message saying "Apple's current stock price is $145" and create the voice response using a speech synthesis system.
[0749] 6. Terminal
[0750] A voice response is played to provide information to the user.
[0751] 2. Providing investment advice
[0752] 1. Users
[0753] User: "What investment advice would you give me?"
[0754] 2. Terminal
[0755] The device captures the audio and sends it to the server.
[0756] 3. Server
[0757] The server converts the speech to text and uses a natural language processing system to identify the intent "provide investment advice."
[0758] Analyze market data and generate investment advice based on current market trends.
[0759] Generate a voice response message that reads, "Based on current market trends, we recommend investing in technology stocks," and use a speech synthesis system to create the voice response.
[0760] 4. Terminal
[0761] A voice response is played to provide investment advice to the user.
[0762] As described above, the present invention provides a system that combines speech recognition, natural language processing, market data acquisition and visual display to enable users to efficiently acquire and understand investment information.
[0763] The processing flow will be explained below.
[0764] When obtaining stock price information
[0765] Step 1:
[0766] The user issues a voice command such as, "What is the current Apple stock price?"
[0767] Step 2:
[0768] The device captures the user's voice with a microphone and converts it into a digital signal.
[0769] Step 3:
[0770] The device transmits the captured audio data to the server.
[0771] Step 4:
[0772] The server analyzes the received voice data using a voice recognition system and converts it into text.
[0773] python
[0774] text_query = self.voice_recognition_system.transcribe(audio_input)
[0775] Step 5:
[0776] The server analyzes the text using a natural language processing system to identify intent and entities.
[0777] python
[0778] intent, entities = self.natural_language_processing.analyze_query(text_query)
[0779] Step 6:
[0780] Based on the analysis results, the server determines that the user's intent is to "get stock price information." At the same time, it extracts the entity "Apple."
[0781] Step 7:
[0782] The server sends a request to a market data service to get current stock price information for Apple.
[0783] python
[0784] market_data = self.market_data_service.get_stock_data(stock_ticker)
[0785] Step 8:
[0786] The server analyzes the acquired market data and sends it to the terminal.
[0787] Step 9:
[0788] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[0789] python
[0790] self.data_visualizer.visualize_data(market_data)
[0791] Step 10:
[0792] The server generates a response message to provide to the user: "Apple's current stock price is $145."
[0793] python
[0794] self.speak_response(f"Here is the market data for {stock_ticker}")
[0795] Step 11:
[0796] The server converts the generated message into a voice response using a speech synthesis system.
[0797] Step 12:
[0798] The terminal plays the voice response sent from the server and provides the information to the user.
[0799] python
[0800] play_audio_response(message)
[0801] ---
[0802] When providing investment advice
[0803] Step 1:
[0804] The user issues the voice command "Give me some investment advice."
[0805] Step 2:
[0806] The device captures the user's voice with a microphone and converts it into a digital signal.
[0807] Step 3:
[0808] The device transmits the captured audio data to the server.
[0809] Step 4:
[0810] The server analyzes the received voice data using a voice recognition system and converts it into text.
[0811] python
[0812] text_query = self.voice_recognition_system.transcribe(audio_input)
[0813] Step 5:
[0814] The server analyzes the text using a natural language processing system to identify the intent.
[0815] python
[0816] intent, entities = self.natural_language_processing.analyze_query(text_query)
[0817] Step 6:
[0818] Based on the analysis results, the server determines that the user's intent is to "provide investment advice."
[0819] Step 7:
[0820] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[0821] python
[0822] analysis = self.market_data_service.analyze_market()
[0823] Step 8:
[0824] The server generates a response message to provide to the user, which might say, "Based on current market trends, we recommend investing in technology stocks."
[0825] python
[0826] self.speak_response(f"My investment advice based on current market trends is: {analysis}")
[0827] Step 9:
[0828] The server converts the generated message into a voice response using a speech synthesis system.
[0829] Step 10:
[0830] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[0831] python
[0832] play_audio_response(message)
[0833] Example 1
[0834] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0835] Conventional investment information acquisition systems make it difficult for users to accurately and quickly acquire and understand market data in real time. Furthermore, since an intuitive method for acquiring information through voice commands was uncommon, improving the user experience was a challenge. Furthermore, the functionality for visually representing acquired data and providing voice responses was insufficient, making it difficult to meet the diverse needs of users.
[0836] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0837] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for outputting a response generated based on the market data by voice, means for using an existing voice recognition system and a natural language processing model for performing voice recognition and natural language processing, means for sending a request to an external data providing service to acquire market data and deserializing the result, means for using a data visualization tool to visually display the acquired market data in the form of graphs and charts, and means for using a voice synthesis system to generate a voice response, thereby enabling a user to intuitively acquire real-time market data through voice commands and quickly understand information visually and audibly.
[0838] "Means for recognizing voice input from a user" refers to a function that uses a voice input device such as a microphone to capture voice commands spoken by a user and obtain them in digital form.
[0839] "Means for converting said voice input to text" refers to a function that uses a voice recognition system to convert captured voice data into text form.
[0840] "Means for analyzing intent and entities from the text" refers to a function that uses a natural language processing model to identify objects (entities) related to a user's intent from text data.
[0841] "Means for obtaining market data based on the analyzed intent" refers to the function of sending a request derived from the analysis result to a market data providing service and obtaining the required data.
[0842] "Means for visually displaying the acquired market data" refers to the ability to use data visualization tools to display the acquired market data to a user in a visual format, such as a graph or chart.
[0843] "Means for outputting a response generated based on the market data in voice" refers to a function for generating a response message based on the acquired market data and providing it to the user in voice format using a voice synthesis system.
[0844] "Means of using existing speech recognition systems and natural language processing models for speech recognition and natural language processing" refers to the ability to apply existing speech recognition and natural language processing technologies to efficiently convert speech data into text and analyze the text.
[0845] "Means for sending requests to external data providers to retrieve market data and deserializing the results" means the functionality for sending requests to external market data providers to retrieve data and parsing the retrieved data into an appropriate format.
[0846] "Means of using data visualization tools to visually display acquired market data in the form of graphs and charts" refers to the ability to display data in a visually easy-to-understand format using dedicated visualization software or libraries.
[0847] "Means for using a speech synthesis system to generate a voice response" refers to the ability to use speech synthesis technology to convert a text message into speech form and provide it to a user.
[0848] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system combines speech recognition, natural language processing, data acquisition, and visual display. Specifically, the system integrates a speech recognition system, a natural language processing model, a market data provider, a data visualization tool, and a speech synthesis system.
[0849] System Configuration
[0850] 1. Speech Recognition and Natural Language Processing
[0851] 1. Users
[0852] The user speaks a voice command such as "What is the current Apple stock price?"
[0853] 2. Terminal
[0854] The device captures the user's voice input using the device's built-in microphone, which may include a smartphone or smart speaker.
[0855] The device sends the captured audio data to the server using an HTTP request.
[0856] 3. Server
[0857] The server uses a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text.
[0858] The server uses a natural language processing model (e.g., OpenAI GPT-3) to identify intents (user intentions) and entities (specific targets) from the text data.
[0859] 2. Obtaining market data
[0860] 1. Server
[0861] The server sends a request based on the intent and entity to an external market data provider (e.g., Yahoo Finance API) using the HTTP GET method.
[0862] The server receives the data returned by the API and deserializes it in JSON format, which contains the latest market information for a specific financial instrument.
[0863] 3. Visual display and audio response
[0864] 1. Server
[0865] The server analyzes the acquired market data, formats it into a format that is easy for users to understand, and sends the formatted data to the terminal.
[0866] 2. Terminal
[0867] The terminal receives the data sent from the server and uses data visualization tools (e.g., Plotly or Matplotlib) to generate and visually display graphs and charts, including line graphs and bar graphs.
[0868] 3. Server
[0869] The server generates a response message such as "Apple's current stock price is $145" and creates the audio data using a speech synthesis system (e.g., Amazon Polly).
[0870] 4. Terminal
[0871] The terminal plays the audio data sent from the server and provides the information to the user, without the user having to perform any specific operation.
[0872] Specific examples
[0873] Get stock quotes
[0874] 1. Users
[0875] User: "What's Apple's stock price right now?"
[0876] 2. Terminal
[0877] The terminal captures the user's voice and transmits the voice data to the server.
[0878] 3. Server
[0879] The server converts the speech to text using the Google Cloud Speech-to-Text API, then uses OpenAI's GPT-3 to identify the intent "get stock quotes" and the entity "Apple" from the text.
[0880] The server sends a request to the Yahoo Finance API to get Apple's current stock price.
[0881] 4. Terminal
[0882] The terminal receives market data sent from the server and displays it as a line graph using Plotly.
[0883] 5. Server
[0884] The server generates a text message saying, "Apple's current stock price is $145," and creates voice data using Amazon Polly.
[0885] 6. Terminal
[0886] The terminal plays back the audio data and provides information to the user.
[0887] This invention allows users to intuitively obtain real-time market data through voice commands, providing information quickly and easily understood both visually and audibly. As described above, the present invention combines voice recognition, natural language processing, market data acquisition and visual display to provide a system that allows users to efficiently obtain and understand investment information.
[0888] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0889] Step 1: Capture and send audio input
[0890] 1. Users
[0891] The user speaks a voice command such as "What is the current Apple stock price?"
[0892] Input: User's voice command
[0893] Output: Audio data
[0894] 2. Terminal
[0895] The terminal uses the device's built-in microphone to capture the user's voice.
[0896] The captured audio is converted into a digital format and data compressed.
[0897] The terminal transmits the compressed audio data to the server as an HTTP request.
[0898] Input: User's voice data
[0899] Output: HTTP request to the server
[0900] Step 2: Convert audio data to text
[0901] 1. Server
[0902] The server sends the received audio data to the Google Cloud Speech-to-Text API.
[0903] The Google Cloud Speech-to-Text API converts the audio data into text format and returns it to the server.
[0904] Input: Compressed audio data
[0905] Output: Text data
[0906] Step 3: Identifying intents and entities through natural language processing
[0907] 1. Server
[0908] The server inputs the acquired text data into the OpenAI GPT-3 model to analyze intent and entities.
[0909] GPT-3 parses text and extracts intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[0910] Input: Text data
[0911] Output: Intents and entities
[0912] Step 4: Obtain market data
[0913] 1. Server
[0914] The server sends an HTTP GET request to the Yahoo Finance API to retrieve market data for a particular financial instrument.
[0915] The Yahoo Finance API returns relevant market data in JSON format upon request.
[0916] The server deserializes the received JSON data and extracts the necessary information.
[0917] Input: Intents and Entities
[0918] Output: Market data
[0919] Step 5: Format and transmit market data
[0920] 1. Server
[0921] The server formats the acquired market data into a format that is easy for users to understand. The formatted data is organized, for example, as time-series data of stock prices.
[0922] The server transmits the formatted market data to the terminal.
[0923] Input: Market Data
[0924] Output: Formatted market data
[0925] Step 6: Generate a visual representation
[0926] 1. Terminal
[0927] The terminal uses the received market data to generate graphs and charts using data visualization tools such as Plotly and Matplotlib.
[0928] The generated visual representation is displayed on the screen of the terminal.
[0929] Input: Formatted market data
[0930] Output: Graphs and charts
[0931] Step 7: Generate and play a voice response
[0932] 1. Server
[0933] The server generates a response message such as "Apple's current stock price is $145."
[0934] The response message is sent to Amazon Polly and converted into voice data.
[0935] The generated voice data is transmitted to the terminal.
[0936] Input: Response message
[0937] Output: Audio data
[0938] 2. Terminal
[0939] The terminal plays back the received audio data and provides the information to the user.
[0940] Input: Audio data
[0941] Output: Audio information
[0942] This specific processing step allows users to obtain real-time market analysis and stock information through voice commands, and to understand it visually and audibly. The system provides an intuitive and easy-to-use interface for users.
[0943] (Application example 1)
[0944] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0945] In modern society, there is a demand for quick and efficient understanding of security risks and market trends. It is also important to enable users to easily access the information they need in real time by displaying that information visually and outputting it audibly. The present invention aims to provide a system that supports the acquisition and understanding of such information, thereby enabling users to quickly grasp risk management and market trends.
[0946] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0947] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for audio outputting a response generated based on the market data, means for acquiring data for analyzing crisis management and market trends in real time, means for visually displaying crisis management information from the data, and means for audio outputting a response generated based on the crisis management information, thereby enabling a user to acquire market analysis and security-related information in real time through voice commands and respond quickly.
[0948] "Means for recognizing voice input from a user" refers to technology that allows a device to receive speech uttered by a user and process it as a digital signal.
[0949] The "means for converting said voice input into text" refers to a technology for converting a voice signal into a string of characters, typically using a voice recognition algorithm.
[0950] The "means for analyzing intent and entities from the text" is a natural language processing technology for extracting a user's intent (intent) and specific target elements (entities) from text data.
[0951] The "means for acquiring market data based on the analyzed intent" is a technology for acquiring market data corresponding to the intent from an external database or API.
[0952] The "means for visually displaying acquired market data" refers to technology for displaying market data on a screen in the form of graphs, charts, etc.
[0953] The "means for outputting a response generated based on the market data as voice" is a technology that uses voice synthesis technology to provide a text response created based on market data to a user.
[0954] "Data acquisition means for analyzing crisis management and market trends in real time" refers to technology for collecting and analyzing current crisis management information and market trends in real time.
[0955] The "means for visually displaying crisis management information from the data" refers to a technology that visually displays collected and analyzed crisis management information in an easy-to-understand manner for users.
[0956] The "means for outputting a response generated based on the crisis management information as voice" is a technology that uses a voice synthesis technology to provide a text response generated based on the crisis management information to a user.
[0957] System Configuration
[0958] 1. Speech Recognition and Natural Language Processing
[0959] 1. Users
[0960] The user uses their smartphone or smart glasses to issue a voice command such as "Tell me the current crisis management information."
[0961] 2. Terminal
[0962] The device (smartphone or smart glasses) captures the user's voice input with a microphone and sends it to a server.
[0963] 3. Server
[0964] The server converts the voice input into text using the Google Speech-to-Text API, then uses a natural language processing system such as spaCy to parse the text and identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends").
[0965] 2. Obtaining market data and crisis management information
[0966] 1. Server
[0967] Use the News API and custom web scrapers to gather up-to-date crisis and market data based on intent and entities.
[0968] 3. Visual display and audio response
[0969] 1. Server
[0970] The acquired information is analyzed and presented in a visually easy-to-understand format using Plotly, and a voice response is generated using the Google Text-to-Speech API, stating, "The current trend in crisis management is an increase in phishing attacks."
[0971] 2. Terminal
[0972] The terminal displays the visual data sent from the server on its screen and plays back audio responses to provide information to the user.
[0973] Specific examples
[0974] 1. Speech Recognition and Natural Language Processing
[0975] A user utters, "What are the current trends in crisis management?"
[0976] The terminal captures the user's voice and sends it to the server.
[0977] The server converts the speech into text, and a natural language processing system identifies "obtaining crisis management information" and "current trends."
[0978] 2. Data Acquisition
[0979] The server collects crisis management information using the News API and custom web scrapers.
[0980] 3. Visual display and audio response
[0981] The server parses the information, formats it for visual display using Plotly, and generates an audio response using the Google Text-to-Speech API.
[0982] The terminal displays the displayed data on a screen and plays audio responses to provide information to the user.
[0983] Prompt Sentence Examples
[0984] Below are some example prompts to be fed into the generative AI model:
[0985] When a user says, "What are the current crisis management trends?" the application performs the following steps:
[0986] 1. Use the Google Speech-to-Text API to convert voice commands into text.
[0987] 2. Using spaCy, we analyzed the intent “Get crisis management information” and the entity “Current trends” from the text.
[0988] 3. Use the News API or a custom web scraper to gather the latest crisis management information.
[0989] 4. Use Plotly to generate visually easy-to-understand graphs and charts and display them on the user screen.
[0990] 5. Using the Google Text-to-Speech API, a voice response was generated that said, "The current trend in crisis management is an increase in phishing attacks," and played on a smartphone or smart glasses.
[0991] This allows users to easily obtain and understand real-time security information through voice commands.
[0992] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0993] Step 1:
[0994] The user says, "Tell me about the current trends in crisis management." The input is the user's voice command, and the output is voice data. This allows the system to recognize the user's request and begin processing.
[0995] Step 2:
[0996] The device captures the user's voice input with a microphone and sends it to the server. The input is voice data, and the output is a digital signal. The voice data is transmitted to the server via the Internet.
[0997] Step 3:
[0998] The server converts the voice input into text using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This process recognizes the voice as text.
[0999] Step 4:
[1000] The server uses a natural language processing system such as spaCy to identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends") from the text data. The input is text data, and the output is the parsed intent and entities. This clarifies the user's intent and specific request.
[1001] Step 5:
[1002] The server collects the latest crisis management information using the News API or a custom web scraper based on the intent and entity. The input is the intent and entity, and the output is the crisis management information data. This allows the appropriate information to be retrieved.
[1003] Step 6:
[1004] The server analyzes the acquired information and uses Plotly to generate graphs and charts in a visually easy-to-understand format. The input is crisis management information data, and the output is visual display data. This converts the information into a format that is easy for users to understand.
[1005] Step 7:
[1006] The server uses the Google Text-to-Speech API to generate a voice response stating, "Current crisis management trends include an increase in phishing attacks." The input is the text response message, and the output is the audio data. This provides audio information along with visual information.
[1007] Step 8:
[1008] The terminal displays the visual data sent from the server on the screen and plays back audio responses to provide information to the user. The input is visual display data and audio data, and the output is the displayed data and played audio. This allows the user to obtain and understand the information they need in real time.
[1009] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1010] The present invention provides more personalized market data and investment advice by combining a system that recognizes user voice input, converts text, and processes natural language with an emotion engine. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[1011] ---
[1012] System Configuration
[1013] 1. Speech Recognition and Natural Language Processing
[1014] 1. Users
[1015] Users can speak voice commands such as "What is the current Apple stock price?" or "I'd like some investment advice."
[1016] 2. Terminal
[1017] The device captures the user's voice input with a microphone and transmits it to the server.
[1018] 3. Server
[1019] The server uses a speech recognition system to convert the voice input into text, and then uses a natural language processing system to parse the text and identify intents and entities.
[1020] 2. Emotion recognition
[1021] 1. Server
[1022] The server recognizes the user's emotions using an emotion engine based on the voice recognition data and text data.
[1023] 3. Obtaining market data
[1024] 1. Server
[1025] Based on the intent and entity, a request is sent to the market data service to retrieve specific stock information or market data.
[1026] 4. Regulating responses based on emotions
[1027] 1. Server
[1028] The emotion engine adjusts the tone and content of responses based on the user's emotions. For example, if it detects that the user is feeling stressed, it will generate a more reassuring response.
[1029] 5. Visual display and audio response
[1030] 1. Server
[1031] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[1032] 2. Terminal
[1033] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[1034] 3. Server
[1035] The server generates a response message to be presented to the user and uses a speech synthesis system to create a voice response.
[1036] 4. Terminal
[1037] The terminal plays the voice response sent from the server and provides the information to the user.
[1038] ---
[1039] Specific examples
[1040] 1. Obtaining stock price information
[1041] 1. Users
[1042] User: "What's Apple's stock price right now?"
[1043] 2. Terminal
[1044] The terminal captures the user's voice and transmits it to the server.
[1045] 3. Server
[1046] The server converts the speech into text using a speech recognition system.
[1047] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[1048] Recognize user emotions with an emotion engine.
[1049] Based on the intent, retrieve Apple's current stock price from a market data service.
[1050] 4. Terminal
[1051] Receive market data sent from the server and display it visually in the data visualizer.
[1052] 5. Server
[1053] Generate a voice response message saying "Apple's current stock price is $145" and adjust the tone depending on the user's emotion.
[1054] Generate a reassuring response like, "Apple stock is currently at $145, don't worry, there are good investment opportunities waiting for you."
[1055] 6. Terminal
[1056] A voice response is played to provide information to the user.
[1057] 2. Providing investment advice
[1058] 1. Users
[1059] User: "What investment advice would you give me?"
[1060] 2. Terminal
[1061] The device captures the audio and sends it to the server.
[1062] 3. Server
[1063] The server converts the speech to text and uses a natural language processing system to identify the intent: "Provide investment advice."
[1064] Recognize user emotions with an emotion engine.
[1065] Comprehensively analyze market data and generate investment advice based on current market trends.
[1066] 4. Server
[1067] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions.
[1068] "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and a good long-term view," he suggests.
[1069] 5. Terminal
[1070] A voice response is played to provide investment advice to the user.
[1071] As described above, the present invention provides a system that enables users to efficiently acquire and understand personalized investment information by integrating and operating speech recognition, natural language processing, market data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[1072] The processing flow will be explained below.
[1073] When obtaining stock price information
[1074] Step 1:
[1075] The user issues a voice command such as, "What is the current Apple stock price?"
[1076] Step 2:
[1077] The device captures the user's voice with a microphone and converts it into a digital signal.
[1078] Step 3:
[1079] The device transmits the captured audio data to the server.
[1080] Step 4:
[1081] The server analyzes the received voice data using a voice recognition system and converts it into text.
[1082] python
[1083] text_query = self.voice_recognition_system.transcribe(audio_input)
[1084] Step 5:
[1085] The server analyzes the text using a natural language processing system to identify intent and entities.
[1086] python
[1087] intent, entities = self.natural_language_processing.analyze_query(text_query)
[1088] Step 6:
[1089] The server verifies that the intent parsed from the text is "Get stock quotes" and identifies "Apple" as the entity.
[1090] Step 7:
[1091] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[1092] python
[1093] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[1094] Step 8:
[1095] The server takes into account the perceived sentiment and sends a request to a market data service to get current stock price information for Apple.
[1096] python
[1097] market_data = self.market_data_service.get_stock_data(stock_ticker)
[1098] Step 9:
[1099] The server analyzes the acquired market data and sends it to the terminal.
[1100] Step 10:
[1101] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[1102] python
[1103] self.data_visualizer.visualize_data(market_data)
[1104] Step 11:
[1105] The server generates a response message to be delivered to the user, such as "Apple's current stock price is $145," and adjusts the tone according to the user's emotion as determined by the emotion engine.
[1106] python
[1107] response_message = f"Apple's current stock price is ${market_data['price']}."
[1108] response_message = self.emotion_engine.adjust_tone(response_message, user_emotion)
[1109] Step 12:
[1110] The server converts the generated message into a voice response using a speech synthesis system.
[1111] Step 13:
[1112] The terminal plays the voice response sent from the server and provides the information to the user.
[1113] python
[1114] play_audio_response(response_message)
[1115] ---
[1116] When providing investment advice
[1117] Step 1:
[1118] The user issues the voice command "Give me some investment advice."
[1119] Step 2:
[1120] The device captures the user's voice with a microphone and converts it into a digital signal.
[1121] Step 3:
[1122] The device transmits the captured audio data to the server.
[1123] Step 4:
[1124] The server analyzes the received voice data using a voice recognition system and converts it into text.
[1125] python
[1126] text_query = self.voice_recognition_system.transcribe(audio_input)
[1127] Step 5:
[1128] The server analyzes the text using a natural language processing system to identify the intent.
[1129] python
[1130] intent, entities = self.natural_language_processing.analyze_query(text_query)
[1131] Step 6:
[1132] The server determines that the intent parsed from the text is "provide investment advice."
[1133] Step 7:
[1134] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[1135] python
[1136] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[1137] Step 8:
[1138] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[1139] python
[1140] analysis = self.market_data_service.analyze_market()
[1141] investment_advice = f"My investment advice based on current market trends is: {analysis}"
[1142] Step 9:
[1143] The server adjusts the content and tone of the investment advice based on the perceived user sentiment.
[1144] python
[1145] investment_advice = self.emotion_engine.adjust_tone(investment_advice, user_emotion)
[1146] Step 10:
[1147] The server converts the generated investment advice message into a voice response using a voice synthesis system.
[1148] Step 11:
[1149] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[1150] python
[1151] play_audio_response(investment_advice)
[1152] Through the above steps, the present invention combines voice input, natural language processing, emotion recognition, data acquisition, and visual display to realize a system that provides users with personalized investment information, allowing them to acquire and understand investment information more quickly and efficiently.
[1153] Example 2
[1154] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1155] Conventional systems simply convert voice input from users into text and provide market data based on specific intents and entities, but are unable to generate personalized responses that take into account the user's emotions. This has led to the challenge of not being able to provide responses with the appropriate tone and content according to the user's emotions.
[1156] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1157] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for recognizing user emotions from the voice input and text data, means for acquiring market data based on the analyzed intent and entities, means for visually displaying the acquired market data, means for adjusting the tone and content of a response generated based on the user emotions, and means for outputting the market data and the adjusted response by voice, thereby enabling a user to efficiently acquire and understand personalized investment information.
[1158] A "user" is a human user who requests information by speaking voice commands to the system.
[1159] "Voice input" is voice data uttered by the user.
[1160] "Speech recognition" is a technology that analyzes a user's voice input and converts the content into text data.
[1161] "Text conversion" is the process of converting audio data into text data.
[1162] "Intent" refers to identifying the user's intention or purpose for text input.
[1163] An "entity" is an element that extracts and identifies specific keywords or items within a text.
[1164] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state from voice input or text data.
[1165] "Market data" refers to financial-related data such as specific stock information and economic indicators.
[1166] "Visual display" refers to converting acquired market data into a visually understandable format, such as a graph or chart, and displaying it.
[1167] "Response generation" is the process of creating a response message to be provided to a user based on acquired market data and user sentiment.
[1168] "Audio output" is the process of playing back the generated response message as audio.
[1169] This invention relates to a system that recognizes user voice input, performs text conversion and natural language processing, and combines it with an emotion engine to provide personalized market data and investment advice. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[1170] System configuration
[1171] The system includes the following main components:
[1172] 1. Acquiring voice input
[1173] User
[1174] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[1175] Terminal
[1176] The device will use a built-in microphone to capture the user's voice input, and at this stage it is advisable to use a high-quality microphone and noise-cancelling technology.
[1177] 2. Sending audio data to the server
[1178] Terminal
[1179] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[1180] 3. Speech to text conversion
[1181] server
[1182] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[1183] 4. Analysis using natural language processing
[1184] server
[1185] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[1186] 5. Emotion recognition
[1187] server
[1188] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[1189] 6. Obtaining Market Data
[1190] server
[1191] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[1192] 7. Response generation and coordination
[1193] server
[1194] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[1195] 8. Visual Indications
[1196] server
[1197] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[1198] Terminal
[1199] The terminal displays the visual data sent from the server on a browser or dedicated application.
[1200] 9. Providing voice responses
[1201] Terminal
[1202] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[1203] Specific examples
[1204] Get stock quotes
[1205] 1. Users
[1206] User: "What's Apple's stock price right now?"
[1207] 2. Terminal
[1208] The terminal captures the user's voice and transmits it to the server.
[1209] 3. Server
[1210] The server converts the speech to text using a speech recognition system. The natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple." The emotion engine recognizes the user's emotions. Based on the intent, the server retrieves Apple's current stock price from a market data service.
[1211] 4. Terminal
[1212] Receive market data sent from the server and display it visually using D3.js, Chart.js, etc.
[1213] 5. Server
[1214] Generate a voice response message such as "Apple's current stock price is $145" and adjust the tone depending on the user's emotions. Generate a reassuring response such as "Apple's current stock price is $145, don't worry, there are good investment opportunities waiting for you."
[1215] 6. Terminal
[1216] A voice response is played to provide information to the user.
[1217] Providing investment advice
[1218] 1. Users
[1219] User: "What investment advice would you give me?"
[1220] 2. Terminal
[1221] The device captures the audio and sends it to the server.
[1222] 3. Server
[1223] The server converts the speech into text and uses a natural language processing system to identify the intent "provide investment advice." The emotion engine recognizes the user's emotions. The server then comprehensively analyzes market data and generates investment advice based on current market trends.
[1224] 4. Server
[1225] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions. It also suggests, "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and are good for the long term."
[1226] 5. Terminal
[1227] A voice response is played to provide investment advice to the user.
[1228] Prompt Sentence Examples
[1229] User says: "What is Apple's stock price right now?"
[1230] Prompt: "What is the process flow if the user says, 'What is Apple's stock price now?'"
[1231] In this way, the system integrates multiple technologies, including user emotion recognition, to provide a way for users to efficiently obtain and understand personalized investment information.
[1232] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1233] Step 1:
[1234] Acquiring voice input
[1235] User
[1236] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[1237] Terminal
[1238] The device uses a built-in microphone to capture the user's voice input, utilizing high-quality microphone and noise-canceling technology to collect voice data in real time.
[1239] Input: User's voice command
[1240] Output: Captured audio data
[1241] Step 2:
[1242] Sending audio data to the server
[1243] Terminal
[1244] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[1245] Input: Captured audio data
[1246] Output: Audio data sent to the server
[1247] Step 3:
[1248] Speech-to-text conversion
[1249] server
[1250] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[1251] Input: Transmitted audio data
[1252] Output: Converted text data
[1253] Step 4:
[1254] Analysis using natural language processing
[1255] server
[1256] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[1257] Input: Text data
[1258] Output: Identified intents and entities
[1259] Step 5:
[1260] emotion recognition
[1261] server
[1262] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[1263] Input: Audio and text data
[1264] Output: Recognized user emotion
[1265] Step 6:
[1266] Obtaining Market Data
[1267] server
[1268] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[1269] Input: Identified intents and entities
[1270] Output: Market data
[1271] Step 7:
[1272] Response generation and coordination
[1273] server
[1274] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[1275] Input: Market data and user sentiment
[1276] Output: Adjusted response message
[1277] Step 8:
[1278] Visual Indication
[1279] server
[1280] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[1281] Input: Market Data
[1282] Output: Visually displayable data
[1283] Terminal
[1284] The terminal displays the visual data sent from the server on a browser or dedicated application.
[1285] Input: Visually displayable data
[1286] Output: Graphs and charts that are displayed to the user
[1287] Step 9:
[1288] Providing voice responses
[1289] Terminal
[1290] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[1291] Input: Tailored response message
[1292] Output: The audio response the user hears
[1293] In this way, each processing step works in conjunction with the others, allowing the user to efficiently obtain and understand personalized investment information.
[1294] (Application example 2)
[1295] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1296] Conventional food delivery services have the problem that they do not suggest personalized menus based on the user's emotions, making it difficult to select the optimal menu based on the user's psychological state. Also, because they do not suggest optimal dishes based on the user's emotions, the improvement in satisfaction is limited.
[1297] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring menu information based on the analyzed intent, means for analyzing the user's emotion based on the intent and entities, means for adjusting the recommended menu based on the emotion information, means for visually displaying the acquired menu information, and means for outputting a response generated based on the menu information by voice. This makes it possible to propose an optimal menu based on the user's emotion, thereby significantly improving user satisfaction.
[1298] A "user" is someone who uses the system.
[1299] "Voice input" is data of the voice that the user utters to the system.
[1300] "Text" refers to speech input converted into character data.
[1301] An "intent" is the purpose or request that a user intends to make to a system.
[1302] "Entity" means a concrete object or subject that is relevant to an intent.
[1303] "Menu information" is data related to the provision of food and drink, and is a list of dishes and drinks suggested to the user.
[1304] "Emotion" refers to the user's mental or emotional state.
[1305] A "recommended menu" is a selection of food and drink options suggested based on the user's intent and emotions.
[1306] The "visual display means" is a method for presenting the acquired menu information and recommended menus to the user in a graphical format.
[1307] "Audio output means" refers to a method for transmitting acquired information or generated responses to the user in audio form.
[1308] The term "system" refers to a series of functional blocks that operate in combination with the above means.
[1309] This invention relates to a food delivery system that recognizes and analyzes voice input from a user to suggest menu items based on the user's emotions. Below, we will explain the program and its detailed processing for specifically implementing the invention, the hardware and software used, and specific examples of use.
[1310] Hardware and Software Used
[1311] Hardware: smartphone, microphone, server
[1312] software:
[1313] speech_recognition library: for speech recognition
[1314] textblob library: for natural language processing
[1315] sentiment_analysis module: for emotion recognition
[1316] delivery_service_api: To obtain and recommend menu information
[1317] Explanation of the program processing flow
[1318] 1. Voice Recognition
[1319] When a user speaks, the microphone connected to the smartphone captures the voice data. The speech_recognition library is used to convert this voice data into text. An example of voice input might be, "I'm tired today, so I want some food to relax me."
[1320] 2. Natural Language Processing
[1321] The converted text data is then analyzed in detail using the textblob library, where the user's intent and the desired object (entity) are extracted. For example, from this utterance, the intent "I want to relax" is identified, and the entity "food" is identified.
[1322] 3. Emotion recognition
[1323] The analyzed text data is then subjected to emotion recognition by the sentiment_analysis module, which determines the user's emotional state as "fatigue" and selects the most appropriate menu based on this.
[1324] 4. Menu information acquisition and recommendation
[1325] Based on the user's sentiment and intent, the delivery_service_api is used to obtain the most suitable menu information. For example, it may recommend herbal tea or Japanese soup as a relaxing food.
[1326] 5. Visual and audio output
[1327] The acquired menu information is sent from the server to the smartphone and displayed visually on the smartphone screen. This information is also output as a voice response. For example, a voice response such as "Today, we recommend relaxing herbal tea or Japanese-style soup."
[1328] Specific examples
[1329] Usage example 1
[1330] User: "I'm tired today and I want some food to help me relax."
[1331] Server: Converts speech to text, identifies intent "relaxation", and determines emotional state as "fatigue".
[1332] System: Retrieves recommended menu items such as herbal tea and Japanese-style soup, and presents them to the user visually and audibly.
[1333] Prompt Sentence Examples
[1334] "User: "I'm tired today and I want some food to help me relax."
[1335] Server: "Thank you for your hard work. To help you relax, we recommend the following: herbal tea, Japanese soup. Please let us know if there's anything else you'd like."
[1336] In this way, the present invention realizes a system that is sensitive to the user's emotions and provides optimal menu suggestions and food delivery services.
[1337] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1338] Step 1:
[1339] The user opens the smartphone application and speaks into the microphone, for example, saying, "I'm tired today, so I want some food to help me relax." The voice recognition module in the smartphone captures this voice data.
[1340] Input: User voice input
[1341] Output: Captured audio data
[1342] Step 2:
[1343] The device converts the captured voice data into text data using the speech_recognition library. The user's statement, "I'm tired today, so I want some food to relax me," is converted into text.
[1344] Input: Captured audio data
[1345] Output: Converted text data
[1346] Step 3:
[1347] The server receives the converted text data and parses it using the textblob library to extract the intent "I want to relax" and the entity "food."
[1348] Input: Converted text data
[1349] Output: Intents and entities
[1350] Step 4:
[1351] The server analyzes the user's emotion using the sentiment_analysis module based on the intent and entities. In this case, the user's emotion is determined to be "fatigue."
[1352] Input: Intents and Entities
[1353] Output: User's emotional information
[1354] Step 5:
[1355] The server uses the preprocessed intent and emotion information to retrieve appropriate menu information via the delivery_service_api. For example, herbal tea or Japanese soup is selected as relaxing food.
[1356] Input: Intent and emotion information
[1357] Output: Recommended menu information
[1358] Step 6:
[1359] The server converts the recommended menu information into a visually displayable format and sends it to the terminal, which then uses its visualization engine to display the received data on the screen.
[1360] Input: Recommended menu information
[1361] Output: Visualized menu information
[1362] Step 7:
[1363] The server generates a voice response message to convey to the user and converts it into voice data using a speech synthesis system. For example, a message such as "Today, we recommend a relaxing herbal tea or Japanese-style soup."
[1364] Input: Recommended menu information
[1365] Output: Voice response message
[1366] Step 8:
[1367] The terminal plays the voice response message sent from the server to provide information to the user, who can then check the details of the relaxing menu through the visual menu display and voice response.
[1368] Input: Voice response message
[1369] Output: Providing audio and visual information to the user
[1370] As described above, a series of processes are carried out, from recognizing the voice input from the user, analyzing emotions based on that, and recommending and providing the most suitable menu.
[1371] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1372] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1373] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1374] [Third embodiment]
[1375] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1376] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1377] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1378] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1379] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1380] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1381] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1382] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1383] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1384] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1385] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1386] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1387] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system works by combining speech recognition, natural language processing, data acquisition, and visual display.
[1388] ---
[1389] System Configuration
[1390] 1. Speech Recognition and Natural Language Processing
[1391] 1. Users
[1392] The user speaks a voice command, such as "What is the current Apple stock price?"
[1393] 2. Terminal
[1394] The device captures the user's voice input with a microphone and transmits it to the server.
[1395] 3. Server
[1396] The server uses a speech recognition system to convert the voice input into text, then uses a natural language processing system to parse the text and identify intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[1397] 2. Obtaining market data
[1398] 1. Server
[1399] Based on the intent and entity, send a request to a market data service to retrieve specific stock information (e.g., "Apple" stock price).
[1400] 3. Visual display and audio response
[1401] 1. Server
[1402] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[1403] 2. Terminal
[1404] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[1405] 3. Server
[1406] The server generates a response message to provide the user with the stock price information by voice, and creates a voice response using a voice synthesis system.
[1407] 4. Terminal
[1408] The terminal plays the voice response sent from the server and provides the information to the user.
[1409] ---
[1410] Specific examples
[1411] 1. Obtaining stock price information
[1412] 1. Users
[1413] User: "What's Apple's stock price right now?"
[1414] 2. Terminal
[1415] The terminal captures the user's voice and transmits it to the server.
[1416] 3. Server
[1417] The server converts the speech into text using a speech recognition system.
[1418] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[1419] Based on the intent, retrieve Apple's current stock price from a market data service.
[1420] 4. Terminal
[1421] Receive market data sent from the server and display it visually in the data visualizer.
[1422] 5. Server
[1423] Generate a voice response message saying "Apple's current stock price is $145" and create the voice response using a speech synthesis system.
[1424] 6. Terminal
[1425] A voice response is played to provide information to the user.
[1426] 2. Providing investment advice
[1427] 1. Users
[1428] User: "What investment advice would you give me?"
[1429] 2. Terminal
[1430] The device captures the audio and sends it to the server.
[1431] 3. Server
[1432] The server converts the speech to text and uses a natural language processing system to identify the intent "provide investment advice."
[1433] Analyze market data and generate investment advice based on current market trends.
[1434] Generate a voice response message that reads, "Based on current market trends, we recommend investing in technology stocks," and use a speech synthesis system to create the voice response.
[1435] 4. Terminal
[1436] A voice response is played to provide investment advice to the user.
[1437] As described above, the present invention provides a system that combines speech recognition, natural language processing, market data acquisition and visual display to enable users to efficiently acquire and understand investment information.
[1438] The processing flow will be explained below.
[1439] When obtaining stock price information
[1440] Step 1:
[1441] The user issues a voice command such as, "What is the current Apple stock price?"
[1442] Step 2:
[1443] The device captures the user's voice with a microphone and converts it into a digital signal.
[1444] Step 3:
[1445] The device transmits the captured audio data to the server.
[1446] Step 4:
[1447] The server analyzes the received voice data using a voice recognition system and converts it into text.
[1448] python
[1449] text_query = self.voice_recognition_system.transcribe(audio_input)
[1450] Step 5:
[1451] The server analyzes the text using a natural language processing system to identify intent and entities.
[1452] python
[1453] intent, entities = self.natural_language_processing.analyze_query(text_query)
[1454] Step 6:
[1455] Based on the analysis results, the server determines that the user's intent is to "get stock price information." At the same time, it extracts the entity "Apple."
[1456] Step 7:
[1457] The server sends a request to a market data service to get current stock price information for Apple.
[1458] python
[1459] market_data = self.market_data_service.get_stock_data(stock_ticker)
[1460] Step 8:
[1461] The server analyzes the acquired market data and sends it to the terminal.
[1462] Step 9:
[1463] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[1464] python
[1465] self.data_visualizer.visualize_data(market_data)
[1466] Step 10:
[1467] The server generates a response message to provide to the user: "Apple's current stock price is $145."
[1468] python
[1469] self.speak_response(f"Here is the market data for {stock_ticker}")
[1470] Step 11:
[1471] The server converts the generated message into a voice response using a speech synthesis system.
[1472] Step 12:
[1473] The terminal plays the voice response sent from the server and provides the information to the user.
[1474] python
[1475] play_audio_response(message)
[1476] ---
[1477] When providing investment advice
[1478] Step 1:
[1479] The user issues the voice command "Give me some investment advice."
[1480] Step 2:
[1481] The device captures the user's voice with a microphone and converts it into a digital signal.
[1482] Step 3:
[1483] The device transmits the captured audio data to the server.
[1484] Step 4:
[1485] The server analyzes the received voice data using a voice recognition system and converts it into text.
[1486] python
[1487] text_query = self.voice_recognition_system.transcribe(audio_input)
[1488] Step 5:
[1489] The server analyzes the text using a natural language processing system to identify the intent.
[1490] python
[1491] intent, entities = self.natural_language_processing.analyze_query(text_query)
[1492] Step 6:
[1493] Based on the analysis results, the server determines that the user's intent is to "provide investment advice."
[1494] Step 7:
[1495] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[1496] python
[1497] analysis = self.market_data_service.analyze_market()
[1498] Step 8:
[1499] The server generates a response message to provide to the user, which might say, "Based on current market trends, we recommend investing in technology stocks."
[1500] python
[1501] self.speak_response(f"My investment advice based on current market trends is: {analysis}")
[1502] Step 9:
[1503] The server converts the generated message into a voice response using a speech synthesis system.
[1504] Step 10:
[1505] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[1506] python
[1507] play_audio_response(message)
[1508] Example 1
[1509] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1510] Conventional investment information acquisition systems make it difficult for users to accurately and quickly acquire and understand market data in real time. Furthermore, since an intuitive method for acquiring information through voice commands was uncommon, improving the user experience was a challenge. Furthermore, the functionality for visually representing acquired data and providing voice responses was insufficient, making it difficult to meet the diverse needs of users.
[1511] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1512] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for outputting a response generated based on the market data by voice, means for using an existing voice recognition system and a natural language processing model for performing voice recognition and natural language processing, means for sending a request to an external data providing service to acquire market data and deserializing the result, means for using a data visualization tool to visually display the acquired market data in the form of graphs and charts, and means for using a voice synthesis system to generate a voice response, thereby enabling a user to intuitively acquire real-time market data through voice commands and quickly understand information visually and audibly.
[1513] "Means for recognizing voice input from a user" refers to a function that uses a voice input device such as a microphone to capture voice commands spoken by a user and obtain them in digital form.
[1514] "Means for converting said voice input to text" refers to a function that uses a voice recognition system to convert captured voice data into text form.
[1515] "Means for analyzing intent and entities from the text" refers to a function that uses a natural language processing model to identify objects (entities) related to a user's intent from text data.
[1516] "Means for obtaining market data based on the analyzed intent" refers to the function of sending a request derived from the analysis result to a market data providing service and obtaining the required data.
[1517] "Means for visually displaying the acquired market data" refers to the ability to use data visualization tools to display the acquired market data to a user in a visual format, such as a graph or chart.
[1518] "Means for outputting a response generated based on the market data in voice" refers to a function for generating a response message based on the acquired market data and providing it to the user in voice format using a voice synthesis system.
[1519] "Means of using existing speech recognition systems and natural language processing models for speech recognition and natural language processing" refers to the ability to apply existing speech recognition and natural language processing technologies to efficiently convert speech data into text and analyze the text.
[1520] "Means for sending requests to external data providers to retrieve market data and deserializing the results" means the functionality for sending requests to external market data providers to retrieve data and parsing the retrieved data into an appropriate format.
[1521] "Means of using data visualization tools to visually display acquired market data in the form of graphs and charts" refers to the ability to display data in a visually easy-to-understand format using dedicated visualization software or libraries.
[1522] "Means for using a speech synthesis system to generate a voice response" refers to the ability to use speech synthesis technology to convert a text message into speech form and provide it to a user.
[1523] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system combines speech recognition, natural language processing, data acquisition, and visual display. Specifically, the system integrates a speech recognition system, a natural language processing model, a market data provider, a data visualization tool, and a speech synthesis system.
[1524] System Configuration
[1525] 1. Speech Recognition and Natural Language Processing
[1526] 1. Users
[1527] The user speaks a voice command such as "What is the current Apple stock price?"
[1528] 2. Terminal
[1529] The device captures the user's voice input using the device's built-in microphone, which may include a smartphone or smart speaker.
[1530] The device sends the captured audio data to the server using an HTTP request.
[1531] 3. Server
[1532] The server uses a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text.
[1533] The server uses a natural language processing model (e.g., OpenAI GPT-3) to identify intents (user intentions) and entities (specific targets) from the text data.
[1534] 2. Obtaining market data
[1535] 1. Server
[1536] The server sends a request based on the intent and entity to an external market data provider (e.g., Yahoo Finance API) using the HTTP GET method.
[1537] The server receives the data returned by the API and deserializes it in JSON format, which contains the latest market information for a specific financial instrument.
[1538] 3. Visual display and audio response
[1539] 1. Server
[1540] The server analyzes the acquired market data, formats it into a format that is easy for users to understand, and sends the formatted data to the terminal.
[1541] 2. Terminal
[1542] The terminal receives the data sent from the server and uses data visualization tools (e.g., Plotly or Matplotlib) to generate and visually display graphs and charts, including line graphs and bar graphs.
[1543] 3. Server
[1544] The server generates a response message such as "Apple's current stock price is $145" and creates the audio data using a speech synthesis system (e.g., Amazon Polly).
[1545] 4. Terminal
[1546] The terminal plays the audio data sent from the server and provides the information to the user, without the user having to perform any specific operation.
[1547] Specific examples
[1548] Get stock quotes
[1549] 1. Users
[1550] User: "What's Apple's stock price right now?"
[1551] 2. Terminal
[1552] The terminal captures the user's voice and transmits the voice data to the server.
[1553] 3. Server
[1554] The server converts the speech to text using the Google Cloud Speech-to-Text API, then uses OpenAI's GPT-3 to identify the intent "get stock quotes" and the entity "Apple" from the text.
[1555] The server sends a request to the Yahoo Finance API to get Apple's current stock price.
[1556] 4. Terminal
[1557] The terminal receives market data sent from the server and displays it as a line graph using Plotly.
[1558] 5. Server
[1559] The server generates a text message saying, "Apple's current stock price is $145," and creates voice data using Amazon Polly.
[1560] 6. Terminal
[1561] The terminal plays back the audio data and provides information to the user.
[1562] This invention allows users to intuitively obtain real-time market data through voice commands, providing information quickly and easily understood both visually and audibly. As described above, the present invention combines voice recognition, natural language processing, market data acquisition and visual display to provide a system that allows users to efficiently obtain and understand investment information.
[1563] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1564] Step 1: Capture and send audio input
[1565] 1. Users
[1566] The user speaks a voice command such as "What is the current Apple stock price?"
[1567] Input: User's voice command
[1568] Output: Audio data
[1569] 2. Terminal
[1570] The terminal uses the device's built-in microphone to capture the user's voice.
[1571] The captured audio is converted into a digital format and data compressed.
[1572] The terminal transmits the compressed audio data to the server as an HTTP request.
[1573] Input: User's voice data
[1574] Output: HTTP request to the server
[1575] Step 2: Convert audio data to text
[1576] 1. Server
[1577] The server sends the received audio data to the Google Cloud Speech-to-Text API.
[1578] The Google Cloud Speech-to-Text API converts the audio data into text format and returns it to the server.
[1579] Input: Compressed audio data
[1580] Output: Text data
[1581] Step 3: Identifying intents and entities through natural language processing
[1582] 1. Server
[1583] The server inputs the acquired text data into the OpenAI GPT-3 model to analyze intent and entities.
[1584] GPT-3 parses text and extracts intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[1585] Input: Text data
[1586] Output: Intents and entities
[1587] Step 4: Obtain market data
[1588] 1. Server
[1589] The server sends an HTTP GET request to the Yahoo Finance API to retrieve market data for a particular financial instrument.
[1590] The Yahoo Finance API returns relevant market data in JSON format upon request.
[1591] The server deserializes the received JSON data and extracts the necessary information.
[1592] Input: Intents and Entities
[1593] Output: Market data
[1594] Step 5: Format and transmit market data
[1595] 1. Server
[1596] The server formats the acquired market data into a format that is easy for users to understand. The formatted data is organized, for example, as time-series data of stock prices.
[1597] The server transmits the formatted market data to the terminal.
[1598] Input: Market Data
[1599] Output: Formatted market data
[1600] Step 6: Generate a visual representation
[1601] 1. Terminal
[1602] The terminal uses the received market data to generate graphs and charts using data visualization tools such as Plotly and Matplotlib.
[1603] The generated visual representation is displayed on the screen of the terminal.
[1604] Input: Formatted market data
[1605] Output: Graphs and charts
[1606] Step 7: Generate and play a voice response
[1607] 1. Server
[1608] The server generates a response message such as "Apple's current stock price is $145."
[1609] The response message is sent to Amazon Polly and converted into voice data.
[1610] The generated voice data is transmitted to the terminal.
[1611] Input: Response message
[1612] Output: Audio data
[1613] 2. Terminal
[1614] The terminal plays back the received audio data and provides the information to the user.
[1615] Input: Audio data
[1616] Output: Audio information
[1617] This specific processing step allows users to obtain real-time market analysis and stock information through voice commands, and to understand it visually and audibly. The system provides an intuitive and easy-to-use interface for users.
[1618] (Application example 1)
[1619] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1620] In modern society, there is a demand for quick and efficient understanding of security risks and market trends. It is also important to enable users to easily access the information they need in real time by displaying that information visually and outputting it audibly. The present invention aims to provide a system that supports the acquisition and understanding of such information, thereby enabling users to quickly grasp risk management and market trends.
[1621] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1622] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for audio outputting a response generated based on the market data, means for acquiring data for analyzing crisis management and market trends in real time, means for visually displaying crisis management information from the data, and means for audio outputting a response generated based on the crisis management information, thereby enabling a user to acquire market analysis and security-related information in real time through voice commands and respond quickly.
[1623] "Means for recognizing voice input from a user" refers to technology that allows a device to receive speech uttered by a user and process it as a digital signal.
[1624] The "means for converting said voice input into text" refers to a technology for converting a voice signal into a string of characters, typically using a voice recognition algorithm.
[1625] The "means for analyzing intent and entities from the text" is a natural language processing technology for extracting a user's intent (intent) and specific target elements (entities) from text data.
[1626] The "means for acquiring market data based on the analyzed intent" is a technology for acquiring market data corresponding to the intent from an external database or API.
[1627] The "means for visually displaying acquired market data" refers to technology for displaying market data on a screen in the form of graphs, charts, etc.
[1628] The "means for outputting a response generated based on the market data as voice" is a technology that uses voice synthesis technology to provide a text response created based on market data to a user.
[1629] "Data acquisition means for analyzing crisis management and market trends in real time" refers to technology for collecting and analyzing current crisis management information and market trends in real time.
[1630] The "means for visually displaying crisis management information from the data" refers to a technology that visually displays collected and analyzed crisis management information in an easy-to-understand manner for users.
[1631] The "means for outputting a response generated based on the crisis management information as voice" is a technology that uses a voice synthesis technology to provide a text response generated based on the crisis management information to a user.
[1632] System Configuration
[1633] 1. Speech Recognition and Natural Language Processing
[1634] 1. Users
[1635] The user uses their smartphone or smart glasses to issue a voice command such as "Tell me the current crisis management information."
[1636] 2. Terminal
[1637] The device (smartphone or smart glasses) captures the user's voice input with a microphone and sends it to a server.
[1638] 3. Server
[1639] The server converts the voice input into text using the Google Speech-to-Text API, then uses a natural language processing system such as spaCy to parse the text and identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends").
[1640] 2. Obtaining market data and crisis management information
[1641] 1. Server
[1642] Use the News API and custom web scrapers to gather up-to-date crisis and market data based on intent and entities.
[1643] 3. Visual display and audio response
[1644] 1. Server
[1645] The acquired information is analyzed and presented in a visually easy-to-understand format using Plotly, and a voice response is generated using the Google Text-to-Speech API, stating, "The current trend in crisis management is an increase in phishing attacks."
[1646] 2. Terminal
[1647] The terminal displays the visual data sent from the server on its screen and plays back audio responses to provide information to the user.
[1648] Specific examples
[1649] 1. Speech Recognition and Natural Language Processing
[1650] A user utters, "What are the current trends in crisis management?"
[1651] The terminal captures the user's voice and sends it to the server.
[1652] The server converts the speech into text, and a natural language processing system identifies "obtaining crisis management information" and "current trends."
[1653] 2. Data Acquisition
[1654] The server collects crisis management information using the News API and custom web scrapers.
[1655] 3. Visual display and audio response
[1656] The server parses the information, formats it for visual display using Plotly, and generates an audio response using the Google Text-to-Speech API.
[1657] The terminal displays the displayed data on a screen and plays audio responses to provide information to the user.
[1658] Prompt Sentence Examples
[1659] Below are some example prompts to be fed into the generative AI model:
[1660] When a user says, "What are the current crisis management trends?" the application performs the following steps:
[1661] 1. Use the Google Speech-to-Text API to convert voice commands into text.
[1662] 2. Using spaCy, we analyzed the intent “Get crisis management information” and the entity “Current trends” from the text.
[1663] 3. Use the News API or a custom web scraper to gather the latest crisis management information.
[1664] 4. Use Plotly to generate visually easy-to-understand graphs and charts and display them on the user screen.
[1665] 5. Using the Google Text-to-Speech API, a voice response was generated that said, "The current trend in crisis management is an increase in phishing attacks," and played on a smartphone or smart glasses.
[1666] This allows users to easily obtain and understand real-time security information through voice commands.
[1667] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1668] Step 1:
[1669] The user says, "Tell me about the current trends in crisis management." The input is the user's voice command, and the output is voice data. This allows the system to recognize the user's request and begin processing.
[1670] Step 2:
[1671] The device captures the user's voice input with a microphone and sends it to the server. The input is voice data, and the output is a digital signal. The voice data is transmitted to the server via the Internet.
[1672] Step 3:
[1673] The server converts the voice input into text using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This process recognizes the voice as text.
[1674] Step 4:
[1675] The server uses a natural language processing system such as spaCy to identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends") from the text data. The input is text data, and the output is the parsed intent and entities. This clarifies the user's intent and specific request.
[1676] Step 5:
[1677] The server collects the latest crisis management information using the News API or a custom web scraper based on the intent and entity. The input is the intent and entity, and the output is the crisis management information data. This allows the appropriate information to be retrieved.
[1678] Step 6:
[1679] The server analyzes the acquired information and uses Plotly to generate graphs and charts in a visually easy-to-understand format. The input is crisis management information data, and the output is visual display data. This converts the information into a format that is easy for users to understand.
[1680] Step 7:
[1681] The server uses the Google Text-to-Speech API to generate a voice response stating, "Current crisis management trends include an increase in phishing attacks." The input is the text response message, and the output is the audio data. This provides audio information along with visual information.
[1682] Step 8:
[1683] The terminal displays the visual data sent from the server on the screen and plays back audio responses to provide information to the user. The input is visual display data and audio data, and the output is the displayed data and played audio. This allows the user to obtain and understand the information they need in real time.
[1684] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1685] The present invention provides more personalized market data and investment advice by combining a system that recognizes user voice input, converts text, and processes natural language with an emotion engine. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[1686] ---
[1687] System Configuration
[1688] 1. Speech Recognition and Natural Language Processing
[1689] 1. Users
[1690] Users can speak voice commands such as "What is the current Apple stock price?" or "I'd like some investment advice."
[1691] 2. Terminal
[1692] The device captures the user's voice input with a microphone and transmits it to the server.
[1693] 3. Server
[1694] The server uses a speech recognition system to convert the voice input into text, and then uses a natural language processing system to parse the text and identify intents and entities.
[1695] 2. Emotion recognition
[1696] 1. Server
[1697] The server recognizes the user's emotions using an emotion engine based on the voice recognition data and text data.
[1698] 3. Obtaining market data
[1699] 1. Server
[1700] Based on the intent and entity, a request is sent to the market data service to retrieve specific stock information or market data.
[1701] 4. Regulating responses based on emotions
[1702] 1. Server
[1703] The emotion engine adjusts the tone and content of responses based on the user's emotions. For example, if it detects that the user is feeling stressed, it will generate a more reassuring response.
[1704] 5. Visual display and audio response
[1705] 1. Server
[1706] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[1707] 2. Terminal
[1708] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[1709] 3. Server
[1710] The server generates a response message to be presented to the user and uses a speech synthesis system to create a voice response.
[1711] 4. Terminal
[1712] The terminal plays the voice response sent from the server and provides the information to the user.
[1713] ---
[1714] Specific examples
[1715] 1. Obtaining stock price information
[1716] 1. Users
[1717] User: "What's Apple's stock price right now?"
[1718] 2. Terminal
[1719] The terminal captures the user's voice and transmits it to the server.
[1720] 3. Server
[1721] The server converts the speech into text using a speech recognition system.
[1722] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[1723] Recognize user emotions with an emotion engine.
[1724] Based on the intent, retrieve Apple's current stock price from a market data service.
[1725] 4. Terminal
[1726] Receive market data sent from the server and display it visually in the data visualizer.
[1727] 5. Server
[1728] Generate a voice response message saying "Apple's current stock price is $145" and adjust the tone depending on the user's emotion.
[1729] Generate a reassuring response like, "Apple stock is currently at $145, don't worry, there are good investment opportunities waiting for you."
[1730] 6. Terminal
[1731] A voice response is played to provide information to the user.
[1732] 2. Providing investment advice
[1733] 1. Users
[1734] User: "What investment advice would you give me?"
[1735] 2. Terminal
[1736] The device captures the audio and sends it to the server.
[1737] 3. Server
[1738] The server converts the speech to text and uses a natural language processing system to identify the intent: "Provide investment advice."
[1739] Recognize user emotions with an emotion engine.
[1740] Comprehensively analyze market data and generate investment advice based on current market trends.
[1741] 4. Server
[1742] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions.
[1743] "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and a good long-term view," he suggests.
[1744] 5. Terminal
[1745] A voice response is played to provide investment advice to the user.
[1746] As described above, the present invention provides a system that enables users to efficiently acquire and understand personalized investment information by integrating and operating speech recognition, natural language processing, market data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[1747] The processing flow will be explained below.
[1748] When obtaining stock price information
[1749] Step 1:
[1750] The user issues a voice command such as, "What is the current Apple stock price?"
[1751] Step 2:
[1752] The device captures the user's voice with a microphone and converts it into a digital signal.
[1753] Step 3:
[1754] The device transmits the captured audio data to the server.
[1755] Step 4:
[1756] The server analyzes the received voice data using a voice recognition system and converts it into text.
[1757] python
[1758] text_query = self.voice_recognition_system.transcribe(audio_input)
[1759] Step 5:
[1760] The server analyzes the text using a natural language processing system to identify intent and entities.
[1761] python
[1762] intent, entities = self.natural_language_processing.analyze_query(text_query)
[1763] Step 6:
[1764] The server verifies that the intent parsed from the text is "Get stock quotes" and identifies "Apple" as the entity.
[1765] Step 7:
[1766] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[1767] python
[1768] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[1769] Step 8:
[1770] The server takes into account the perceived sentiment and sends a request to a market data service to get current stock price information for Apple.
[1771] python
[1772] market_data = self.market_data_service.get_stock_data(stock_ticker)
[1773] Step 9:
[1774] The server analyzes the acquired market data and sends it to the terminal.
[1775] Step 10:
[1776] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[1777] python
[1778] self.data_visualizer.visualize_data(market_data)
[1779] Step 11:
[1780] The server generates a response message to be delivered to the user, such as "Apple's current stock price is $145," and adjusts the tone according to the user's emotion as determined by the emotion engine.
[1781] python
[1782] response_message = f"Apple's current stock price is ${market_data['price']}."
[1783] response_message = self.emotion_engine.adjust_tone(response_message, user_emotion)
[1784] Step 12:
[1785] The server converts the generated message into a voice response using a speech synthesis system.
[1786] Step 13:
[1787] The terminal plays the voice response sent from the server and provides the information to the user.
[1788] python
[1789] play_audio_response(response_message)
[1790] ---
[1791] When providing investment advice
[1792] Step 1:
[1793] The user issues the voice command "Give me some investment advice."
[1794] Step 2:
[1795] The device captures the user's voice with a microphone and converts it into a digital signal.
[1796] Step 3:
[1797] The device transmits the captured audio data to the server.
[1798] Step 4:
[1799] The server analyzes the received voice data using a voice recognition system and converts it into text.
[1800] python
[1801] text_query = self.voice_recognition_system.transcribe(audio_input)
[1802] Step 5:
[1803] The server analyzes the text using a natural language processing system to identify the intent.
[1804] python
[1805] intent, entities = self.natural_language_processing.analyze_query(text_query)
[1806] Step 6:
[1807] The server determines that the intent parsed from the text is "provide investment advice."
[1808] Step 7:
[1809] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[1810] python
[1811] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[1812] Step 8:
[1813] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[1814] python
[1815] analysis = self.market_data_service.analyze_market()
[1816] investment_advice = f"My investment advice based on current market trends is: {analysis}"
[1817] Step 9:
[1818] The server adjusts the content and tone of the investment advice based on the perceived user sentiment.
[1819] python
[1820] investment_advice = self.emotion_engine.adjust_tone(investment_advice, user_emotion)
[1821] Step 10:
[1822] The server converts the generated investment advice message into a voice response using a voice synthesis system.
[1823] Step 11:
[1824] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[1825] python
[1826] play_audio_response(investment_advice)
[1827] Through the above steps, the present invention combines voice input, natural language processing, emotion recognition, data acquisition, and visual display to realize a system that provides users with personalized investment information, allowing them to acquire and understand investment information more quickly and efficiently.
[1828] Example 2
[1829] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1830] Conventional systems simply convert voice input from users into text and provide market data based on specific intents and entities, but are unable to generate personalized responses that take into account the user's emotions. This has led to the challenge of not being able to provide responses with the appropriate tone and content according to the user's emotions.
[1831] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1832] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for recognizing user emotions from the voice input and text data, means for acquiring market data based on the analyzed intent and entities, means for visually displaying the acquired market data, means for adjusting the tone and content of a response generated based on the user emotions, and means for outputting the market data and the adjusted response by voice, thereby enabling a user to efficiently acquire and understand personalized investment information.
[1833] A "user" is a human user who requests information by speaking voice commands to the system.
[1834] "Voice input" is voice data uttered by the user.
[1835] "Speech recognition" is a technology that analyzes a user's voice input and converts the content into text data.
[1836] "Text conversion" is the process of converting audio data into text data.
[1837] "Intent" refers to identifying the user's intention or purpose for text input.
[1838] An "entity" is an element that extracts and identifies specific keywords or items within a text.
[1839] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state from voice input or text data.
[1840] "Market data" refers to financial-related data such as specific stock information and economic indicators.
[1841] "Visual display" refers to converting acquired market data into a visually understandable format, such as a graph or chart, and displaying it.
[1842] "Response generation" is the process of creating a response message to be provided to a user based on acquired market data and user sentiment.
[1843] "Audio output" is the process of playing back the generated response message as audio.
[1844] This invention relates to a system that recognizes user voice input, performs text conversion and natural language processing, and combines it with an emotion engine to provide personalized market data and investment advice. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[1845] System configuration
[1846] The system includes the following main components:
[1847] 1. Acquiring voice input
[1848] User
[1849] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[1850] Terminal
[1851] The device will use a built-in microphone to capture the user's voice input, and at this stage it is advisable to use a high-quality microphone and noise-cancelling technology.
[1852] 2. Sending audio data to the server
[1853] Terminal
[1854] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[1855] 3. Speech to text conversion
[1856] server
[1857] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[1858] 4. Analysis using natural language processing
[1859] server
[1860] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[1861] 5. Emotion recognition
[1862] server
[1863] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[1864] 6. Obtaining Market Data
[1865] server
[1866] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[1867] 7. Response generation and coordination
[1868] server
[1869] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[1870] 8. Visual Indications
[1871] server
[1872] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[1873] Terminal
[1874] The terminal displays the visual data sent from the server on a browser or dedicated application.
[1875] 9. Providing voice responses
[1876] Terminal
[1877] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[1878] Specific examples
[1879] Get stock quotes
[1880] 1. Users
[1881] User: "What's Apple's stock price right now?"
[1882] 2. Terminal
[1883] The terminal captures the user's voice and transmits it to the server.
[1884] 3. Server
[1885] The server converts the speech to text using a speech recognition system. The natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple." The emotion engine recognizes the user's emotions. Based on the intent, the server retrieves Apple's current stock price from a market data service.
[1886] 4. Terminal
[1887] Receive market data sent from the server and display it visually using D3.js, Chart.js, etc.
[1888] 5. Server
[1889] Generate a voice response message such as "Apple's current stock price is $145" and adjust the tone depending on the user's emotions. Generate a reassuring response such as "Apple's current stock price is $145, don't worry, there are good investment opportunities waiting for you."
[1890] 6. Terminal
[1891] A voice response is played to provide information to the user.
[1892] Providing investment advice
[1893] 1. Users
[1894] User: "What investment advice would you give me?"
[1895] 2. Terminal
[1896] The device captures the audio and sends it to the server.
[1897] 3. Server
[1898] The server converts the speech into text and uses a natural language processing system to identify the intent "provide investment advice." The emotion engine recognizes the user's emotions. The server then comprehensively analyzes market data and generates investment advice based on current market trends.
[1899] 4. Server
[1900] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions. It also suggests, "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and are good for the long term."
[1901] 5. Terminal
[1902] A voice response is played to provide investment advice to the user.
[1903] Prompt Sentence Examples
[1904] User says: "What is Apple's stock price right now?"
[1905] Prompt: "What is the process flow if the user says, 'What is Apple's stock price now?'"
[1906] In this way, the system integrates multiple technologies, including user emotion recognition, to provide a way for users to efficiently obtain and understand personalized investment information.
[1907] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1908] Step 1:
[1909] Acquiring voice input
[1910] User
[1911] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[1912] Terminal
[1913] The device uses a built-in microphone to capture the user's voice input, utilizing high-quality microphone and noise-canceling technology to collect voice data in real time.
[1914] Input: User's voice command
[1915] Output: Captured audio data
[1916] Step 2:
[1917] Sending audio data to the server
[1918] Terminal
[1919] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[1920] Input: Captured audio data
[1921] Output: Audio data sent to the server
[1922] Step 3:
[1923] Speech-to-text conversion
[1924] server
[1925] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[1926] Input: Transmitted audio data
[1927] Output: Converted text data
[1928] Step 4:
[1929] Analysis using natural language processing
[1930] server
[1931] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[1932] Input: Text data
[1933] Output: Identified intents and entities
[1934] Step 5:
[1935] emotion recognition
[1936] server
[1937] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[1938] Input: Audio and text data
[1939] Output: Recognized user emotion
[1940] Step 6:
[1941] Obtaining Market Data
[1942] server
[1943] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[1944] Input: Identified intents and entities
[1945] Output: Market data
[1946] Step 7:
[1947] Response generation and coordination
[1948] server
[1949] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[1950] Input: Market data and user sentiment
[1951] Output: Adjusted response message
[1952] Step 8:
[1953] Visual Indication
[1954] server
[1955] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[1956] Input: Market Data
[1957] Output: Visually displayable data
[1958] Terminal
[1959] The terminal displays the visual data sent from the server on a browser or dedicated application.
[1960] Input: Visually displayable data
[1961] Output: Graphs and charts that are displayed to the user
[1962] Step 9:
[1963] Providing voice responses
[1964] Terminal
[1965] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[1966] Input: Tailored response message
[1967] Output: The audio response the user hears
[1968] In this way, each processing step works in conjunction with the others, allowing the user to efficiently obtain and understand personalized investment information.
[1969] (Application example 2)
[1970] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1971] Conventional food delivery services have the problem that they do not suggest personalized menus based on the user's emotions, making it difficult to select the optimal menu based on the user's psychological state. Also, because they do not suggest optimal dishes based on the user's emotions, the improvement in satisfaction is limited.
[1972] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring menu information based on the analyzed intent, means for analyzing the user's emotion based on the intent and entities, means for adjusting the recommended menu based on the emotion information, means for visually displaying the acquired menu information, and means for outputting a response generated based on the menu information by voice. This makes it possible to propose an optimal menu based on the user's emotion, thereby significantly improving user satisfaction.
[1973] A "user" is someone who uses the system.
[1974] "Voice input" is data of the voice that the user utters to the system.
[1975] "Text" refers to speech input converted into character data.
[1976] An "intent" is the purpose or request that a user intends to make to a system.
[1977] "Entity" means a concrete object or subject that is relevant to an intent.
[1978] "Menu information" is data related to the provision of food and drink, and is a list of dishes and drinks suggested to the user.
[1979] "Emotion" refers to the user's mental or emotional state.
[1980] A "recommended menu" is a selection of food and drink options suggested based on the user's intent and emotions.
[1981] The "visual display means" is a method for presenting the acquired menu information and recommended menus to the user in a graphical format.
[1982] "Audio output means" refers to a method for transmitting acquired information or generated responses to the user in audio form.
[1983] The term "system" refers to a series of functional blocks that operate in combination with the above means.
[1984] This invention relates to a food delivery system that recognizes and analyzes voice input from a user to suggest menu items based on the user's emotions. Below, we will explain the program and its detailed processing for specifically implementing the invention, the hardware and software used, and specific examples of use.
[1985] Hardware and Software Used
[1986] Hardware: smartphone, microphone, server
[1987] software:
[1988] speech_recognition library: for speech recognition
[1989] textblob library: for natural language processing
[1990] sentiment_analysis module: for emotion recognition
[1991] delivery_service_api: To obtain and recommend menu information
[1992] Explanation of the program processing flow
[1993] 1. Voice Recognition
[1994] When a user speaks, the microphone connected to the smartphone captures the voice data. The speech_recognition library is used to convert this voice data into text. An example of voice input might be, "I'm tired today, so I want some food to relax me."
[1995] 2. Natural Language Processing
[1996] The converted text data is then analyzed in detail using the textblob library, where the user's intent and the desired object (entity) are extracted. For example, from this utterance, the intent "I want to relax" is identified, and the entity "food" is identified.
[1997] 3. Emotion recognition
[1998] The analyzed text data is then subjected to emotion recognition by the sentiment_analysis module, which determines the user's emotional state as "fatigue" and selects the most appropriate menu based on this.
[1999] 4. Menu information acquisition and recommendation
[2000] Based on the user's sentiment and intent, the delivery_service_api is used to obtain the most suitable menu information. For example, it may recommend herbal tea or Japanese soup as a relaxing food.
[2001] 5. Visual and audio output
[2002] The acquired menu information is sent from the server to the smartphone and displayed visually on the smartphone screen. This information is also output as a voice response. For example, a voice response such as "Today, we recommend relaxing herbal tea or Japanese-style soup."
[2003] Specific examples
[2004] Usage example 1
[2005] User: "I'm tired today and I want some food to help me relax."
[2006] Server: Converts speech to text, identifies intent "relaxation", and determines emotional state as "fatigue".
[2007] System: Retrieves recommended menu items such as herbal tea and Japanese-style soup, and presents them to the user visually and audibly.
[2008] Prompt Sentence Examples
[2009] "User: "I'm tired today and I want some food to help me relax."
[2010] Server: "Thank you for your hard work. To help you relax, we recommend the following: herbal tea, Japanese soup. Please let us know if there's anything else you'd like."
[2011] In this way, the present invention realizes a system that is sensitive to the user's emotions and provides optimal menu suggestions and food delivery services.
[2012] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2013] Step 1:
[2014] The user opens the smartphone application and speaks into the microphone, for example, saying, "I'm tired today, so I want some food to help me relax." The voice recognition module in the smartphone captures this voice data.
[2015] Input: User voice input
[2016] Output: Captured audio data
[2017] Step 2:
[2018] The device converts the captured voice data into text data using the speech_recognition library. The user's statement, "I'm tired today, so I want some food to relax me," is converted into text.
[2019] Input: Captured audio data
[2020] Output: Converted text data
[2021] Step 3:
[2022] The server receives the converted text data and parses it using the textblob library to extract the intent "I want to relax" and the entity "food."
[2023] Input: Converted text data
[2024] Output: Intents and entities
[2025] Step 4:
[2026] The server analyzes the user's emotion using the sentiment_analysis module based on the intent and entities. In this case, the user's emotion is determined to be "fatigue."
[2027] Input: Intents and Entities
[2028] Output: User's emotional information
[2029] Step 5:
[2030] The server uses the preprocessed intent and emotion information to retrieve appropriate menu information via the delivery_service_api. For example, herbal tea or Japanese soup is selected as relaxing food.
[2031] Input: Intent and emotion information
[2032] Output: Recommended menu information
[2033] Step 6:
[2034] The server converts the recommended menu information into a visually displayable format and sends it to the terminal, which then uses its visualization engine to display the received data on the screen.
[2035] Input: Recommended menu information
[2036] Output: Visualized menu information
[2037] Step 7:
[2038] The server generates a voice response message to convey to the user and converts it into voice data using a speech synthesis system. For example, a message such as "Today, we recommend a relaxing herbal tea or Japanese-style soup."
[2039] Input: Recommended menu information
[2040] Output: Voice response message
[2041] Step 8:
[2042] The terminal plays the voice response message sent from the server to provide information to the user, who can then check the details of the relaxing menu through the visual menu display and voice response.
[2043] Input: Voice response message
[2044] Output: Providing audio and visual information to the user
[2045] As described above, a series of processes are carried out, from recognizing the voice input from the user, analyzing emotions based on that, and recommending and providing the most suitable menu.
[2046] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2047] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2048] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2049] [Fourth embodiment]
[2050] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2051] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2052] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2053] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2054] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2055] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2056] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2057] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2058] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2059] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2060] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2061] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2062] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2063] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system works by combining speech recognition, natural language processing, data acquisition, and visual display.
[2064] ---
[2065] System Configuration
[2066] 1. Speech Recognition and Natural Language Processing
[2067] 1. Users
[2068] The user speaks a voice command, such as "What is the current Apple stock price?"
[2069] 2. Terminal
[2070] The device captures the user's voice input with a microphone and transmits it to the server.
[2071] 3. Server
[2072] The server uses a speech recognition system to convert the voice input into text, then uses a natural language processing system to parse the text and identify intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[2073] 2. Obtaining market data
[2074] 1. Server
[2075] Based on the intent and entity, send a request to a market data service to retrieve specific stock information (e.g., "Apple" stock price).
[2076] 3. Visual display and audio response
[2077] 1. Server
[2078] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[2079] 2. Terminal
[2080] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[2081] 3. Server
[2082] The server generates a response message to provide the user with the stock price information by voice, and creates a voice response using a voice synthesis system.
[2083] 4. Terminal
[2084] The terminal plays the voice response sent from the server and provides the information to the user.
[2085] ---
[2086] Specific examples
[2087] 1. Obtaining stock price information
[2088] 1. Users
[2089] User: "What's Apple's stock price right now?"
[2090] 2. Terminal
[2091] The terminal captures the user's voice and transmits it to the server.
[2092] 3. Server
[2093] The server converts the speech into text using a speech recognition system.
[2094] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[2095] Based on the intent, retrieve Apple's current stock price from a market data service.
[2096] 4. Terminal
[2097] Receive market data sent from the server and display it visually in the data visualizer.
[2098] 5. Server
[2099] Generate a voice response message saying "Apple's current stock price is $145" and create the voice response using a speech synthesis system.
[2100] 6. Terminal
[2101] A voice response is played to provide information to the user.
[2102] 2. Providing investment advice
[2103] 1. Users
[2104] User: "What investment advice would you give me?"
[2105] 2. Terminal
[2106] The device captures the audio and sends it to the server.
[2107] 3. Server
[2108] The server converts the speech to text and uses a natural language processing system to identify the intent "provide investment advice."
[2109] Analyze market data and generate investment advice based on current market trends.
[2110] Generate a voice response message that reads, "Based on current market trends, we recommend investing in technology stocks," and use a speech synthesis system to create the voice response.
[2111] 4. Terminal
[2112] A voice response is played to provide investment advice to the user.
[2113] As described above, the present invention provides a system that combines speech recognition, natural language processing, market data acquisition and visual display to enable users to efficiently acquire and understand investment information.
[2114] The processing flow will be explained below.
[2115] When obtaining stock price information
[2116] Step 1:
[2117] The user issues a voice command such as, "What is the current Apple stock price?"
[2118] Step 2:
[2119] The device captures the user's voice with a microphone and converts it into a digital signal.
[2120] Step 3:
[2121] The device transmits the captured audio data to the server.
[2122] Step 4:
[2123] The server analyzes the received voice data using a voice recognition system and converts it into text.
[2124] python
[2125] text_query = self.voice_recognition_system.transcribe(audio_input)
[2126] Step 5:
[2127] The server analyzes the text using a natural language processing system to identify intent and entities.
[2128] python
[2129] intent, entities = self.natural_language_processing.analyze_query(text_query)
[2130] Step 6:
[2131] Based on the analysis results, the server determines that the user's intent is to "get stock price information." At the same time, it extracts the entity "Apple."
[2132] Step 7:
[2133] The server sends a request to a market data service to get current stock price information for Apple.
[2134] python
[2135] market_data = self.market_data_service.get_stock_data(stock_ticker)
[2136] Step 8:
[2137] The server analyzes the acquired market data and sends it to the terminal.
[2138] Step 9:
[2139] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[2140] python
[2141] self.data_visualizer.visualize_data(market_data)
[2142] Step 10:
[2143] The server generates a response message to provide to the user: "Apple's current stock price is $145."
[2144] python
[2145] self.speak_response(f"Here is the market data for {stock_ticker}")
[2146] Step 11:
[2147] The server converts the generated message into a voice response using a speech synthesis system.
[2148] Step 12:
[2149] The terminal plays the voice response sent from the server and provides the information to the user.
[2150] python
[2151] play_audio_response(message)
[2152] ---
[2153] When providing investment advice
[2154] Step 1:
[2155] The user issues the voice command "Give me some investment advice."
[2156] Step 2:
[2157] The device captures the user's voice with a microphone and converts it into a digital signal.
[2158] Step 3:
[2159] The device transmits the captured audio data to the server.
[2160] Step 4:
[2161] The server analyzes the received voice data using a voice recognition system and converts it into text.
[2162] python
[2163] text_query = self.voice_recognition_system.transcribe(audio_input)
[2164] Step 5:
[2165] The server analyzes the text using a natural language processing system to identify the intent.
[2166] python
[2167] intent, entities = self.natural_language_processing.analyze_query(text_query)
[2168] Step 6:
[2169] Based on the analysis results, the server determines that the user's intent is to "provide investment advice."
[2170] Step 7:
[2171] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[2172] python
[2173] analysis = self.market_data_service.analyze_market()
[2174] Step 8:
[2175] The server generates a response message to provide to the user, which might say, "Based on current market trends, we recommend investing in technology stocks."
[2176] python
[2177] self.speak_response(f"My investment advice based on current market trends is: {analysis}")
[2178] Step 9:
[2179] The server converts the generated message into a voice response using a speech synthesis system.
[2180] Step 10:
[2181] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[2182] python
[2183] play_audio_response(message)
[2184] Example 1
[2185] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2186] Conventional investment information acquisition systems make it difficult for users to accurately and quickly acquire and understand market data in real time. Furthermore, since an intuitive method for acquiring information through voice commands was uncommon, improving the user experience was a challenge. Furthermore, the functionality for visually representing acquired data and providing voice responses was insufficient, making it difficult to meet the diverse needs of users.
[2187] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2188] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for outputting a response generated based on the market data by voice, means for using an existing voice recognition system and a natural language processing model for performing voice recognition and natural language processing, means for sending a request to an external data providing service to acquire market data and deserializing the result, means for using a data visualization tool to visually display the acquired market data in the form of graphs and charts, and means for using a voice synthesis system to generate a voice response, thereby enabling a user to intuitively acquire real-time market data through voice commands and quickly understand information visually and audibly.
[2189] "Means for recognizing voice input from a user" refers to a function that uses a voice input device such as a microphone to capture voice commands spoken by a user and obtain them in digital form.
[2190] "Means for converting said voice input to text" refers to a function that uses a voice recognition system to convert captured voice data into text form.
[2191] "Means for analyzing intent and entities from the text" refers to a function that uses a natural language processing model to identify objects (entities) related to a user's intent from text data.
[2192] "Means for obtaining market data based on the analyzed intent" refers to the function of sending a request derived from the analysis result to a market data providing service and obtaining the required data.
[2193] "Means for visually displaying the acquired market data" refers to the ability to use data visualization tools to display the acquired market data to a user in a visual format, such as a graph or chart.
[2194] "Means for outputting a response generated based on the market data in voice" refers to a function for generating a response message based on the acquired market data and providing it to the user in voice format using a voice synthesis system.
[2195] "Means of using existing speech recognition systems and natural language processing models for speech recognition and natural language processing" refers to the ability to apply existing speech recognition and natural language processing technologies to efficiently convert speech data into text and analyze the text.
[2196] "Means for sending requests to external data providers to retrieve market data and deserializing the results" means the functionality for sending requests to external market data providers to retrieve data and parsing the retrieved data into an appropriate format.
[2197] "Means of using data visualization tools to visually display acquired market data in the form of graphs and charts" refers to the ability to display data in a visually easy-to-understand format using dedicated visualization software or libraries.
[2198] "Means for using a speech synthesis system to generate a voice response" refers to the ability to use speech synthesis technology to convert a text message into speech form and provide it to a user.
[2199] The present invention is a system that allows users to receive real-time market analysis, stock information, and investment advice through voice commands. The system combines speech recognition, natural language processing, data acquisition, and visual display. Specifically, the system integrates a speech recognition system, a natural language processing model, a market data provider, a data visualization tool, and a speech synthesis system.
[2200] System Configuration
[2201] 1. Speech Recognition and Natural Language Processing
[2202] 1. Users
[2203] The user speaks a voice command such as "What is the current Apple stock price?"
[2204] 2. Terminal
[2205] The device captures the user's voice input using the device's built-in microphone, which may include a smartphone or smart speaker.
[2206] The device sends the captured audio data to the server using an HTTP request.
[2207] 3. Server
[2208] The server uses a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text.
[2209] The server uses a natural language processing model (e.g., OpenAI GPT-3) to identify intents (user intentions) and entities (specific targets) from the text data.
[2210] 2. Obtaining market data
[2211] 1. Server
[2212] The server sends a request based on the intent and entity to an external market data provider (e.g., Yahoo Finance API) using the HTTP GET method.
[2213] The server receives the data returned by the API and deserializes it in JSON format, which contains the latest market information for a specific financial instrument.
[2214] 3. Visual display and audio response
[2215] 1. Server
[2216] The server analyzes the acquired market data, formats it into a format that is easy for users to understand, and sends the formatted data to the terminal.
[2217] 2. Terminal
[2218] The terminal receives the data sent from the server and uses data visualization tools (e.g., Plotly or Matplotlib) to generate and visually display graphs and charts, including line graphs and bar graphs.
[2219] 3. Server
[2220] The server generates a response message such as "Apple's current stock price is $145" and creates the audio data using a speech synthesis system (e.g., Amazon Polly).
[2221] 4. Terminal
[2222] The terminal plays the audio data sent from the server and provides the information to the user, without the user having to perform any specific operation.
[2223] Specific examples
[2224] Get stock quotes
[2225] 1. Users
[2226] User: "What's Apple's stock price right now?"
[2227] 2. Terminal
[2228] The terminal captures the user's voice and transmits the voice data to the server.
[2229] 3. Server
[2230] The server converts the speech to text using the Google Cloud Speech-to-Text API, then uses OpenAI's GPT-3 to identify the intent "get stock quotes" and the entity "Apple" from the text.
[2231] The server sends a request to the Yahoo Finance API to get Apple's current stock price.
[2232] 4. Terminal
[2233] The terminal receives market data sent from the server and displays it as a line graph using Plotly.
[2234] 5. Server
[2235] The server generates a text message saying, "Apple's current stock price is $145," and creates voice data using Amazon Polly.
[2236] 6. Terminal
[2237] The terminal plays back the audio data and provides information to the user.
[2238] This invention allows users to intuitively obtain real-time market data through voice commands, providing information quickly and easily understood both visually and audibly. As described above, the present invention combines voice recognition, natural language processing, market data acquisition and visual display to provide a system that allows users to efficiently obtain and understand investment information.
[2239] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2240] Step 1: Capture and send audio input
[2241] 1. Users
[2242] The user speaks a voice command such as "What is the current Apple stock price?"
[2243] Input: User's voice command
[2244] Output: Audio data
[2245] 2. Terminal
[2246] The terminal uses the device's built-in microphone to capture the user's voice.
[2247] The captured audio is converted into a digital format and data compressed.
[2248] The terminal transmits the compressed audio data to the server as an HTTP request.
[2249] Input: User's voice data
[2250] Output: HTTP request to the server
[2251] Step 2: Convert audio data to text
[2252] 1. Server
[2253] The server sends the received audio data to the Google Cloud Speech-to-Text API.
[2254] The Google Cloud Speech-to-Text API converts the audio data into text format and returns it to the server.
[2255] Input: Compressed audio data
[2256] Output: Text data
[2257] Step 3: Identifying intents and entities through natural language processing
[2258] 1. Server
[2259] The server inputs the acquired text data into the OpenAI GPT-3 model to analyze intent and entities.
[2260] GPT-3 parses text and extracts intents (e.g., "get stock quotes") and entities (e.g., "Apple").
[2261] Input: Text data
[2262] Output: Intents and entities
[2263] Step 4: Obtain market data
[2264] 1. Server
[2265] The server sends an HTTP GET request to the Yahoo Finance API to retrieve market data for a particular financial instrument.
[2266] The Yahoo Finance API returns relevant market data in JSON format upon request.
[2267] The server deserializes the received JSON data and extracts the necessary information.
[2268] Input: Intents and Entities
[2269] Output: Market data
[2270] Step 5: Format and transmit market data
[2271] 1. Server
[2272] The server formats the acquired market data into a format that is easy for users to understand. The formatted data is organized, for example, as time-series data of stock prices.
[2273] The server transmits the formatted market data to the terminal.
[2274] Input: Market Data
[2275] Output: Formatted market data
[2276] Step 6: Generate a visual representation
[2277] 1. Terminal
[2278] The terminal uses the received market data to generate graphs and charts using data visualization tools such as Plotly and Matplotlib.
[2279] The generated visual representation is displayed on the screen of the terminal.
[2280] Input: Formatted market data
[2281] Output: Graphs and charts
[2282] Step 7: Generate and play a voice response
[2283] 1. Server
[2284] The server generates a response message such as "Apple's current stock price is $145."
[2285] The response message is sent to Amazon Polly and converted into voice data.
[2286] The generated voice data is transmitted to the terminal.
[2287] Input: Response message
[2288] Output: Audio data
[2289] 2. Terminal
[2290] The terminal plays back the received audio data and provides the information to the user.
[2291] Input: Audio data
[2292] Output: Audio information
[2293] This specific processing step allows users to obtain real-time market analysis and stock information through voice commands, and to understand it visually and audibly. The system provides an intuitive and easy-to-use interface for users.
[2294] (Application example 1)
[2295] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2296] In modern society, there is a demand for quick and efficient understanding of security risks and market trends. It is also important to enable users to easily access the information they need in real time by displaying that information visually and outputting it audibly. The present invention aims to provide a system that supports the acquisition and understanding of such information, thereby enabling users to quickly grasp risk management and market trends.
[2297] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2298] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring market data based on the analyzed intent, means for visually displaying the acquired market data, means for audio outputting a response generated based on the market data, means for acquiring data for analyzing crisis management and market trends in real time, means for visually displaying crisis management information from the data, and means for audio outputting a response generated based on the crisis management information, thereby enabling a user to acquire market analysis and security-related information in real time through voice commands and respond quickly.
[2299] "Means for recognizing voice input from a user" refers to technology that allows a device to receive speech uttered by a user and process it as a digital signal.
[2300] The "means for converting said voice input into text" refers to a technology for converting a voice signal into a string of characters, typically using a voice recognition algorithm.
[2301] The "means for analyzing intent and entities from the text" is a natural language processing technology for extracting a user's intent (intent) and specific target elements (entities) from text data.
[2302] The "means for acquiring market data based on the analyzed intent" is a technology for acquiring market data corresponding to the intent from an external database or API.
[2303] The "means for visually displaying acquired market data" refers to technology for displaying market data on a screen in the form of graphs, charts, etc.
[2304] The "means for outputting a response generated based on the market data as voice" is a technology that uses voice synthesis technology to provide a text response created based on market data to a user.
[2305] "Data acquisition means for analyzing crisis management and market trends in real time" refers to technology for collecting and analyzing current crisis management information and market trends in real time.
[2306] The "means for visually displaying crisis management information from the data" refers to a technology that visually displays collected and analyzed crisis management information in an easy-to-understand manner for users.
[2307] The "means for outputting a response generated based on the crisis management information as voice" is a technology that uses a voice synthesis technology to provide a text response generated based on the crisis management information to a user.
[2308] System Configuration
[2309] 1. Speech Recognition and Natural Language Processing
[2310] 1. Users
[2311] The user uses their smartphone or smart glasses to issue a voice command such as "Tell me the current crisis management information."
[2312] 2. Terminal
[2313] The device (smartphone or smart glasses) captures the user's voice input with a microphone and sends it to a server.
[2314] 3. Server
[2315] The server converts the voice input into text using the Google Speech-to-Text API, then uses a natural language processing system such as spaCy to parse the text and identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends").
[2316] 2. Obtaining market data and crisis management information
[2317] 1. Server
[2318] Use the News API and custom web scrapers to gather up-to-date crisis and market data based on intent and entities.
[2319] 3. Visual display and audio response
[2320] 1. Server
[2321] The acquired information is analyzed and presented in a visually easy-to-understand format using Plotly, and a voice response is generated using the Google Text-to-Speech API, stating, "The current trend in crisis management is an increase in phishing attacks."
[2322] 2. Terminal
[2323] The terminal displays the visual data sent from the server on its screen and plays back audio responses to provide information to the user.
[2324] Specific examples
[2325] 1. Speech Recognition and Natural Language Processing
[2326] A user utters, "What are the current trends in crisis management?"
[2327] The terminal captures the user's voice and sends it to the server.
[2328] The server converts the speech into text, and a natural language processing system identifies "obtaining crisis management information" and "current trends."
[2329] 2. Data Acquisition
[2330] The server collects crisis management information using the News API and custom web scrapers.
[2331] 3. Visual display and audio response
[2332] The server parses the information, formats it for visual display using Plotly, and generates an audio response using the Google Text-to-Speech API.
[2333] The terminal displays the displayed data on a screen and plays audio responses to provide information to the user.
[2334] Prompt Sentence Examples
[2335] Below are some example prompts to be fed into the generative AI model:
[2336] When a user says, "What are the current crisis management trends?" the application performs the following steps:
[2337] 1. Use the Google Speech-to-Text API to convert voice commands into text.
[2338] 2. Using spaCy, we analyzed the intent “Get crisis management information” and the entity “Current trends” from the text.
[2339] 3. Use the News API or a custom web scraper to gather the latest crisis management information.
[2340] 4. Use Plotly to generate visually easy-to-understand graphs and charts and display them on the user screen.
[2341] 5. Using the Google Text-to-Speech API, a voice response was generated that said, "The current trend in crisis management is an increase in phishing attacks," and played on a smartphone or smart glasses.
[2342] This allows users to easily obtain and understand real-time security information through voice commands.
[2343] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2344] Step 1:
[2345] The user says, "Tell me about the current trends in crisis management." The input is the user's voice command, and the output is voice data. This allows the system to recognize the user's request and begin processing.
[2346] Step 2:
[2347] The device captures the user's voice input with a microphone and sends it to the server. The input is voice data, and the output is a digital signal. The voice data is transmitted to the server via the Internet.
[2348] Step 3:
[2349] The server converts the voice input into text using the Google Speech-to-Text API. The input is digital audio data, and the output is text data. This process recognizes the voice as text.
[2350] Step 4:
[2351] The server uses a natural language processing system such as spaCy to identify intents (e.g., "Get crisis management information") and entities (e.g., "Current trends") from the text data. The input is text data, and the output is the parsed intent and entities. This clarifies the user's intent and specific request.
[2352] Step 5:
[2353] The server collects the latest crisis management information using the News API or a custom web scraper based on the intent and entity. The input is the intent and entity, and the output is the crisis management information data. This allows the appropriate information to be retrieved.
[2354] Step 6:
[2355] The server analyzes the acquired information and uses Plotly to generate graphs and charts in a visually easy-to-understand format. The input is crisis management information data, and the output is visual display data. This converts the information into a format that is easy for users to understand.
[2356] Step 7:
[2357] The server uses the Google Text-to-Speech API to generate a voice response stating, "Current crisis management trends include an increase in phishing attacks." The input is the text response message, and the output is the audio data. This provides audio information along with visual information.
[2358] Step 8:
[2359] The terminal displays the visual data sent from the server on the screen and plays back audio responses to provide information to the user. The input is visual display data and audio data, and the output is the displayed data and played audio. This allows the user to obtain and understand the information they need in real time.
[2360] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2361] The present invention provides more personalized market data and investment advice by combining a system that recognizes user voice input, converts text, and processes natural language with an emotion engine. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[2362] ---
[2363] System Configuration
[2364] 1. Speech Recognition and Natural Language Processing
[2365] 1. Users
[2366] Users can speak voice commands such as "What is the current Apple stock price?" or "I'd like some investment advice."
[2367] 2. Terminal
[2368] The device captures the user's voice input with a microphone and transmits it to the server.
[2369] 3. Server
[2370] The server uses a speech recognition system to convert the voice input into text, and then uses a natural language processing system to parse the text and identify intents and entities.
[2371] 2. Emotion recognition
[2372] 1. Server
[2373] The server recognizes the user's emotions using an emotion engine based on the voice recognition data and text data.
[2374] 3. Obtaining market data
[2375] 1. Server
[2376] Based on the intent and entity, a request is sent to the market data service to retrieve specific stock information or market data.
[2377] 4. Regulating responses based on emotions
[2378] 1. Server
[2379] The emotion engine adjusts the tone and content of responses based on the user's emotions. For example, if it detects that the user is feeling stressed, it will generate a more reassuring response.
[2380] 5. Visual display and audio response
[2381] 1. Server
[2382] The acquired market data is analyzed and sent to the terminal in a visually easy-to-understand format.
[2383] 2. Terminal
[2384] The terminal uses a data visualizer to display market data in visual formats such as graphs and charts.
[2385] 3. Server
[2386] The server generates a response message to be presented to the user and uses a speech synthesis system to create a voice response.
[2387] 4. Terminal
[2388] The terminal plays the voice response sent from the server and provides the information to the user.
[2389] ---
[2390] Specific examples
[2391] 1. Obtaining stock price information
[2392] 1. Users
[2393] User: "What's Apple's stock price right now?"
[2394] 2. Terminal
[2395] The terminal captures the user's voice and transmits it to the server.
[2396] 3. Server
[2397] The server converts the speech into text using a speech recognition system.
[2398] A natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple."
[2399] Recognize user emotions with an emotion engine.
[2400] Based on the intent, retrieve Apple's current stock price from a market data service.
[2401] 4. Terminal
[2402] Receive market data sent from the server and display it visually in the data visualizer.
[2403] 5. Server
[2404] Generate a voice response message saying "Apple's current stock price is $145" and adjust the tone depending on the user's emotion.
[2405] Generate a reassuring response like, "Apple stock is currently at $145, don't worry, there are good investment opportunities waiting for you."
[2406] 6. Terminal
[2407] A voice response is played to provide information to the user.
[2408] 2. Providing investment advice
[2409] 1. Users
[2410] User: "What investment advice would you give me?"
[2411] 2. Terminal
[2412] The device captures the audio and sends it to the server.
[2413] 3. Server
[2414] The server converts the speech to text and uses a natural language processing system to identify the intent: "Provide investment advice."
[2415] Recognize user emotions with an emotion engine.
[2416] Comprehensively analyze market data and generate investment advice based on current market trends.
[2417] 4. Server
[2418] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions.
[2419] "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and a good long-term view," he suggests.
[2420] 5. Terminal
[2421] A voice response is played to provide investment advice to the user.
[2422] As described above, the present invention provides a system that enables users to efficiently acquire and understand personalized investment information by integrating and operating speech recognition, natural language processing, market data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[2423] The processing flow will be explained below.
[2424] When obtaining stock price information
[2425] Step 1:
[2426] The user issues a voice command such as, "What is the current Apple stock price?"
[2427] Step 2:
[2428] The device captures the user's voice with a microphone and converts it into a digital signal.
[2429] Step 3:
[2430] The device transmits the captured audio data to the server.
[2431] Step 4:
[2432] The server analyzes the received voice data using a voice recognition system and converts it into text.
[2433] python
[2434] text_query = self.voice_recognition_system.transcribe(audio_input)
[2435] Step 5:
[2436] The server analyzes the text using a natural language processing system to identify intent and entities.
[2437] python
[2438] intent, entities = self.natural_language_processing.analyze_query(text_query)
[2439] Step 6:
[2440] The server verifies that the intent parsed from the text is "Get stock quotes" and identifies "Apple" as the entity.
[2441] Step 7:
[2442] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[2443] python
[2444] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[2445] Step 8:
[2446] The server takes into account the perceived sentiment and sends a request to a market data service to get current stock price information for Apple.
[2447] python
[2448] market_data = self.market_data_service.get_stock_data(stock_ticker)
[2449] Step 9:
[2450] The server analyzes the acquired market data and sends it to the terminal.
[2451] Step 10:
[2452] The terminal uses a data visualizer to visually display the received market data in graphs and charts.
[2453] python
[2454] self.data_visualizer.visualize_data(market_data)
[2455] Step 11:
[2456] The server generates a response message to be delivered to the user, such as "Apple's current stock price is $145," and adjusts the tone according to the user's emotion as determined by the emotion engine.
[2457] python
[2458] response_message = f"Apple's current stock price is ${market_data['price']}."
[2459] response_message = self.emotion_engine.adjust_tone(response_message, user_emotion)
[2460] Step 12:
[2461] The server converts the generated message into a voice response using a speech synthesis system.
[2462] Step 13:
[2463] The terminal plays the voice response sent from the server and provides the information to the user.
[2464] python
[2465] play_audio_response(response_message)
[2466] ---
[2467] When providing investment advice
[2468] Step 1:
[2469] The user issues the voice command "Give me some investment advice."
[2470] Step 2:
[2471] The device captures the user's voice with a microphone and converts it into a digital signal.
[2472] Step 3:
[2473] The device transmits the captured audio data to the server.
[2474] Step 4:
[2475] The server analyzes the received voice data using a voice recognition system and converts it into text.
[2476] python
[2477] text_query = self.voice_recognition_system.transcribe(audio_input)
[2478] Step 5:
[2479] The server analyzes the text using a natural language processing system to identify the intent.
[2480] python
[2481] intent, entities = self.natural_language_processing.analyze_query(text_query)
[2482] Step 6:
[2483] The server determines that the intent parsed from the text is "provide investment advice."
[2484] Step 7:
[2485] The server uses an emotion engine to recognize the user's emotions based on the voice data and text data.
[2486] python
[2487] user_emotion = self.emotion_engine.analyze(audio_input, text_query)
[2488] Step 8:
[2489] The server performs comprehensive analysis of market data and generates investment advice based on current market trends.
[2490] python
[2491] analysis = self.market_data_service.analyze_market()
[2492] investment_advice = f"My investment advice based on current market trends is: {analysis}"
[2493] Step 9:
[2494] The server adjusts the content and tone of the investment advice based on the perceived user sentiment.
[2495] python
[2496] investment_advice = self.emotion_engine.adjust_tone(investment_advice, user_emotion)
[2497] Step 10:
[2498] The server converts the generated investment advice message into a voice response using a voice synthesis system.
[2499] Step 11:
[2500] The terminal plays back the voice response sent from the server and provides investment advice to the user.
[2501] python
[2502] play_audio_response(investment_advice)
[2503] Through the above steps, the present invention combines voice input, natural language processing, emotion recognition, data acquisition, and visual display to realize a system that provides users with personalized investment information, allowing them to acquire and understand investment information more quickly and efficiently.
[2504] Example 2
[2505] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2506] Conventional systems simply convert voice input from users into text and provide market data based on specific intents and entities, but are unable to generate personalized responses that take into account the user's emotions. This has led to the challenge of not being able to provide responses with the appropriate tone and content according to the user's emotions.
[2507] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2508] In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for recognizing user emotions from the voice input and text data, means for acquiring market data based on the analyzed intent and entities, means for visually displaying the acquired market data, means for adjusting the tone and content of a response generated based on the user emotions, and means for outputting the market data and the adjusted response by voice, thereby enabling a user to efficiently acquire and understand personalized investment information.
[2509] A "user" is a human user who requests information by speaking voice commands to the system.
[2510] "Voice input" is voice data uttered by the user.
[2511] "Speech recognition" is a technology that analyzes a user's voice input and converts the content into text data.
[2512] "Text conversion" is the process of converting audio data into text data.
[2513] "Intent" refers to identifying the user's intention or purpose for text input.
[2514] An "entity" is an element that extracts and identifies specific keywords or items within a text.
[2515] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state from voice input or text data.
[2516] "Market data" refers to financial-related data such as specific stock information and economic indicators.
[2517] "Visual display" refers to converting acquired market data into a visually understandable format, such as a graph or chart, and displaying it.
[2518] "Response generation" is the process of creating a response message to be provided to a user based on acquired market data and user sentiment.
[2519] "Audio output" is the process of playing back the generated response message as audio.
[2520] This invention relates to a system that recognizes user voice input, performs text conversion and natural language processing, and combines it with an emotion engine to provide personalized market data and investment advice. The system integrates speech recognition, natural language processing, data acquisition, visual display, voice output, and an emotion engine that recognizes user emotions.
[2521] System configuration
[2522] The system includes the following main components:
[2523] 1. Acquiring voice input
[2524] User
[2525] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[2526] Terminal
[2527] The device will use a built-in microphone to capture the user's voice input, and at this stage it is advisable to use a high-quality microphone and noise-cancelling technology.
[2528] 2. Sending audio data to the server
[2529] Terminal
[2530] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[2531] 3. Speech to text conversion
[2532] server
[2533] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[2534] 4. Analysis using natural language processing
[2535] server
[2536] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[2537] 5. Emotion recognition
[2538] server
[2539] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[2540] 6. Obtaining Market Data
[2541] server
[2542] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[2543] 7. Response generation and coordination
[2544] server
[2545] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[2546] 8. Visual Indications
[2547] server
[2548] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[2549] Terminal
[2550] The terminal displays the visual data sent from the server on a browser or dedicated application.
[2551] 9. Providing voice responses
[2552] Terminal
[2553] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[2554] Specific examples
[2555] Get stock quotes
[2556] 1. Users
[2557] User: "What's Apple's stock price right now?"
[2558] 2. Terminal
[2559] The terminal captures the user's voice and transmits it to the server.
[2560] 3. Server
[2561] The server converts the speech to text using a speech recognition system. The natural language processing system analyzes the text and identifies the intent "get stock quotes" and the entity "Apple." The emotion engine recognizes the user's emotions. Based on the intent, the server retrieves Apple's current stock price from a market data service.
[2562] 4. Terminal
[2563] Receive market data sent from the server and display it visually using D3.js, Chart.js, etc.
[2564] 5. Server
[2565] Generate a voice response message such as "Apple's current stock price is $145" and adjust the tone depending on the user's emotions. Generate a reassuring response such as "Apple's current stock price is $145, don't worry, there are good investment opportunities waiting for you."
[2566] 6. Terminal
[2567] A voice response is played to provide information to the user.
[2568] Providing investment advice
[2569] 1. Users
[2570] User: "What investment advice would you give me?"
[2571] 2. Terminal
[2572] The device captures the audio and sends it to the server.
[2573] 3. Server
[2574] The server converts the speech into text and uses a natural language processing system to identify the intent "provide investment advice." The emotion engine recognizes the user's emotions. The server then comprehensively analyzes market data and generates investment advice based on current market trends.
[2575] 4. Server
[2576] It generates a voice response message such as, "Based on current market trends, we recommend investing in technology stocks," and adjusts the tone and content depending on the user's emotions. It also suggests, "If you're worried about taking a little risk, consider technology stocks that offer peace of mind and are good for the long term."
[2577] 5. Terminal
[2578] A voice response is played to provide investment advice to the user.
[2579] Prompt Sentence Examples
[2580] User says: "What is Apple's stock price right now?"
[2581] Prompt: "What is the process flow if the user says, 'What is Apple's stock price now?'"
[2582] In this way, the system integrates multiple technologies, including user emotion recognition, to provide a way for users to efficiently obtain and understand personalized investment information.
[2583] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2584] Step 1:
[2585] Acquiring voice input
[2586] User
[2587] Users can speak voice commands such as "What is the current Apple stock price?" or "Give me some investment advice."
[2588] Terminal
[2589] The device uses a built-in microphone to capture the user's voice input, utilizing high-quality microphone and noise-canceling technology to collect voice data in real time.
[2590] Input: User's voice command
[2591] Output: Captured audio data
[2592] Step 2:
[2593] Sending audio data to the server
[2594] Terminal
[2595] The device sends the captured audio data to the server via an HTTP request, which encodes and encrypts the data to ensure security.
[2596] Input: Captured audio data
[2597] Output: Audio data sent to the server
[2598] Step 3:
[2599] Speech-to-text conversion
[2600] server
[2601] The server uses a speech recognition system to convert the received voice data into text data, using speech recognition technology such as the Google Cloud Speech-to-Text API.
[2602] Input: Transmitted audio data
[2603] Output: Converted text data
[2604] Step 4:
[2605] Analysis using natural language processing
[2606] server
[2607] The server uses a natural language processing system to analyze the text data, specifically using technologies like OpenAI's GPT-3 to identify intent and entities.
[2608] Input: Text data
[2609] Output: Identified intents and entities
[2610] Step 5:
[2611] emotion recognition
[2612] server
[2613] The server uses an emotion engine to recognize the user's emotions based on the voice recognition data and text data. For this, Affectiva's SDK can be used.
[2614] Input: Audio and text data
[2615] Output: Recognized user emotion
[2616] Step 6:
[2617] Obtaining Market Data
[2618] server
[2619] The server sends a request to a service that provides specific stock information or market data (e.g., the Yahoo Finance API) and retrieves data based on the intent and entities.
[2620] Input: Identified intents and entities
[2621] Output: Market data
[2622] Step 7:
[2623] Response generation and coordination
[2624] server
[2625] The server generates an appropriate response message based on the acquired data and the user's emotions, using a text-to-speech system such as Amazon Polly to convert the text into speech.
[2626] Input: Market data and user sentiment
[2627] Output: Adjusted response message
[2628] Step 8:
[2629] Visual Indication
[2630] server
[2631] The server uses libraries such as D3.js and Chart.js to convert the acquired market data into a format that is easy to understand visually.
[2632] Input: Market Data
[2633] Output: Visually displayable data
[2634] Terminal
[2635] The terminal displays the visual data sent from the server on a browser or dedicated application.
[2636] Input: Visually displayable data
[2637] Output: Graphs and charts that are displayed to the user
[2638] Step 9:
[2639] Providing voice responses
[2640] Terminal
[2641] The device plays back the voice response sent from the server, using a high-quality speaker to provide information to the user.
[2642] Input: Tailored response message
[2643] Output: The audio response the user hears
[2644] In this way, each processing step works in conjunction with the others, allowing the user to efficiently obtain and understand personalized investment information.
[2645] (Application example 2)
[2646] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2647] Conventional food delivery services have the problem that they do not suggest personalized menus based on the user's emotions, making it difficult to select the optimal menu based on the user's psychological state. Also, because they do not suggest optimal dishes based on the user's emotions, the improvement in satisfaction is limited.
[2648] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a voice input from a user, means for converting the voice input into text, means for analyzing intent and entities from the text, means for acquiring menu information based on the analyzed intent, means for analyzing the user's emotion based on the intent and entities, means for adjusting the recommended menu based on the emotion information, means for visually displaying the acquired menu information, and means for outputting a response generated based on the menu information by voice. This makes it possible to propose an optimal menu based on the user's emotion, thereby significantly improving user satisfaction.
[2649] A "user" is someone who uses the system.
[2650] "Voice input" is data of the voice that the user utters to the system.
[2651] "Text" refers to speech input converted into character data.
[2652] An "intent" is the purpose or request that a user intends to make to a system.
[2653] "Entity" means a concrete object or subject that is relevant to an intent.
[2654] "Menu information" is data related to the provision of food and drink, and is a list of dishes and drinks suggested to the user.
[2655] "Emotion" refers to the user's mental or emotional state.
[2656] A "recommended menu" is a selection of food and drink options suggested based on the user's intent and emotions.
[2657] The "visual display means" is a method for presenting the acquired menu information and recommended menus to the user in a graphical format.
[2658] "Audio output means" refers to a method for transmitting acquired information or generated responses to the user in audio form.
[2659] The term "system" refers to a series of functional blocks that operate in combination with the above means.
[2660] This invention relates to a food delivery system that recognizes and analyzes voice input from a user to suggest menu items based on the user's emotions. Below, we will explain the program and its detailed processing for specifically implementing the invention, the hardware and software used, and specific examples of use.
[2661] Hardware and Software Used
[2662] Hardware: smartphone, microphone, server
[2663] software:
[2664] speech_recognition library: for speech recognition
[2665] textblob library: for natural language processing
[2666] sentiment_analysis module: for emotion recognition
[2667] delivery_service_api: To obtain and recommend menu information
[2668] Explanation of the program processing flow
[2669] 1. Voice Recognition
[2670] When a user speaks, the microphone connected to the smartphone captures the voice data. The speech_recognition library is used to convert this voice data into text. An example of voice input might be, "I'm tired today, so I want some food to relax me."
[2671] 2. Natural Language Processing
[2672] The converted text data is then analyzed in detail using the textblob library, where the user's intent and the desired object (entity) are extracted. For example, from this utterance, the intent "I want to relax" is identified, and the entity "food" is identified.
[2673] 3. Emotion recognition
[2674] The analyzed text data is then subjected to emotion recognition by the sentiment_analysis module, which determines the user's emotional state as "fatigue" and selects the most appropriate menu based on this.
[2675] 4. Menu information acquisition and recommendation
[2676] Based on the user's sentiment and intent, the delivery_service_api is used to obtain the most suitable menu information. For example, it may recommend herbal tea or Japanese soup as a relaxing food.
[2677] 5. Visual and audio output
[2678] The acquired menu information is sent from the server to the smartphone and displayed visually on the smartphone screen. This information is also output as a voice response. For example, a voice response such as "Today, we recommend relaxing herbal tea or Japanese-style soup."
[2679] Specific examples
[2680] Usage example 1
[2681] User: "I'm tired today and I want some food to help me relax."
[2682] Server: Converts speech to text, identifies intent "relaxation", and determines emotional state as "fatigue".
[2683] System: Retrieves recommended menu items such as herbal tea and Japanese-style soup, and presents them to the user visually and audibly.
[2684] Prompt Sentence Examples
[2685] "User: "I'm tired today and I want some food to help me relax."
[2686] Server: "Thank you for your hard work. To help you relax, we recommend the following: herbal tea, Japanese soup. Please let us know if there's anything else you'd like."
[2687] In this way, the present invention realizes a system that is sensitive to the user's emotions and provides optimal menu suggestions and food delivery services.
[2688] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2689] Step 1:
[2690] The user opens the smartphone application and speaks into the microphone, for example, saying, "I'm tired today, so I want some food to help me relax." The voice recognition module in the smartphone captures this voice data.
[2691] Input: User voice input
[2692] Output: Captured audio data
[2693] Step 2:
[2694] The device converts the captured voice data into text data using the speech_recognition library. The user's statement, "I'm tired today, so I want some food to relax me," is converted into text.
[2695] Input: Captured audio data
[2696] Output: Converted text data
[2697] Step 3:
[2698] The server receives the converted text data and parses it using the textblob library to extract the intent "I want to relax" and the entity "food."
[2699] Input: Converted text data
[2700] Output: Intents and entities
[2701] Step 4:
[2702] The server analyzes the user's emotion using the sentiment_analysis module based on the intent and entities. In this case, the user's emotion is determined to be "fatigue."
[2703] Input: Intents and Entities
[2704] Output: User's emotional information
[2705] Step 5:
[2706] The server uses the preprocessed intent and emotion information to retrieve appropriate menu information via the delivery_service_api. For example, herbal tea or Japanese soup is selected as relaxing food.
[2707] Input: Intent and emotion information
[2708] Output: Recommended menu information
[2709] Step 6:
[2710] The server converts the recommended menu information into a visually displayable format and sends it to the terminal, which then uses its visualization engine to display the received data on the screen.
[2711] Input: Recommended menu information
[2712] Output: Visualized menu information
[2713] Step 7:
[2714] The server generates a voice response message to convey to the user and converts it into voice data using a speech synthesis system. For example, a message such as "Today, we recommend a relaxing herbal tea or Japanese-style soup."
[2715] Input: Recommended menu information
[2716] Output: Voice response message
[2717] Step 8:
[2718] The terminal plays the voice response message sent from the server to provide information to the user, who can then check the details of the relaxing menu through the visual menu display and voice response.
[2719] Input: Voice response message
[2720] Output: Providing audio and visual information to the user
[2721] As described above, a series of processes are carried out, from recognizing the voice input from the user, analyzing emotions based on that, and recommending and providing the most suitable menu.
[2722] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2723] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2724] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2725] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2726] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2727] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2728] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2729] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2730] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2731] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2732] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2733] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2734] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2735] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2736] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2737] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2738] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2739] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2740] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2741] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2742] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2743] The following is further disclosed regarding the above embodiment.
[2744] (Claim 1)
[2745] means for recognizing speech input from a user;
[2746] means for converting said speech input into text;
[2747] means for parsing intent and entities from the text;
[2748] means for obtaining market data based on the parsed intent;
[2749] a means for visually displaying the obtained market data;
[2750] means for audibly outputting a response generated based on said market data;
[2751] A system including:
[2752] (Claim 2)
[2753] 10. The system of claim 1, further comprising means for processing multimodal input including images and video as part of said means.
[2754] (Claim 3)
[2755] 10. The system of claim 1, further comprising means for providing specific stock information based on entities parsed from the text.
[2756] "Example 1"
[2757] (Claim 1)
[2758] means for recognizing speech input from a user;
[2759] means for converting said speech input into text;
[2760] means for parsing intent and entities from the text;
[2761] means for obtaining market data based on the parsed intent;
[2762] a means for visually displaying the obtained market data;
[2763] means for audibly outputting a response generated based on said market data;
[2764] a means for using existing speech recognition systems and natural language processing models to perform speech recognition and natural language processing;
[2765] a means for sending requests to an external data provider to obtain market data and deserializing the results;
[2766] The means to use data visualization tools to visually display acquired market data in the form of graphs and charts;
[2767] means for using a speech synthesis system to generate a voice response;
[2768] A system including:
[2769] (Claim 2)
[2770] 10. The system of claim 1, further comprising means for processing multimodal input including images and video as part of said means.
[2771] (Claim 3)
[2772] 10. The system of claim 1, further comprising: means for providing information about a particular financial instrument based on entities parsed from the text.
[2773] "Application Example 1"
[2774] Rewriting claims to fit the new invention
[2775] (Claim 1)
[2776] means for recognizing speech input from a user;
[2777] means for converting said speech input into text;
[2778] means for parsing intent and entities from the text;
[2779] means for obtaining market data based on the parsed intent;
[2780] a means for visually displaying the obtained market data;
[2781] means for audibly outputting a response generated based on said market data;
[2782] Data acquisition methods for crisis management and real-time analysis of market trends,
[2783] means for visually displaying crisis management information from said data;
[2784] a means for outputting a response generated based on the crisis management information by voice;
[2785] A system including:
[2786] (Claim 2)
[2787] 10. The system of claim 1, further comprising means for processing multimodal input including images and video as part of said means.
[2788] (Claim 3)
[2789] 10. The system of claim 1, further comprising means for providing specific crisis management information based on entities parsed from the text.
[2790] "Example 2: Combining Emotion Engines"
[2791] (Claim 1)
[2792] means for recognizing speech input from a user;
[2793] means for converting said speech input into text;
[2794] means for parsing intent and entities from the text;
[2795] means for recognizing a user's emotion from the voice input and text data;
[2796] means for obtaining market data based on the parsed intent and entities;
[2797] a means for visually displaying the obtained market data;
[2798] means for adjusting the tone and content of the generated response based on the user's emotions;
[2799] means for audibly outputting said market data and adjusted responses;
[2800] A system including:
[2801] (Claim 2)
[2802] 10. The system of claim 1, further comprising means for processing multimodal input including images and video as part of said means.
[2803] (Claim 3)
[2804] 10. The system of claim 1, further comprising means for providing specific market information based on entities parsed from the text.
[2805] "Application example 2 when combining emotion engines"
[2806] (Claim 1)
[2807] means for recognizing speech input from a user;
[2808] means for converting said speech input into text;
[2809] means for parsing intent and entities from the text;
[2810] means for obtaining menu information based on the parsed intent;
[2811] means for analyzing a user's sentiment based on the intent and entities;
[2812] means for adjusting a recommendation menu based on the emotion information;
[2813] means for visually displaying the acquired menu information;
[2814] a means for outputting a response generated based on the menu information by voice;
[2815] A system including:
[2816] (Claim 2)
[2817] 10. The system of claim 1, further comprising means for processing multimodal input including images and video as part of said means.
[2818] (Claim 3)
[2819] 10. The system of claim 1, further comprising: means for providing specific cuisine information based on entities parsed from the text. [Explanation of symbols]
[2820] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for recognizing speech input from a user; means for converting said speech input into text; means for parsing intent and entities from the text; means for obtaining market data based on the parsed intent; a means for visually displaying the obtained market data; means for audibly outputting a response generated based on said market data; A system including:
2. The system of claim 1 further comprising means for processing multimodal input including images and video as part of said means.
3. The system of claim 1 further comprising means for providing specific stock information based on entities parsed from the text.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A