system
A voice-operated navigation system for motorcyclists converts voice input into text, analyzes intent, retrieves real-time data, and outputs audio guidance, addressing the challenge of safely and efficiently obtaining navigation information while riding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Motorcyclists face challenges in safely and efficiently obtaining real-time navigation information due to the need to manually operate navigation systems with limited hands-free options, compromising safety and convenience.
A voice-operated navigation system that converts voice input into text, analyzes user intent, retrieves real-time information, generates responses, and outputs audio guidance, allowing riders to access traffic, weather, and facility information without using their hands.
Enables safe and efficient travel by providing timely navigation information through voice control, improving safety and convenience for motorcyclists.
Smart Images

Figure 2026064602000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including: receiving a user utterance; adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When riding a motorcycle, the rider is concentrating both hands on handling the handlebars, making it difficult to manually operate the navigation system. In addition, since there are limited means to provide the rider with information that is updated in real time, such as traffic congestion, weather changes, and facility information near the destination, safe and comfortable travel is hindered. The purpose of this invention is to solve these problems by providing the rider with real-time information using voice operation and improving the efficiency of navigation.
Means for Solving the Problems
[0005] The present invention is a system that includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data and determining the user's intent, means for acquiring real-time information based on the analysis results, means for generating a response based on the acquired information, and means for converting the generated response into voice data and outputting it to the user. With this system, the rider can acquire real-time traffic information, weather information, facility information near the destination, etc., through voice operation without using their hands, and receive appropriate navigation. Furthermore, it is also possible to acquire real-time traffic information and suggest detour routes based on congestion information, and provide information on recommended spots around the destination. This enables safe and efficient travel.
[0006] "Voice input" is a method of inputting the user's voice into the system.
[0007] "Text data" refers to character information that is converted from speech input after analysis.
[0008] "Analysis" is the process of analyzing input text data to determine the user's intent.
[0009] "Real-time information" refers to the latest data on traffic, weather, facilities, and other relevant information as of the current time.
[0010] A "response" refers to the instructions or information provided by the system to the user based on the analysis results.
[0011] "Audio data" refers to data in which the generated response has been converted into an audio format.
[0012] "Output" refers to the process of playing audio data to the user and conveying information.
[0013] "Traffic information" refers to data related to traffic, such as road congestion and accident information.
[0014] The "detour route" is an alternative route proposed to avoid traffic jams and accidents.
[0015] The "recommended spot information" is information on facilities and places proposed based on the interests and needs of the user.
[0016] The "navigation system" is an electronic system that provides the user with guidance on the moving route and destination.
Brief Explanation of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Modes for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] This invention relates to a navigation system for motorcycle riders, enabling them to obtain real-time information through voice control without using their hands, and to travel safely and efficiently. The processing and operation of the system's program are described in detail below.
[0039] Overall flow and process overview
[0040] 1. Acquisition of voice input:
[0041] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[0042] Example: "Check if the current road is congested."
[0043] 2. Speech-to-text conversion:
[0044] The device converts the received voice input into text data using a speech recognition API.
[0045] 3. Information Analysis:
[0046] The server receives text data and performs natural language processing to analyze the user's intent.
[0047] Based on the analysis results, appropriate real-time information is obtained for keywords such as "traffic congestion."
[0048] 4. Obtaining real-time information:
[0049] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[0050] The acquired traffic and weather information is integrated with the analysis results.
[0051] 5. Generating the response:
[0052] Generative AI generates responses to the user based on analysis results and acquired real-time information.
[0053] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0054] 6. Conversion to and output of audio data:
[0055] The generative AI converts the generated response text into speech data using a speech synthesis API.
[0056] The device outputs audio data to the user.
[0057] Processing of specific examples
[0058] 1. Proposal of an alternative route
[0059] User: "Check if the current road is congested."
[0060] The terminal converts voice input into text data and sends it to the server.
[0061] The server analyzes the text data and calls a traffic information API to obtain real-time traffic information.
[0062] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[0063] The generative AI generates a response such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts it into audio data.
[0064] The device outputs audio data to the user.
[0065] 2. Providing information on recommended spots
[0066] User: "I'm hungry, can you recommend some good restaurants nearby?"
[0067] The terminal converts voice input into text data and sends it to the server.
[0068] The server analyzes the text data and calls a restaurant information API to retrieve the latest information.
[0069] The server sends the acquired restaurant information to a generative AI, which then generates a response suggesting the most suitable restaurant.
[0070] The generative AI generates the response, "There is a highly-rated Italian restaurant 1 kilometer away," and converts it into audio data.
[0071] The device outputs audio data to the user.
[0072] System Configuration
[0073] This system includes the following:
[0074] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[0075] A means of converting audio data into text data (speech recognition API).
[0076] A means of analyzing text data and determining the user's intent (natural language processing engine).
[0077] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[0078] A means of generating a response based on acquired information (generative AI).
[0079] A means of converting a response into audio data and outputting it (speech synthesis API).
[0080] These features allow users to travel safely and efficiently through voice control. This system is particularly effective when riding a motorcycle, as it can provide timely and appropriate information even when both hands are occupied.
[0081] The following describes the processing flow.
[0082] Step 1:
[0083] The user gives voice commands through a helmet with a microphone or a Bluetooth headset.
[0084] Example: "Check if the current road is congested."
[0085] Step 2:
[0086] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[0087] It calls a speech recognition API (for example, Google® Cloud Speech-to-Text API) to convert speech to text.
[0088] The converted data is temporarily stored on the device.
[0089] Step 3:
[0090] The terminal sends text data to the server.
[0091] Send an HTTP request to the server, including text data.
[0092] Step 4:
[0093] The server receives text data.
[0094] The natural language processing (NLP) engine is invoked to analyze the received text data.
[0095] Step 5:
[0096] The server analyzes the text data using a natural language processing engine to determine the user's intent.
[0097] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[0098] Step 6:
[0099] The server determines the need to acquire real-time information based on the analysis results.
[0100] The analysis results indicate a request to "check traffic congestion information."
[0101] Step 7:
[0102] The server calls a traffic information API to obtain real-time traffic information.
[0103] Example: Use the Google Maps API to retrieve current traffic information.
[0104] The acquired traffic information is temporarily stored.
[0105] Step 8:
[0106] The server integrates real-time traffic information with analysis results and sends it to the generative AI.
[0107] Based on the analysis results and traffic information, the data is sent to a generative AI engine.
[0108] Step 9:
[0109] The generative AI generates a response to the user.
[0110] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0111] Generate response text.
[0112] Step 10:
[0113] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[0114] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[0115] Step 11:
[0116] The device receives audio data and outputs it to the user.
[0117] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[0118] Step 12:
[0119] The user listens to voice guidance and follows the suggested detour route.
[0120] Choose a safe new route and head towards your destination.
[0121] (Example 1)
[0122] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0123] It is not easy for motorcyclists to safely and efficiently obtain real-time traffic information and recommended spots without using both hands while riding. Conventional navigation systems require screen operation, which can compromise safety while driving. Therefore, there is a need for a more intuitive and safe information acquisition system that uses voice control.
[0124] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0125] In this invention, the server includes means for the user to give instructions by voice, means for converting the given voice instructions into text format, means for analyzing the text and determining the user's intent, means for acquiring real-time information based on the determined intent, a generative model for generating a detailed response based on the real-time information, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire necessary information in real time through voice operation and to move safely and efficiently.
[0126] "Voice input" refers to the act of capturing the user's voice into the system via a microphone or headset.
[0127] "Means of converting to text data" refers to technology or equipment for converting voice input into strings of characters.
[0128] "Means for determining user intent" refers to technologies or devices that analyze the content of received text data to understand the content of user requests or questions.
[0129] "Means for obtaining real-time information" refers to technologies or devices that obtain the latest data, such as current traffic conditions and weather, from appropriate sources based on analysis results.
[0130] "Means for generating responses" refers to technologies or devices that create appropriate answers or suggestions for users based on acquired real-time information.
[0131] "Means for converting into audio data and outputting to the user" refers to a technology or device that converts generated text data into audio and outputs it to the user in a format that is audible.
[0132] A "generative model" is a machine learning model that generates appropriate responses in natural language based on input prompts.
[0133] This invention relates to a voice-operated navigation system for use while riding a motorcycle, enabling the user to obtain real-time information without using both hands and to travel safely and efficiently. The following describes specific embodiments of this invention.
[0134] Hardware and software to be used
[0135] This system utilizes the following hardware and software.
[0136] 1. Hardware:
[0137] Helmet with microphone or Bluetooth headset: A device for capturing the user's voice.
[0138] Smartphone or GPS device: A device for processing audio data, acquiring real-time information, and playing back generated audio output.
[0139] 2. Software:
[0140] Speech recognition APIs (e.g., Google Cloud Speech-to-Text API): Convert voice input into text data.
[0141] Natural language processing engine (e.g., IBM Watson® NLP): Analyzes text data and determines the user's intent.
[0142] Traffic information APIs and weather information APIs (e.g., Google Maps API, Weather API): Obtain real-time information.
[0143] Generative models (e.g., GPT-3®): Generate responses to the user based on acquired information.
[0144] Speech synthesis API (e.g., Amazon Polly): Converts the generated response into speech data.
[0145] Example of an operation flow
[0146] 1. Proposal of an alternative route
[0147] When a user uses voice input to say, "Check if the current road is congested," the system operates as follows:
[0148] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[0149] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[0150] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to check traffic information."
[0151] Server: Based on the determination result, it retrieves real-time traffic information from the Google Maps API.
[0152] Server: Passes the acquired traffic information to the generation model (GPT-3) and generates alternative route suggestions.
[0153] Generative model: Generates the response text "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[0154] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[0155] Terminal: Outputs the converted audio data to the user.
[0156] 2. Providing information on recommended spots
[0157] When a user uses voice input to say, "I'm hungry, can you recommend a good restaurant nearby?", the following process takes place.
[0158] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[0159] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[0160] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to find recommended restaurants."
[0161] Server: Based on the determination result, retrieves information about nearby restaurants from the corresponding database (e.g., Yelp API).
[0162] Server: Passes the acquired restaurant information to the generative model (GPT-3) to generate suggestions for the best restaurants.
[0163] Generative model: Generates the response text "There is a highly-rated Italian restaurant 1 kilometer away."
[0164] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[0165] Terminal: Outputs the converted audio data to the user.
[0166] Effects of implementation
[0167] This allows users to obtain real-time traffic information and recommended spots via voice control without using their hands while riding a motorcycle, enabling safe and efficient travel. Furthermore, the intuitive operation improves safety while driving.
[0168] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0169] Step 1:
[0170] The user wears a helmet with a microphone or a Bluetooth headset and provides voice input. For example, they might say, "Check if this road is congested." This input is captured as analog voice data. The terminal receives this voice data and begins processing it in the appropriate format.
[0171] Step 2:
[0172] The device sends the received audio data to the Google Cloud Speech-to-Text API. The audio data is sent as input, and the speech recognition API converts it into text data. The output is text data that says, "Check if the current road is congested." This text data is prepared for further processing.
[0173] Step 3:
[0174] The terminal sends the acquired text data to the server. The server passes the text data to IBM Watson NLP, which then starts the process of analyzing the user's intent. The input is text data, and the output is the analysis result indicating the user's intent. For example, the analysis result might be "check traffic information."
[0175] Step 4:
[0176] Based on the analysis results, the server retrieves real-time traffic information from the Google Maps API. The current location and surrounding area information are input as part of the API request. The returned output includes data on current traffic conditions and congestion information. Specifically, API requests and responses are processed.
[0177] Step 5:
[0178] The server sends the acquired traffic information as a prompt to the generative model (GPT-3). The input to the generative model is a prompt statement that combines the user's intent with real-time information. The output is the response text to the user. For example, the generated text might say, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[0179] Step 6:
[0180] The server sends the generated response text to the device and initiates the process of converting it into speech data using Amazon Polly. The input is the generated text, and the output is speech data. The speech synthesis API converts the text to speech, which the device receives.
[0181] Step 7:
[0182] The device plays the converted audio data to the user. The user receives voice guidance via a Bluetooth headset or helmet with a microphone, such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time." The specific action performed is the playback of the voice guidance.
[0183] (Application Example 1)
[0184] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0185] The objective of this invention is to efficiently acquire real-time inventory information and machine status within a factory using voice control, and to quickly provide optimal work instructions. Currently, factory workers have to manually check multiple data sets, which is inefficient and time-consuming.
[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0187] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, and means for analyzing the text data to determine the user's intent. This makes it possible to acquire real-time inventory information and machine status within the factory and generate optimal work instructions based on this information.
[0188] "Means for receiving voice input" refers to a combination of hardware and software for collecting voice commands from the user.
[0189] "Means for converting voice input into text data" refers to technology that recognizes collected voice input and converts it into corresponding text data.
[0190] "Means for analyzing text data and determining user intent" refers to algorithms or engines that analyze text data to understand the intent behind the instructions or information requested by the user.
[0191] "Means for obtaining real-time information based on analysis results" refers to a mechanism that obtains real-time information from relevant databases and information sources according to the analyzed user's intent.
[0192] "Means for generating responses based on acquired information" refers to technology that generates appropriate responses to provide to users based on acquired real-time information.
[0193] "A means of converting generated responses into audio data and outputting it to the user" refers to a technology that converts generated text responses into audio and conveys that audio to the user.
[0194] "A means of acquiring real-time inventory information and machine status within a factory and generating optimal work instructions based on this" refers to a system that collects the latest data on inventory information and machine operating status within a factory and provides optimal instructions to workers via voice based on that data.
[0195] This invention is a system that uses voice input to acquire inventory information and machine status in real time within a factory and provides optimal work instructions to the user via voice. This system is realized by combining technologies such as speech recognition, natural language processing, real-time data acquisition, generative AI, and speech synthesis.
[0196] Specifically, when a user provides voice input through a microphone, the system converts that voice input into text data. A speech recognition API is used for this conversion. Next, the converted text data is sent to a server and analyzed by a natural language processing engine. Based on the analysis results, the server calls various APIs to obtain inventory information and machine status in real time.
[0197] Based on the acquired information, the generative AI generates an appropriate response to the user. This generated response is temporarily stored as text data and then converted into audio data using a speech synthesis API. Finally, the generated audio data is output to the user.
[0198] The system components are as follows:
[0199] Means of receiving voice input: microphone or Bluetooth headset.
[0200] A means of converting voice input into text data: Speech recognition API.
[0201] Means for analyzing text data and determining user intent: Natural language processing engine (e.g., OpenAI® API).
[0202] Means of obtaining real-time information: Inventory management APIs and APIs for monitoring machine status.
[0203] Means for converting the generated response into audio data and outputting it to the user: Generative AI and speech synthesis APIs (e.g., Google Text-to-Speech).
[0204] As a concrete example in a factory, consider a scenario where a worker gives a voice command saying, "Tell me the current inventory count." In this case, the system converts the voice to text and sends that text data to a server. The server calls an inventory management API to obtain real-time inventory information and, based on the analysis results, uses generative AI to generate an appropriate response such as "The current inventory count is 100 units." This response is then converted back into voice data and output to the worker, providing real-time information.
[0205] Similarly, in the case of a voice command such as "Tell me the status of the machine," the system can call a machine monitoring API to obtain the latest status of the machine, and then generate and provide an appropriate response to the user based on that information.
[0206] By utilizing generative AI models, it is possible to generate appropriate responses even to complex questions. For example, the following example prompt can be used:
[0207] text
[0208] User: What is the current stock quantity?
[0209] response:
[0210] As a result, this system can improve work efficiency within the factory and enable the rapid and safe provision of information.
[0211] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0212] Step 1:
[0213] The user provides voice input through a microphone. Specifically, the user gives instructions such as, "Tell me the current stock quantity." This voice input serves as the initial trigger.
[0214] Input: User's voice instructions
[0215] Output: Audio data
[0216] Operation: Uses the microphone to capture audio.
[0217] Step 2:
[0218] The device converts the acquired audio data into text data using a speech recognition API.
[0219] Input: Audio data
[0220] Output: Text data
[0221] Operation: Calls a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text.
[0222] Step 3:
[0223] The terminal sends the converted text data to the server.
[0224] Input: Text data
[0225] Output: API request to the server
[0226] Operation: Sends text data to the server using an HTTP request.
[0227] Step 4:
[0228] The server receives text data and uses a natural language processing engine to analyze the user's intent.
[0229] Input: Text data
[0230] Output: Analysis results
[0231] Operation: Analyzes text data using a natural language processing engine (e.g., OpenAI) to extract the user's intent.
[0232] Step 5:
[0233] Based on the analysis results, the server calls APIs to obtain real-time inventory information and machine status.
[0234] Input: Analysis results
[0235] Output: Real-time data
[0236] Operation: Calls inventory management APIs and machine monitoring APIs to obtain necessary information.
[0237] Step 6:
[0238] Based on the real-time data it acquires, the server uses generative AI to generate appropriate responses for the user.
[0239] Input: Real-time data
[0240] Output: Response text
[0241] Operation: Uses generative AI (e.g., OpenAI) to generate response text based on real-time data.
[0242] Step 7:
[0243] The server calls a speech synthesis API to convert the generated response text into speech data.
[0244] Input: Response text
[0245] Output: Audio data
[0246] Operation: Converts text into speech data using a speech synthesis API (e.g., Google Text-to-Speech).
[0247] Step 8:
[0248] The terminal outputs the generated audio data to the user.
[0249] Input: Audio data
[0250] Output: Voice notification to the user
[0251] Operation: Plays audio data to the user using the speaker or headset connected to the device.
[0252] The system is designed to ensure that each step smoothly executes the entire process, from acquiring voice input and analyzing the data to obtaining real-time information and generating and converting appropriate responses into speech.
[0253] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0254] This invention is a system that allows motorcyclists to obtain real-time navigation information using voice commands while riding, enabling them to reach their destination safely and efficiently. In particular, this invention achieves more appropriate information provision and navigation by combining it with an emotion engine that recognizes the user's emotions.
[0255] Overall flow and process overview
[0256] 1. Acquisition of voice input:
[0257] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[0258] Example: "Check if the current road is congested."
[0259] 2. Speech-to-text conversion:
[0260] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[0261] 3. Recognition of emotions:
[0262] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[0263] Use an emotion recognition API (such as IBM Watson Tone Analyzer) to identify the user's emotions from their voice.
[0264] 4. Information Analysis:
[0265] The server receives text data and uses a natural language processing (NLP) engine to analyze the user's intent.
[0266] Based on text data and sentiment recognition results, the system determines the user's needs and emotional state.
[0267] 5. Obtaining real-time information:
[0268] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[0269] Example: Obtain current traffic and weather information.
[0270] 6. Generating the response:
[0271] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[0272] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0273] 7. Conversion to and output of audio data:
[0274] The response text generated by the generative AI is converted into voice data using a text-to-speech API.
[0275] The terminal outputs the voice data to the user.
[0276] Processing of specific examples
[0277] 1. Proposal of a detour route
[0278] User: "Check if the current road is congested."
[0279] The terminal converts the voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[0280] The server analyzes the text data and the emotion recognition results, calls the traffic information API to obtain real-time traffic information.
[0281] The server sends the analysis results to the generative AI based on the congestion information and calculates a detour route.
[0282] The generative AI generates a response such as "Turn left at the next intersection and enter the highway 3 kilometers ahead to save time" while considering the emotion recognition results, and converts it into voice data.
[0283] The terminal outputs the voice data to the user.
[0284] 2. Provision of recommended spot information
[0285] User: "I'm hungry, please tell me a recommended restaurant nearby."
[0286] The terminal converts the voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[0287] The server analyzes the text data and the emotion recognition results, calls the restaurant information API to obtain the latest information.
[0288] Based on the emotion recognition results, the generative AI generates a response such as "There is a highly-rated Italian restaurant 1 kilometer away" for users in a relaxed state, and converts it into audio data.
[0289] The device outputs audio data to the user.
[0290] System Configuration
[0291] This system includes the following:
[0292] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[0293] A means of converting audio data into text data (speech recognition API).
[0294] A means of analyzing text data and determining the user's intent (natural language processing engine).
[0295] A means of recognizing a user's emotions (emotion recognition API).
[0296] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[0297] A means of generating responses based on acquired information and emotion recognition results (generative AI).
[0298] A means of converting a response into audio data and outputting it (speech synthesis API).
[0299] These features allow users to receive appropriate information and navigation based on their emotional state, enabling safe and efficient travel even while riding a motorcycle. This system, in particular, significantly improves the user experience and supports stress-free travel by combining it with an emotion engine.
[0300] The following describes the processing flow.
[0301] Step 1:
[0302] The user gives voice instructions to the system through a microphone-equipped helmet or a Bluetooth headset.
[0303] Example: "Check if the current road is congested."
[0304] Step 2:
[0305] The terminal acquires the voice input and uses a voice recognition API to convert the voice into text data.
[0306] Call a voice recognition API (e.g., Google Cloud Speech-to-Text API) to convert the voice into text.
[0307] Temporarily save the converted text data in the terminal.
[0308] Step 3:
[0309] The terminal sends the voice data it acquired to an emotion recognition API to analyze the user's emotion.
[0310] Call an emotion recognition API (e.g., IBM Watson Tone Analyzer) to identify the user's emotion from the voice.
[0311] Temporarily save the analysis result (e.g., stress, anxiety, relaxation, etc.) in the terminal.
[0312] Step 4:
[0313] The terminal sends the text data and the emotion recognition result to the server.
[0314] Send an HTTP request to the server, including the text data and the emotion recognition result.
[0315] Step 5:
[0316] The server receives text data and emotion recognition results.
[0317] The received data is passed to the natural language processing (NLP) engine.
[0318] Step 6:
[0319] The server uses a natural language processing engine to analyze text data and determine the user's intent.
[0320] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[0321] Step 7:
[0322] The server determines the need for real-time information acquisition based on the analysis results and emotion recognition results.
[0323] Consider the case where the analysis results request "checking traffic information" and the emotion recognition result indicates "stress."
[0324] Step 8:
[0325] The server calls a traffic information API to obtain real-time traffic information.
[0326] Example: Use the Google Maps API to retrieve current traffic information.
[0327] The acquired traffic information is temporarily stored.
[0328] Step 9:
[0329] The server integrates real-time traffic information acquired by the system with analysis results and sends it to the generative AI.
[0330] Based on the analysis results, traffic information, and emotion recognition results, the data is sent to a generative AI engine.
[0331] Step 10:
[0332] The generative AI generates a response to the user.
[0333] For example, considering traffic information and the user's stress level, it can generate responses such as, "There is currently traffic congestion, so we suggest a relaxing detour. Turn left at the next intersection and take the scenic route 3 kilometers ahead."
[0334] Step 11:
[0335] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[0336] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[0337] Step 12:
[0338] The device receives audio data and outputs it to the user.
[0339] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[0340] Step 13:
[0341] The user listens to voice guidance and follows the suggested detour route.
[0342] Choose a safe new route and head towards your destination.
[0343] (Example 2)
[0344] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0345] Conventional navigation systems made it difficult for motorcyclists to operate them using voice commands, posing a risk of reduced safety. Furthermore, they lacked the ability to provide information tailored to the user's emotional state, resulting in a limited user experience. Additionally, the difficulty in providing real-time traffic information and recommended spots meant that efficient travel was challenging.
[0346] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0347] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data and determining the user's intent, means for recognizing the user's emotions from the text data, means for acquiring real-time information based on the analysis results and emotion recognition, means for generating a response based on the acquired information and emotion recognition, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire navigation information by voice operation even while riding a motorcycle, allowing for safe and efficient travel. Furthermore, it enables the provision of appropriate information according to the user's emotional state, significantly improving the user experience and realizing stress-free navigation.
[0348] 1. "Voice input" refers to instructions or data provided by the user via voice.
[0349] 2. "Text data" refers to digital information obtained by converting audio into text format.
[0350] 3. "Analysis" is the process of analyzing data to determine specific intentions or emotions.
[0351] 4. "User intent" refers to the actions a user wants to take or the information they want to obtain.
[0352] 5. "Emotion recognition" is a method of determining a user's emotional state from voice data.
[0353] 6. "Real-time information" refers to information that provides the latest data based on the current situation.
[0354] 7. A "response" is the answer that the system generates in response to user input.
[0355] 8. "Conversion to audio data" refers to the process of converting text data into an audio format.
[0356] 9. "Traffic information" refers to real-time traffic-related data such as current road conditions, congestion information, and accident information.
[0357] 10. A "detour route" is a proposed alternative route to avoid congestion or obstacles.
[0358] 11. "Recommended Spot Information" refers to information about places and facilities near the destination that are useful to the user.
[0359] 12. "Various information sources" refers to different information providers and systems used to collect data.
[0360] 13. "Generative AI" refers to artificial intelligence technology that generates responses to users based on analysis results and emotion recognition.
[0361] This invention provides a system that allows motorcycle riders to obtain real-time navigation information using voice commands and reach their destination safely and efficiently. This system is implemented by combining voice input, speech recognition, emotion recognition, natural language processing, real-time information acquisition, response generation, and speech synthesis.
[0362] Hardware and software configuration to be used
[0363] 1. User input means:
[0364] Users give voice commands using a helmet with a microphone or a Bluetooth headset.
[0365] 2. Voice recognition means:
[0366] The device receives voice input and uses a speech recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.
[0367] 3. Emotion recognition means:
[0368] The device sends the acquired text data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions.
[0369] 4. Natural language processing tools:
[0370] The server receives text data and sentiment recognition results, uses a natural language processing (NLP) engine (e.g., Google Natural Language API) to analyze the user's intent, and determines the user's needs and emotional state.
[0371] 5. Means of acquiring real-time information:
[0372] The server retrieves real-time data from APIs such as traffic information APIs (e.g., Google Maps API) and weather information APIs (e.g., OpenWeatherMap).
[0373] 6. Response generation means:
[0374] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[0375] 7. Speech synthesis means:
[0376] The response text generated by the generative AI is converted into speech data using a speech synthesis API (e.g., Amazon Polly).
[0377] The device outputs the generated audio data to the user.
[0378] Specific example behavior
[0379] Proposed detour route
[0380] 1. User: "Please check if the current road is congested."
[0381] The device converts voice input into text data and analyzes the emotion "irritated" through an emotion recognition API.
[0382] The server analyzes text data and sentiment recognition results, and uses a traffic information API to obtain real-time traffic congestion information.
[0383] The generative AI generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[0384] The device outputs audio data to the user.
[0385] Providing information on recommended spots
[0386] 1. User: "I'm hungry, can you recommend some good restaurants nearby?"
[0387] The device converts voice input into text data and analyzes the emotion of "being relaxed" through an emotion recognition API.
[0388] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[0389] The generative AI generates responses such as "There is a highly-rated Italian restaurant 1 kilometer away" and converts them into audio data.
[0390] The device outputs audio data to the user.
[0391] Example of a prompt
[0392] Prompt message for suggesting a detour route
[0393] Generate a response to the user's request, "Check if the current road is congested." The user's emotional state is "frustrated." Based on the current traffic information, suggest an alternative route.
[0394] Prompt message for providing recommended spot information
[0395] Generate a response to a user who asks, "I'm hungry, can you recommend a nearby restaurant?" The user's emotional state is "relaxed." Based on the latest restaurant information, recommend a suitable restaurant.
[0396] As described above, the present invention allows users to obtain real-time navigation information and recommended spot information by voice control even while riding a motorcycle. Furthermore, the introduction of an emotion engine enables the provision of appropriate information according to the user's emotional state, thereby providing a safe and comfortable navigation experience.
[0397] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0398] Step 1: Obtaining voice input
[0399] The user gives voice commands using a helmet with a microphone or a Bluetooth headset. The user can issue voice commands such as, "Check if this road is congested."
[0400] Input: User voice commands
[0401] Output: Audio data
[0402] Step 2: Convert speech to text
[0403] The device uses a speech recognition API (e.g., a speech recognition API) to convert the input speech data into text data.
[0404] Specifically, the speech recognition engine analyzes the audio data and generates corresponding text data.
[0405] Input: Audio data
[0406] Output: Text data (Example: "Check if the current road is congested")
[0407] Step 3: Recognizing Emotions
[0408] The device sends the acquired text data to an emotion recognition API (e.g., an emotion recognition API), which analyzes the user's emotions from the text. Specifically, the emotion recognition engine analyzes the text data and identifies emotional states such as "irritated."
[0409] Input: Text data
[0410] Output: Emotion recognition result (e.g., "irritated")
[0411] Step 4: Information Analysis
[0412] Based on the text data received by the server and the sentiment recognition results, a natural language processing (NLP) engine (e.g., a natural language processing engine) is used to analyze the user's intent. Specifically, the natural language processing engine analyzes the text data and identifies the user's instructions.
[0413] Input: Text data, sentiment recognition results
[0414] Output: User intent (e.g., "Check traffic information")
[0415] Step 5: Obtaining real-time information
[0416] Based on the user's intent, the server calls APIs such as traffic information APIs (e.g., Traffic Information API) and weather information APIs (e.g., Weather Information API) to obtain real-time data. Specifically, it sends requests to the corresponding API endpoints to obtain current traffic conditions and weather information.
[0417] Input: User intent
[0418] Output: Real-time information (e.g., "Traffic information for the current route")
[0419] Step 6: Generating the response
[0420] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results. Specifically, the generative AI considers the user's emotional state while generating an appropriate response (e.g., "Turning left at the next intersection and getting onto the highway 3 kilometers ahead will save you time").
[0421] Input: Analysis results, real-time information, emotion recognition results
[0422] Output: Response text
[0423] Step 7: Convert to audio data and output
[0424] The terminal converts the generated response text into audio data using a speech synthesis API (e.g., a speech synthesis API). Specifically, the speech synthesis engine analyzes the text data and generates an audio file. The generated audio data is then output to the user.
[0425] Input: Response text
[0426] Output: Audio data (Example: "There is traffic congestion on your current route, so you can save time by turning left at the next intersection and taking the highway 3 kilometers ahead.")
[0427] (Application Example 2)
[0428] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0429] Efficient and safe work navigation within factories is a critical issue directly linked to reducing worker stress and improving productivity. Conventional systems struggle to instantly acquire and provide real-time information on equipment status and work stations, and furthermore, they lack appropriate feedback tailored to the emotional state of workers, leading to the accumulation of stress and fatigue.
[0430] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data to determine the user's intent, means for recognizing the user's emotions, means for acquiring real-time information based on the analysis results, means for generating a response based on the acquired information, and means for converting the generated response into voice data and outputting it to the user. This enables efficient and safe work navigation within the factory, reducing worker stress and improving productivity.
[0431] "Means for receiving voice input" refers to a system element that acquires user voice commands through input devices such as microphones or head-mounted displays.
[0432] "Means for converting voice input into text data" refers to a system element that analyzes acquired voice and converts it into text data using speech recognition technology.
[0433] "Means for analyzing text data and determining user intent" refers to a system element that analyzes text data using natural language processing technology to understand and determine what the user wants.
[0434] "Means for recognizing user emotions" refers to system elements that include technologies for analyzing and recognizing a user's emotional state based on voice and text data.
[0435] "Means of acquiring real-time information" refers to elements of a system that acquires current conditions and the latest data in real time through networks and sensors.
[0436] "Means for generating a response based on acquired information" refers to a system element that includes technology for generating the optimal response to be provided to the user based on analysis results and information acquired in real time.
[0437] "Means for converting the generated response into audio data and outputting it to the user" refers to a system element that converts the generated text-based response into an audio format using speech synthesis technology and outputs it to the user.
[0438] As an example of carrying out this invention, we will take a work navigation system within a factory. The user wears a head-mounted display (HMD) and gives voice instructions to the system. The system uses the following hardware and software to provide real-time optimal work procedures and route guidance based on the user's instructions.
[0439] Hardware:
[0440] 1. Head-mounted display (HMD): It has a built-in microphone and speaker for voice input and output, and also provides visual information to the user.
[0441] 2. Server: Processes voice input from the user and performs speech recognition and response generation.
[0442] software:
[0443] 1. Speech Recognition API: For example, use Google Cloud Speech-to-Text. Convert voice input into text data.
[0444] 2. Emotion Recognition API: For example, using IBM Watson Tone Analyzer. Recognizes user emotions from text data.
[0445] 3. Natural Language Processing Engine (NLP Engine): For example, OpenAI GPT-3 is used. It analyzes text data and determines the user's intent.
[0446] 4. Real-time information acquisition method: Real-time information is acquired from sensor data within the factory and from an internal database.
[0447] 5. Generative AI: Generates appropriate responses based on the user's intent and emotion recognition results.
[0448] 6. Text-to-Speech API: For example, use Google Cloud Text-to-Speech. Convert the generated text response into speech data.
[0449] Specific processing steps:
[0450] The user inputs voice commands into the system through the HMD's built-in microphone. The HMD converts this voice input into text data, which is then sent to the server. A speech recognition API on the server converts this voice data back into text data, and then an emotion recognition API analyzes the user's emotions. The analysis results are passed to a natural language processing engine to determine the user's intent.
[0451] Next, the server acquires real-time equipment status and work station information through sensor data within the factory and an internal database. Based on the acquired real-time information and emotion recognition results, a generative AI generates an appropriate response. The generated response is converted into audio data using a speech synthesis API and output to the user through the HMD's built-in speaker. Visual information is also displayed on the HMD's screen.
[0452] Specific example:
[0453] Here's an example of a user checking the location of their next work station.
[0454] User: "Where is the next work station?"
[0455] 1. Voice recognition: "Where is the next work station?"
[0456] 2. Emotion Recognition: Calm tone (Emotion recognition result)
[0457] 3. Server analysis: Request ⇒ Search for next station information
[0458] 4. Real-time information acquisition: Obtain current location and work station status.
[0459] 5. Response generation: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[0460] 6. Speech Synthesis and Output: Output information via speech and display.
[0461] This system allows users to instantly grasp real-time factory information and the status of work stations through voice control, enabling them to work efficiently and safely.
[0462] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0463] Step 1:
[0464] The user inputs voice commands through the microphone on their head-mounted display (HMD). For example, they might say, "Where is the next work station?" This voice input becomes the system's initial input data.
[0465] Step 2:
[0466] The device (HMD) receives voice input and sends the voice data to a speech recognition API. The speech recognition API uses Google Cloud Speech-to-Text to convert the voice data into text data. The converted text data is output and used for the next process.
[0467] Step 3:
[0468] The terminal sends the converted text data to the emotion recognition API. The emotion recognition API uses IBM Watson Tone Analyzer to recognize the user's emotions from the text data. The resulting emotion recognition result is output, and this data is used for the next processing step.
[0469] Step 4:
[0470] The server receives text data and sentiment recognition results, and analyzes the data using a natural language processing engine (e.g., OpenAI GPT-3). The analysis determines the user's intent, and the result is output.
[0471] Step 5:
[0472] The server retrieves necessary real-time information according to the user's intent. It retrieves necessary information from sensor data within the factory and from internal databases, for example, obtaining information on the current equipment status and the next work station. This real-time information is output, and that data is used for the next processing.
[0473] Step 6:
[0474] Based on emotion recognition results and real-time information, the server uses generative AI to generate appropriate responses. For example, it might output the generated response text: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[0475] Step 7:
[0476] The server sends the generated response text to a text-to-speech API (e.g., Google Cloud Text-to-Speech). The text-to-speech API converts the response text into speech data, and that speech data is output.
[0477] Step 8:
[0478] The terminal outputs the generated audio data to the user via the HMD's speakers. It also displays the generated response text on the HMD's display, providing visual information simultaneously.
[0479] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0480] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0481] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0482] [Second Embodiment]
[0483] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0484] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0485] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0486] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0487] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0488] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0489] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0490] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0491] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0492] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0493] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0494] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0495] This invention relates to a navigation system for motorcycle riders, enabling them to obtain real-time information through voice control without using their hands, and to travel safely and efficiently. The processing and operation of the system's program are described in detail below.
[0496] Overall flow and process overview
[0497] 1. Acquisition of voice input:
[0498] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[0499] Example: "Check if the current road is congested."
[0500] 2. Speech-to-text conversion:
[0501] The device converts the received voice input into text data using a speech recognition API.
[0502] 3. Information Analysis:
[0503] The server receives text data and performs natural language processing to analyze the user's intent.
[0504] Based on the analysis results, appropriate real-time information is obtained for keywords such as "traffic congestion."
[0505] 4. Obtaining real-time information:
[0506] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[0507] The acquired traffic and weather information is integrated with the analysis results.
[0508] 5. Generating the response:
[0509] Generative AI generates responses to the user based on analysis results and acquired real-time information.
[0510] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0511] 6. Conversion to and output of audio data:
[0512] The generative AI converts the generated response text into speech data using a speech synthesis API.
[0513] The device outputs audio data to the user.
[0514] Processing of specific examples
[0515] 1. Proposal of an alternative route
[0516] User: "Check if the current road is congested."
[0517] The terminal converts voice input into text data and sends it to the server.
[0518] The server analyzes the text data and calls a traffic information API to obtain real-time traffic information.
[0519] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[0520] The generative AI generates a response such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts it into audio data.
[0521] The device outputs audio data to the user.
[0522] 2. Providing information on recommended spots
[0523] User: "I'm hungry, can you recommend some good restaurants nearby?"
[0524] The terminal converts voice input into text data and sends it to the server.
[0525] The server analyzes the text data and calls a restaurant information API to retrieve the latest information.
[0526] The server sends the acquired restaurant information to a generative AI, which then generates a response suggesting the most suitable restaurant.
[0527] The generative AI generates the response, "There is a highly-rated Italian restaurant 1 kilometer away," and converts it into audio data.
[0528] The device outputs audio data to the user.
[0529] System Configuration
[0530] This system includes the following:
[0531] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[0532] A means of converting audio data into text data (speech recognition API).
[0533] A means of analyzing text data and determining the user's intent (natural language processing engine).
[0534] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[0535] A means of generating a response based on acquired information (generative AI).
[0536] A means of converting a response into audio data and outputting it (speech synthesis API).
[0537] These features allow users to travel safely and efficiently through voice control. This system is particularly effective when riding a motorcycle, as it can provide timely and appropriate information even when both hands are occupied.
[0538] The following describes the processing flow.
[0539] Step 1:
[0540] The user gives voice commands through a helmet with a microphone or a Bluetooth headset.
[0541] Example: "Check if the current road is congested."
[0542] Step 2:
[0543] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[0544] It calls a speech recognition API (for example, the Google Cloud Speech-to-Text API) to convert speech to text.
[0545] The converted data is temporarily stored on the device.
[0546] Step 3:
[0547] The terminal sends text data to the server.
[0548] Send an HTTP request to the server, including text data.
[0549] Step 4:
[0550] The server receives text data.
[0551] The natural language processing (NLP) engine is invoked to analyze the received text data.
[0552] Step 5:
[0553] The server analyzes the text data using a natural language processing engine to determine the user's intent.
[0554] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[0555] Step 6:
[0556] The server determines the need to acquire real-time information based on the analysis results.
[0557] The analysis results indicate a request to "check traffic congestion information."
[0558] Step 7:
[0559] The server calls a traffic information API to obtain real-time traffic information.
[0560] Example: Use the Google Maps API to retrieve current traffic information.
[0561] The acquired traffic information is temporarily stored.
[0562] Step 8:
[0563] The server integrates real-time traffic information with analysis results and sends it to the generative AI.
[0564] Based on the analysis results and traffic information, the data is sent to a generative AI engine.
[0565] Step 9:
[0566] The generative AI generates a response to the user.
[0567] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0568] Generate response text.
[0569] Step 10:
[0570] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[0571] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[0572] Step 11:
[0573] The device receives audio data and outputs it to the user.
[0574] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[0575] Step 12:
[0576] The user listens to voice guidance and follows the suggested detour route.
[0577] Choose a safe new route and head towards your destination.
[0578] (Example 1)
[0579] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0580] It is not easy for motorcyclists to safely and efficiently obtain real-time traffic information and recommended spots without using both hands while riding. Conventional navigation systems require screen operation, which can compromise safety while driving. Therefore, there is a need for a more intuitive and safe information acquisition system that uses voice control.
[0581] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0582] In this invention, the server includes means for the user to give instructions by voice, means for converting the given voice instructions into text format, means for analyzing the text and determining the user's intent, means for acquiring real-time information based on the determined intent, a generative model for generating a detailed response based on the real-time information, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire necessary information in real time through voice operation and to move safely and efficiently.
[0583] "Voice input" refers to the act of capturing the user's voice into the system via a microphone or headset.
[0584] "Means of converting to text data" refers to technology or equipment for converting voice input into strings of characters.
[0585] "Means for determining user intent" refers to technologies or devices that analyze the content of received text data to understand the content of user requests or questions.
[0586] "Means for obtaining real-time information" refers to technologies or devices that obtain the latest data, such as current traffic conditions and weather, from appropriate sources based on analysis results.
[0587] "Means for generating responses" refers to technologies or devices that create appropriate answers or suggestions for users based on acquired real-time information.
[0588] "Means for converting into audio data and outputting to the user" refers to a technology or device that converts generated text data into audio and outputs it to the user in a format that is audible.
[0589] A "generative model" is a machine learning model that generates appropriate responses in natural language based on input prompts.
[0590] This invention relates to a voice-operated navigation system for use while riding a motorcycle, enabling the user to obtain real-time information without using both hands and to travel safely and efficiently. The following describes specific embodiments of this invention.
[0591] Hardware and software to be used
[0592] This system utilizes the following hardware and software.
[0593] 1. Hardware:
[0594] Helmet with microphone or Bluetooth headset: A device for capturing the user's voice.
[0595] Smartphone or GPS device: A device for processing audio data, acquiring real-time information, and playing back generated audio output.
[0596] 2. Software:
[0597] Speech recognition APIs (e.g., Google Cloud Speech-to-Text API): Convert voice input into text data.
[0598] Natural language processing engine (e.g., IBM Watson NLP): Analyzes text data to determine the user's intent.
[0599] Traffic information APIs and weather information APIs (e.g., Google Maps API, Weather API): Obtain real-time information.
[0600] Generative models (e.g., GPT-3): These models generate responses to the user based on the information they acquire.
[0601] Speech synthesis API (e.g., Amazon Polly): Converts the generated response into speech data.
[0602] Example of an operation flow
[0603] 1. Proposal of an alternative route
[0604] When a user uses voice input to say, "Check if the current road is congested," the system operates as follows:
[0605] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[0606] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[0607] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to check traffic information."
[0608] Server: Based on the determination result, it retrieves real-time traffic information from the Google Maps API.
[0609] Server: Passes the acquired traffic information to the generation model (GPT-3) and generates alternative route suggestions.
[0610] Generative model: Generates the response text "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[0611] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[0612] Terminal: Outputs the converted audio data to the user.
[0613] 2. Providing information on recommended spots
[0614] When a user uses voice input to say, "I'm hungry, can you recommend a good restaurant nearby?", the following process takes place.
[0615] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[0616] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[0617] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to find recommended restaurants."
[0618] Server: Based on the determination result, retrieves information about nearby restaurants from the corresponding database (e.g., Yelp API).
[0619] Server: Passes the acquired restaurant information to the generative model (GPT-3) to generate suggestions for the best restaurants.
[0620] Generative model: Generates the response text "There is a highly-rated Italian restaurant 1 kilometer away."
[0621] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[0622] Terminal: Outputs the converted audio data to the user.
[0623] Effects of implementation
[0624] This allows users to obtain real-time traffic information and recommended spots via voice control without using their hands while riding a motorcycle, enabling safe and efficient travel. Furthermore, the intuitive operation improves safety while driving.
[0625] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0626] Step 1:
[0627] The user wears a helmet with a microphone or a Bluetooth headset and provides voice input. For example, they might say, "Check if this road is congested." This input is captured as analog voice data. The terminal receives this voice data and begins processing it in the appropriate format.
[0628] Step 2:
[0629] The device sends the received audio data to the Google Cloud Speech-to-Text API. The audio data is sent as input, and the speech recognition API converts it into text data. The output is text data that says, "Check if the current road is congested." This text data is prepared for further processing.
[0630] Step 3:
[0631] The terminal sends the acquired text data to the server. The server passes the text data to IBM Watson NLP, which then starts the process of analyzing the user's intent. The input is text data, and the output is the analysis result indicating the user's intent. For example, the analysis result might be "check traffic information."
[0632] Step 4:
[0633] Based on the analysis results, the server retrieves real-time traffic information from the Google Maps API. The current location and surrounding area information are input as part of the API request. The returned output includes data on current traffic conditions and congestion information. Specifically, API requests and responses are processed.
[0634] Step 5:
[0635] The server sends the acquired traffic information as a prompt to the generative model (GPT-3). The input to the generative model is a prompt statement that combines the user's intent with real-time information. The output is the response text to the user. For example, the generated text might say, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[0636] Step 6:
[0637] The server sends the generated response text to the device and initiates the process of converting it into speech data using Amazon Polly. The input is the generated text, and the output is speech data. The speech synthesis API converts the text to speech, which the device receives.
[0638] Step 7:
[0639] The device plays the converted audio data to the user. The user receives voice guidance via a Bluetooth headset or helmet with a microphone, such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time." The specific action performed is the playback of the voice guidance.
[0640] (Application Example 1)
[0641] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0642] The objective of this invention is to efficiently acquire real-time inventory information and machine status within a factory using voice control, and to quickly provide optimal work instructions. Currently, factory workers have to manually check multiple data sets, which is inefficient and time-consuming.
[0643] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0644] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, and means for analyzing the text data to determine the user's intent. This makes it possible to acquire real-time inventory information and machine status within the factory and generate optimal work instructions based on this information.
[0645] "Means for receiving voice input" refers to a combination of hardware and software for collecting voice commands from the user.
[0646] "Means for converting voice input into text data" refers to technology that recognizes collected voice input and converts it into corresponding text data.
[0647] "Means for analyzing text data and determining user intent" refers to algorithms or engines that analyze text data to understand the intent behind the instructions or information requested by the user.
[0648] "Means for obtaining real-time information based on analysis results" refers to a mechanism that obtains real-time information from relevant databases and information sources according to the analyzed user's intent.
[0649] "Means for generating responses based on acquired information" refers to technology that generates appropriate responses to provide to users based on acquired real-time information.
[0650] "A means of converting generated responses into audio data and outputting it to the user" refers to a technology that converts generated text responses into audio and conveys that audio to the user.
[0651] "A means of acquiring real-time inventory information and machine status within a factory and generating optimal work instructions based on this" refers to a system that collects the latest data on inventory information and machine operating status within a factory and provides optimal instructions to workers via voice based on that data.
[0652] This invention is a system that uses voice input to acquire inventory information and machine status in real time within a factory and provides optimal work instructions to the user via voice. This system is realized by combining technologies such as speech recognition, natural language processing, real-time data acquisition, generative AI, and speech synthesis.
[0653] Specifically, when a user provides voice input through a microphone, the system converts that voice input into text data. A speech recognition API is used for this conversion. Next, the converted text data is sent to a server and analyzed by a natural language processing engine. Based on the analysis results, the server calls various APIs to obtain inventory information and machine status in real time.
[0654] Based on the acquired information, the generative AI generates an appropriate response to the user. This generated response is temporarily stored as text data and then converted into audio data using a speech synthesis API. Finally, the generated audio data is output to the user.
[0655] The system components are as follows:
[0656] Means of receiving voice input: microphone or Bluetooth headset.
[0657] A means of converting voice input into text data: Speech recognition API.
[0658] A means of analyzing text data and determining user intent: a natural language processing engine (e.g., OpenAI's API).
[0659] Means of obtaining real-time information: Inventory management APIs and APIs for monitoring machine status.
[0660] Means for converting the generated response into audio data and outputting it to the user: Generative AI and speech synthesis APIs (e.g., Google Text-to-Speech).
[0661] As a concrete example in a factory, consider a scenario where a worker gives a voice command saying, "Tell me the current inventory count." In this case, the system converts the voice to text and sends that text data to a server. The server calls an inventory management API to obtain real-time inventory information and, based on the analysis results, uses generative AI to generate an appropriate response such as "The current inventory count is 100 units." This response is then converted back into voice data and output to the worker, providing real-time information.
[0662] Similarly, in the case of a voice command such as "Tell me the status of the machine," the system can call a machine monitoring API to obtain the latest status of the machine, and then generate and provide an appropriate response to the user based on that information.
[0663] By utilizing generative AI models, it is possible to generate appropriate responses even to complex questions. For example, the following example prompt can be used:
[0664] text
[0665] User: What is the current stock quantity?
[0666] response:
[0667] As a result, this system can improve work efficiency within the factory and enable the rapid and safe provision of information.
[0668] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0669] Step 1:
[0670] The user provides voice input through a microphone. Specifically, the user gives instructions such as, "Tell me the current stock quantity." This voice input serves as the initial trigger.
[0671] Input: User's voice instructions
[0672] Output: Audio data
[0673] Operation: Uses the microphone to capture audio.
[0674] Step 2:
[0675] The device converts the acquired audio data into text data using a speech recognition API.
[0676] Input: Audio data
[0677] Output: Text data
[0678] Operation: Calls a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text.
[0679] Step 3:
[0680] The terminal sends the converted text data to the server.
[0681] Input: Text data
[0682] Output: API request to the server
[0683] Operation: Sends text data to the server using an HTTP request.
[0684] Step 4:
[0685] The server receives text data and uses a natural language processing engine to analyze the user's intent.
[0686] Input: Text data
[0687] Output: Analysis results
[0688] Operation: Analyzes text data using a natural language processing engine (e.g., OpenAI) to extract the user's intent.
[0689] Step 5:
[0690] Based on the analysis results, the server calls APIs to obtain real-time inventory information and machine status.
[0691] Input: Analysis results
[0692] Output: Real-time data
[0693] Operation: Calls inventory management APIs and machine monitoring APIs to obtain necessary information.
[0694] Step 6:
[0695] Based on the real-time data it acquires, the server uses generative AI to generate appropriate responses for the user.
[0696] Input: Real-time data
[0697] Output: Response text
[0698] Operation: Uses generative AI (e.g., OpenAI) to generate response text based on real-time data.
[0699] Step 7:
[0700] The server calls a speech synthesis API to convert the generated response text into speech data.
[0701] Input: Response text
[0702] Output: Audio data
[0703] Operation: Converts text into speech data using a speech synthesis API (e.g., Google Text-to-Speech).
[0704] Step 8:
[0705] The terminal outputs the generated audio data to the user.
[0706] Input: Audio data
[0707] Output: Voice notification to the user
[0708] Operation: Plays audio data to the user using the speaker or headset connected to the device.
[0709] The system is designed to ensure that each step smoothly executes the entire process, from acquiring voice input and analyzing the data to obtaining real-time information and generating and converting appropriate responses into speech.
[0710] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0711] This invention is a system that allows motorcyclists to obtain real-time navigation information using voice commands while riding, enabling them to reach their destination safely and efficiently. In particular, this invention achieves more appropriate information provision and navigation by combining it with an emotion engine that recognizes the user's emotions.
[0712] Overall flow and process overview
[0713] 1. Acquisition of voice input:
[0714] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[0715] Example: "Check if the current road is congested."
[0716] 2. Speech-to-text conversion:
[0717] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[0718] 3. Recognition of emotions:
[0719] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[0720] Use an emotion recognition API (such as IBM Watson Tone Analyzer) to identify the user's emotions from their voice.
[0721] 4. Information Analysis:
[0722] The server receives text data and uses a natural language processing (NLP) engine to analyze the user's intent.
[0723] Based on text data and sentiment recognition results, the system determines the user's needs and emotional state.
[0724] 5. Obtaining real-time information:
[0725] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[0726] Example: Obtain current traffic and weather information.
[0727] 6. Generating the response:
[0728] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[0729] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0730] 7. Conversion to and output of audio data:
[0731] The response text generated by the generative AI is converted into speech data using a speech synthesis API.
[0732] The device outputs audio data to the user.
[0733] Processing of specific examples
[0734] 1. Proposal of an alternative route
[0735] User: "Check if the current road is congested."
[0736] The device converts voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[0737] The server analyzes text data and emotion recognition results, and then calls a traffic information API to obtain real-time traffic information.
[0738] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[0739] The generative AI takes emotion recognition results into account and generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[0740] The device outputs audio data to the user.
[0741] 2. Providing information on recommended spots
[0742] User: "I'm hungry, can you recommend some good restaurants nearby?"
[0743] The device converts voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[0744] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[0745] Based on the emotion recognition results, the generative AI generates a response such as "There is a highly-rated Italian restaurant 1 kilometer away" for users in a relaxed state, and converts it into audio data.
[0746] The device outputs audio data to the user.
[0747] System Configuration
[0748] This system includes the following:
[0749] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[0750] A means of converting audio data into text data (speech recognition API).
[0751] A means of analyzing text data and determining the user's intent (natural language processing engine).
[0752] A means of recognizing a user's emotions (emotion recognition API).
[0753] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[0754] A means of generating responses based on acquired information and emotion recognition results (generative AI).
[0755] A means of converting a response into audio data and outputting it (speech synthesis API).
[0756] These features allow users to receive appropriate information and navigation based on their emotional state, enabling safe and efficient travel even while riding a motorcycle. This system, in particular, significantly improves the user experience and supports stress-free travel by combining it with an emotion engine.
[0757] The following describes the processing flow.
[0758] Step 1:
[0759] The user gives voice commands to the system via a helmet with a microphone or a Bluetooth headset.
[0760] Example: "Check if the current road is congested."
[0761] Step 2:
[0762] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[0763] This converts speech to text by calling a speech recognition API (for example, the Google Cloud Speech-to-Text API).
[0764] The converted text data is temporarily stored on the device.
[0765] Step 3:
[0766] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[0767] An emotion recognition API (for example, IBM Watson Tone Analyzer) is called to identify the user's emotions from their voice.
[0768] The analysis results (e.g., stress, anxiety, relaxation, etc.) are temporarily stored on the device.
[0769] Step 4:
[0770] The device sends text data and emotion recognition results to the server.
[0771] An HTTP request is sent to the server, including text data and sentiment recognition results.
[0772] Step 5:
[0773] The server receives text data and emotion recognition results.
[0774] The received data is passed to the natural language processing (NLP) engine.
[0775] Step 6:
[0776] The server uses a natural language processing engine to analyze text data and determine the user's intent.
[0777] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[0778] Step 7:
[0779] The server determines the need for real-time information acquisition based on the analysis results and emotion recognition results.
[0780] Consider the case where the analysis results request "checking traffic information" and the emotion recognition result indicates "stress."
[0781] Step 8:
[0782] The server calls a traffic information API to obtain real-time traffic information.
[0783] Example: Use the Google Maps API to retrieve current traffic information.
[0784] The acquired traffic information is temporarily stored.
[0785] Step 9:
[0786] The server integrates real-time traffic information acquired by the system with analysis results and sends it to the generative AI.
[0787] Based on the analysis results, traffic information, and emotion recognition results, the data is sent to a generative AI engine.
[0788] Step 10:
[0789] The generative AI generates a response to the user.
[0790] For example, considering traffic information and the user's stress level, it can generate responses such as, "There is currently traffic congestion, so we suggest a relaxing detour. Turn left at the next intersection and take the scenic route 3 kilometers ahead."
[0791] Step 11:
[0792] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[0793] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[0794] Step 12:
[0795] The device receives audio data and outputs it to the user.
[0796] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[0797] Step 13:
[0798] The user listens to voice guidance and follows the suggested detour route.
[0799] Choose a safe new route and head towards your destination.
[0800] (Example 2)
[0801] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0802] Conventional navigation systems made it difficult for motorcyclists to operate them using voice commands, posing a risk of reduced safety. Furthermore, they lacked the ability to provide information tailored to the user's emotional state, resulting in a limited user experience. Additionally, the difficulty in providing real-time traffic information and recommended spots meant that efficient travel was challenging.
[0803] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0804] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data and determining the user's intent, means for recognizing the user's emotions from the text data, means for acquiring real-time information based on the analysis results and emotion recognition, means for generating a response based on the acquired information and emotion recognition, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire navigation information by voice operation even while riding a motorcycle, allowing for safe and efficient travel. Furthermore, it enables the provision of appropriate information according to the user's emotional state, significantly improving the user experience and realizing stress-free navigation.
[0805] 1. "Voice input" refers to instructions or data provided by the user via voice.
[0806] 2. "Text data" refers to digital information obtained by converting audio into text format.
[0807] 3. "Analysis" is the process of analyzing data to determine specific intentions or emotions.
[0808] 4. "User intent" refers to the actions a user wants to take or the information they want to obtain.
[0809] 5. "Emotion recognition" is a method of determining a user's emotional state from voice data.
[0810] 6. "Real-time information" refers to information that provides the latest data based on the current situation.
[0811] 7. A "response" is the answer that the system generates in response to user input.
[0812] 8. "Conversion to audio data" refers to the process of converting text data into an audio format.
[0813] 9. "Traffic information" refers to real-time traffic-related data such as current road conditions, congestion information, and accident information.
[0814] 10. A "detour route" is a proposed alternative route to avoid congestion or obstacles.
[0815] 11. "Recommended Spot Information" refers to information about places and facilities near the destination that are useful to the user.
[0816] 12. "Various information sources" refers to different information providers and systems used to collect data.
[0817] 13. "Generative AI" refers to artificial intelligence technology that generates responses to users based on analysis results and emotion recognition.
[0818] This invention provides a system that allows motorcycle riders to obtain real-time navigation information using voice commands and reach their destination safely and efficiently. This system is implemented by combining voice input, speech recognition, emotion recognition, natural language processing, real-time information acquisition, response generation, and speech synthesis.
[0819] Hardware and software configuration to be used
[0820] 1. User input means:
[0821] Users give voice commands using a helmet with a microphone or a Bluetooth headset.
[0822] 2. Voice recognition means:
[0823] The device receives voice input and uses a speech recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.
[0824] 3. Emotion recognition means:
[0825] The device sends the acquired text data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions.
[0826] 4. Natural language processing tools:
[0827] The server receives text data and sentiment recognition results, uses a natural language processing (NLP) engine (e.g., Google Natural Language API) to analyze the user's intent, and determines the user's needs and emotional state.
[0828] 5. Means of acquiring real-time information:
[0829] The server retrieves real-time data from APIs such as traffic information APIs (e.g., Google Maps API) and weather information APIs (e.g., OpenWeatherMap).
[0830] 6. Response generation means:
[0831] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[0832] 7. Speech synthesis means:
[0833] The response text generated by the generative AI is converted into speech data using a speech synthesis API (e.g., Amazon Polly).
[0834] The device outputs the generated audio data to the user.
[0835] Specific example behavior
[0836] Proposed detour route
[0837] 1. User: "Please check if the current road is congested."
[0838] The device converts voice input into text data and analyzes the emotion "irritated" through an emotion recognition API.
[0839] The server analyzes text data and sentiment recognition results, and uses a traffic information API to obtain real-time traffic congestion information.
[0840] The generative AI generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[0841] The device outputs audio data to the user.
[0842] Providing information on recommended spots
[0843] 1. User: "I'm hungry, can you recommend some good restaurants nearby?"
[0844] The device converts voice input into text data and analyzes the emotion of "being relaxed" through an emotion recognition API.
[0845] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[0846] The generative AI generates responses such as "There is a highly-rated Italian restaurant 1 kilometer away" and converts them into audio data.
[0847] The device outputs audio data to the user.
[0848] Example of a prompt
[0849] Prompt message for suggesting a detour route
[0850] Generate a response to the user's request, "Check if the current road is congested." The user's emotional state is "frustrated." Based on the current traffic information, suggest an alternative route.
[0851] Prompt message for providing recommended spot information
[0852] Generate a response to a user who asks, "I'm hungry, can you recommend a nearby restaurant?" The user's emotional state is "relaxed." Based on the latest restaurant information, recommend a suitable restaurant.
[0853] As described above, the present invention allows users to obtain real-time navigation information and recommended spot information by voice control even while riding a motorcycle. Furthermore, the introduction of an emotion engine enables the provision of appropriate information according to the user's emotional state, thereby providing a safe and comfortable navigation experience.
[0854] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0855] Step 1: Obtaining voice input
[0856] The user gives voice commands using a helmet with a microphone or a Bluetooth headset. The user can issue voice commands such as, "Check if this road is congested."
[0857] Input: User voice commands
[0858] Output: Audio data
[0859] Step 2: Convert speech to text
[0860] The device uses a speech recognition API (e.g., a speech recognition API) to convert the input speech data into text data.
[0861] Specifically, the speech recognition engine analyzes the audio data and generates corresponding text data.
[0862] Input: Audio data
[0863] Output: Text data (Example: "Check if the current road is congested")
[0864] Step 3: Recognizing Emotions
[0865] The device sends the acquired text data to an emotion recognition API (e.g., an emotion recognition API), which analyzes the user's emotions from the text. Specifically, the emotion recognition engine analyzes the text data and identifies emotional states such as "irritated."
[0866] Input: Text data
[0867] Output: Emotion recognition result (e.g., "irritated")
[0868] Step 4: Information Analysis
[0869] Based on the text data received by the server and the sentiment recognition results, a natural language processing (NLP) engine (e.g., a natural language processing engine) is used to analyze the user's intent. Specifically, the natural language processing engine analyzes the text data and identifies the user's instructions.
[0870] Input: Text data, sentiment recognition results
[0871] Output: User intent (e.g., "Check traffic information")
[0872] Step 5: Obtaining real-time information
[0873] Based on the user's intent, the server calls APIs such as traffic information APIs (e.g., Traffic Information API) and weather information APIs (e.g., Weather Information API) to obtain real-time data. Specifically, it sends requests to the corresponding API endpoints to obtain current traffic conditions and weather information.
[0874] Input: User intent
[0875] Output: Real-time information (e.g., "Traffic information for the current route")
[0876] Step 6: Generating the response
[0877] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results. Specifically, the generative AI considers the user's emotional state while generating an appropriate response (e.g., "Turning left at the next intersection and getting onto the highway 3 kilometers ahead will save you time").
[0878] Input: Analysis results, real-time information, emotion recognition results
[0879] Output: Response text
[0880] Step 7: Convert to audio data and output
[0881] The terminal converts the generated response text into audio data using a speech synthesis API (e.g., a speech synthesis API). Specifically, the speech synthesis engine analyzes the text data and generates an audio file. The generated audio data is then output to the user.
[0882] Input: Response text
[0883] Output: Audio data (Example: "There is traffic congestion on your current route, so you can save time by turning left at the next intersection and taking the highway 3 kilometers ahead.")
[0884] (Application Example 2)
[0885] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0886] Efficient and safe work navigation within factories is a critical issue directly linked to reducing worker stress and improving productivity. Conventional systems struggle to instantly acquire and provide real-time information on equipment status and work stations, and furthermore, they lack appropriate feedback tailored to the emotional state of workers, leading to the accumulation of stress and fatigue.
[0887] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data to determine the user's intent, means for recognizing the user's emotions, means for acquiring real-time information based on the analysis results, means for generating a response based on the acquired information, and means for converting the generated response into voice data and outputting it to the user. This enables efficient and safe work navigation within the factory, reducing worker stress and improving productivity.
[0888] "Means for receiving voice input" refers to a system element that acquires user voice commands through input devices such as microphones or head-mounted displays.
[0889] "Means for converting voice input into text data" refers to a system element that analyzes acquired voice and converts it into text data using speech recognition technology.
[0890] "Means for analyzing text data and determining user intent" refers to a system element that analyzes text data using natural language processing technology to understand and determine what the user wants.
[0891] "Means for recognizing user emotions" refers to system elements that include technologies for analyzing and recognizing a user's emotional state based on voice and text data.
[0892] "Means of acquiring real-time information" refers to elements of a system that acquires current conditions and the latest data in real time through networks and sensors.
[0893] "Means for generating a response based on acquired information" refers to a system element that includes technology for generating the optimal response to be provided to the user based on analysis results and information acquired in real time.
[0894] "Means for converting the generated response into audio data and outputting it to the user" refers to a system element that converts the generated text-based response into an audio format using speech synthesis technology and outputs it to the user.
[0895] As an example of carrying out this invention, we will take a work navigation system within a factory. The user wears a head-mounted display (HMD) and gives voice instructions to the system. The system uses the following hardware and software to provide real-time optimal work procedures and route guidance based on the user's instructions.
[0896] Hardware:
[0897] 1. Head-mounted display (HMD): It has a built-in microphone and speaker for voice input and output, and also provides visual information to the user.
[0898] 2. Server: Processes voice input from the user and performs speech recognition and response generation.
[0899] software:
[0900] 1. Speech Recognition API: For example, use Google Cloud Speech-to-Text. Convert voice input into text data.
[0901] 2. Emotion Recognition API: For example, using IBM Watson Tone Analyzer. Recognizes user emotions from text data.
[0902] 3. Natural Language Processing Engine (NLP Engine): For example, OpenAI GPT-3 is used. It analyzes text data and determines the user's intent.
[0903] 4. Real-time information acquisition method: Real-time information is acquired from sensor data within the factory and from an internal database.
[0904] 5. Generative AI: Generates appropriate responses based on the user's intent and emotion recognition results.
[0905] 6. Text-to-Speech API: For example, use Google Cloud Text-to-Speech. Convert the generated text response into speech data.
[0906] Specific processing steps:
[0907] The user inputs voice commands into the system through the HMD's built-in microphone. The HMD converts this voice input into text data, which is then sent to the server. A speech recognition API on the server converts this voice data back into text data, and then an emotion recognition API analyzes the user's emotions. The analysis results are passed to a natural language processing engine to determine the user's intent.
[0908] Next, the server acquires real-time equipment status and work station information through sensor data within the factory and an internal database. Based on the acquired real-time information and emotion recognition results, a generative AI generates an appropriate response. The generated response is converted into audio data using a speech synthesis API and output to the user through the HMD's built-in speaker. Visual information is also displayed on the HMD's screen.
[0909] Specific example:
[0910] Here's an example of a user checking the location of their next work station.
[0911] User: "Where is the next work station?"
[0912] 1. Voice recognition: "Where is the next work station?"
[0913] 2. Emotion Recognition: Calm tone (Emotion recognition result)
[0914] 3. Server analysis: Request ⇒ Search for next station information
[0915] 4. Real-time information acquisition: Obtain current location and work station status.
[0916] 5. Response generation: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[0917] 6. Speech Synthesis and Output: Output information via speech and display.
[0918] This system allows users to instantly grasp real-time factory information and the status of work stations through voice control, enabling them to work efficiently and safely.
[0919] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0920] Step 1:
[0921] The user inputs voice commands through the microphone on their head-mounted display (HMD). For example, they might say, "Where is the next work station?" This voice input becomes the system's initial input data.
[0922] Step 2:
[0923] The device (HMD) receives voice input and sends the voice data to a speech recognition API. The speech recognition API uses Google Cloud Speech-to-Text to convert the voice data into text data. The converted text data is output and used for the next process.
[0924] Step 3:
[0925] The terminal sends the converted text data to the emotion recognition API. The emotion recognition API uses IBM Watson Tone Analyzer to recognize the user's emotions from the text data. The resulting emotion recognition result is output, and this data is used for the next processing step.
[0926] Step 4:
[0927] The server receives text data and sentiment recognition results, and analyzes the data using a natural language processing engine (e.g., OpenAI GPT-3). The analysis determines the user's intent, and the result is output.
[0928] Step 5:
[0929] The server retrieves necessary real-time information according to the user's intent. It retrieves necessary information from sensor data within the factory and from internal databases, for example, obtaining information on the current equipment status and the next work station. This real-time information is output, and that data is used for the next processing.
[0930] Step 6:
[0931] Based on emotion recognition results and real-time information, the server uses generative AI to generate appropriate responses. For example, it might output the generated response text: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[0932] Step 7:
[0933] The server sends the generated response text to a text-to-speech API (e.g., Google Cloud Text-to-Speech). The text-to-speech API converts the response text into speech data, and that speech data is output.
[0934] Step 8:
[0935] The terminal outputs the generated audio data to the user via the HMD's speakers. It also displays the generated response text on the HMD's display, providing visual information simultaneously.
[0936] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0937] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0938] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0939] [Third Embodiment]
[0940] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0941] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0942] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0943] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0944] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0945] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0946] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0947] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0948] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0949] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0950] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0951] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0952] This invention relates to a navigation system for motorcycle riders, enabling them to obtain real-time information through voice control without using their hands, and to travel safely and efficiently. The processing and operation of the system's program are described in detail below.
[0953] Overall flow and process overview
[0954] 1. Acquisition of voice input:
[0955] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[0956] Example: "Check if the current road is congested."
[0957] 2. Speech-to-text conversion:
[0958] The device converts the received voice input into text data using a speech recognition API.
[0959] 3. Information Analysis:
[0960] The server receives text data and performs natural language processing to analyze the user's intent.
[0961] Based on the analysis results, appropriate real-time information is obtained for keywords such as "traffic congestion."
[0962] 4. Obtaining real-time information:
[0963] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[0964] The acquired traffic and weather information is integrated with the analysis results.
[0965] 5. Generating the response:
[0966] Generative AI generates responses to the user based on analysis results and acquired real-time information.
[0967] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[0968] 6. Conversion to and output of audio data:
[0969] The generative AI converts the generated response text into speech data using a speech synthesis API.
[0970] The device outputs audio data to the user.
[0971] Processing of specific examples
[0972] 1. Proposal of an alternative route
[0973] User: "Check if the current road is congested."
[0974] The terminal converts voice input into text data and sends it to the server.
[0975] The server analyzes the text data and calls a traffic information API to obtain real-time traffic information.
[0976] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[0977] The generative AI generates a response such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts it into audio data.
[0978] The device outputs audio data to the user.
[0979] 2. Providing information on recommended spots
[0980] User: "I'm hungry, can you recommend some good restaurants nearby?"
[0981] The terminal converts voice input into text data and sends it to the server.
[0982] The server analyzes the text data and calls a restaurant information API to retrieve the latest information.
[0983] The server sends the acquired restaurant information to a generative AI, which then generates a response suggesting the most suitable restaurant.
[0984] The generative AI generates the response, "There is a highly-rated Italian restaurant 1 kilometer away," and converts it into audio data.
[0985] The device outputs audio data to the user.
[0986] System Configuration
[0987] This system includes the following:
[0988] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[0989] A means of converting audio data into text data (speech recognition API).
[0990] A means of analyzing text data and determining the user's intent (natural language processing engine).
[0991] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[0992] A means of generating a response based on acquired information (generative AI).
[0993] A means of converting a response into audio data and outputting it (speech synthesis API).
[0994] These features allow users to travel safely and efficiently through voice control. This system is particularly effective when riding a motorcycle, as it can provide timely and appropriate information even when both hands are occupied.
[0995] The following describes the processing flow.
[0996] Step 1:
[0997] The user gives voice commands through a helmet with a microphone or a Bluetooth headset.
[0998] Example: "Check if the current road is congested."
[0999] Step 2:
[1000] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[1001] It calls a speech recognition API (for example, the Google Cloud Speech-to-Text API) to convert speech to text.
[1002] The converted data is temporarily stored on the device.
[1003] Step 3:
[1004] The terminal sends text data to the server.
[1005] Send an HTTP request to the server, including text data.
[1006] Step 4:
[1007] The server receives text data.
[1008] The natural language processing (NLP) engine is invoked to analyze the received text data.
[1009] Step 5:
[1010] The server analyzes the text data using a natural language processing engine to determine the user's intent.
[1011] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[1012] Step 6:
[1013] The server determines the need to acquire real-time information based on the analysis results.
[1014] The analysis results indicate a request to "check traffic congestion information."
[1015] Step 7:
[1016] The server calls a traffic information API to obtain real-time traffic information.
[1017] Example: Use the Google Maps API to retrieve current traffic information.
[1018] The acquired traffic information is temporarily stored.
[1019] Step 8:
[1020] The server integrates real-time traffic information with analysis results and sends it to the generative AI.
[1021] Based on the analysis results and traffic information, the data is sent to a generative AI engine.
[1022] Step 9:
[1023] The generative AI generates a response to the user.
[1024] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[1025] Generate response text.
[1026] Step 10:
[1027] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[1028] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[1029] Step 11:
[1030] The device receives audio data and outputs it to the user.
[1031] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[1032] Step 12:
[1033] The user listens to voice guidance and follows the suggested detour route.
[1034] Choose a safe new route and head towards your destination.
[1035] (Example 1)
[1036] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1037] It is not easy for motorcyclists to safely and efficiently obtain real-time traffic information and recommended spots without using both hands while riding. Conventional navigation systems require screen operation, which can compromise safety while driving. Therefore, there is a need for a more intuitive and safe information acquisition system that uses voice control.
[1038] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1039] In this invention, the server includes means for the user to give instructions by voice, means for converting the given voice instructions into text format, means for analyzing the text and determining the user's intent, means for acquiring real-time information based on the determined intent, a generative model for generating a detailed response based on the real-time information, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire necessary information in real time through voice operation and to move safely and efficiently.
[1040] "Voice input" refers to the act of capturing the user's voice into the system via a microphone or headset.
[1041] "Means of converting to text data" refers to technology or equipment for converting voice input into strings of characters.
[1042] "Means for determining user intent" refers to technologies or devices that analyze the content of received text data to understand the content of user requests or questions.
[1043] "Means for obtaining real-time information" refers to technologies or devices that obtain the latest data, such as current traffic conditions and weather, from appropriate sources based on analysis results.
[1044] "Means for generating responses" refers to technologies or devices that create appropriate answers or suggestions for users based on acquired real-time information.
[1045] "Means for converting into audio data and outputting to the user" refers to a technology or device that converts generated text data into audio and outputs it to the user in a format that is audible.
[1046] A "generative model" is a machine learning model that generates appropriate responses in natural language based on input prompts.
[1047] This invention relates to a voice-operated navigation system for use while riding a motorcycle, enabling the user to obtain real-time information without using both hands and to travel safely and efficiently. The following describes specific embodiments of this invention.
[1048] Hardware and software to be used
[1049] This system utilizes the following hardware and software.
[1050] 1. Hardware:
[1051] Helmet with microphone or Bluetooth headset: A device for capturing the user's voice.
[1052] Smartphone or GPS device: A device for processing audio data, acquiring real-time information, and playing back generated audio output.
[1053] 2. Software:
[1054] Speech recognition APIs (e.g., Google Cloud Speech-to-Text API): Convert voice input into text data.
[1055] Natural language processing engine (e.g., IBM Watson NLP): Analyzes text data to determine the user's intent.
[1056] Traffic information APIs and weather information APIs (e.g., Google Maps API, Weather API): Obtain real-time information.
[1057] Generative models (e.g., GPT-3): These models generate responses to the user based on the information they acquire.
[1058] Speech synthesis API (e.g., Amazon Polly): Converts the generated response into speech data.
[1059] Example of an operation flow
[1060] 1. Proposal of an alternative route
[1061] When a user uses voice input to say, "Check if the current road is congested," the system operates as follows:
[1062] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[1063] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[1064] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to check traffic information."
[1065] Server: Based on the determination result, it retrieves real-time traffic information from the Google Maps API.
[1066] Server: Passes the acquired traffic information to the generation model (GPT-3) and generates alternative route suggestions.
[1067] Generative model: Generates the response text "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[1068] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[1069] Terminal: Outputs the converted audio data to the user.
[1070] 2. Providing information on recommended spots
[1071] When a user uses voice input to say, "I'm hungry, can you recommend a good restaurant nearby?", the following process takes place.
[1072] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[1073] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[1074] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to find recommended restaurants."
[1075] Server: Based on the determination result, retrieves information about nearby restaurants from the corresponding database (e.g., Yelp API).
[1076] Server: Passes the acquired restaurant information to the generative model (GPT-3) to generate suggestions for the best restaurants.
[1077] Generative model: Generates the response text "There is a highly-rated Italian restaurant 1 kilometer away."
[1078] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[1079] Terminal: Outputs the converted audio data to the user.
[1080] Effects of implementation
[1081] This allows users to obtain real-time traffic information and recommended spots via voice control without using their hands while riding a motorcycle, enabling safe and efficient travel. Furthermore, the intuitive operation improves safety while driving.
[1082] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1083] Step 1:
[1084] The user wears a helmet with a microphone or a Bluetooth headset and provides voice input. For example, they might say, "Check if this road is congested." This input is captured as analog voice data. The terminal receives this voice data and begins processing it in the appropriate format.
[1085] Step 2:
[1086] The device sends the received audio data to the Google Cloud Speech-to-Text API. The audio data is sent as input, and the speech recognition API converts it into text data. The output is text data that says, "Check if the current road is congested." This text data is prepared for further processing.
[1087] Step 3:
[1088] The terminal sends the acquired text data to the server. The server passes the text data to IBM Watson NLP, which then starts the process of analyzing the user's intent. The input is text data, and the output is the analysis result indicating the user's intent. For example, the analysis result might be "check traffic information."
[1089] Step 4:
[1090] Based on the analysis results, the server retrieves real-time traffic information from the Google Maps API. The current location and surrounding area information are input as part of the API request. The returned output includes data on current traffic conditions and congestion information. Specifically, API requests and responses are processed.
[1091] Step 5:
[1092] The server sends the acquired traffic information as a prompt to the generative model (GPT-3). The input to the generative model is a prompt statement that combines the user's intent with real-time information. The output is the response text to the user. For example, the generated text might say, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[1093] Step 6:
[1094] The server sends the generated response text to the device and initiates the process of converting it into speech data using Amazon Polly. The input is the generated text, and the output is speech data. The speech synthesis API converts the text to speech, which the device receives.
[1095] Step 7:
[1096] The device plays the converted audio data to the user. The user receives voice guidance via a Bluetooth headset or helmet with a microphone, such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time." The specific action performed is the playback of the voice guidance.
[1097] (Application Example 1)
[1098] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1099] The objective of this invention is to efficiently acquire real-time inventory information and machine status within a factory using voice control, and to quickly provide optimal work instructions. Currently, factory workers have to manually check multiple data sets, which is inefficient and time-consuming.
[1100] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1101] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, and means for analyzing the text data to determine the user's intent. This makes it possible to acquire real-time inventory information and machine status within the factory and generate optimal work instructions based on this information.
[1102] "Means for receiving voice input" refers to a combination of hardware and software for collecting voice commands from the user.
[1103] "Means for converting voice input into text data" refers to technology that recognizes collected voice input and converts it into corresponding text data.
[1104] "Means for analyzing text data and determining user intent" refers to algorithms or engines that analyze text data to understand the intent behind the instructions or information requested by the user.
[1105] "Means for obtaining real-time information based on analysis results" refers to a mechanism that obtains real-time information from relevant databases and information sources according to the analyzed user's intent.
[1106] "Means for generating responses based on acquired information" refers to technology that generates appropriate responses to provide to users based on acquired real-time information.
[1107] "A means of converting generated responses into audio data and outputting it to the user" refers to a technology that converts generated text responses into audio and conveys that audio to the user.
[1108] "A means of acquiring real-time inventory information and machine status within a factory and generating optimal work instructions based on this" refers to a system that collects the latest data on inventory information and machine operating status within a factory and provides optimal instructions to workers via voice based on that data.
[1109] This invention is a system that uses voice input to acquire inventory information and machine status in real time within a factory and provides optimal work instructions to the user via voice. This system is realized by combining technologies such as speech recognition, natural language processing, real-time data acquisition, generative AI, and speech synthesis.
[1110] Specifically, when a user provides voice input through a microphone, the system converts that voice input into text data. A speech recognition API is used for this conversion. Next, the converted text data is sent to a server and analyzed by a natural language processing engine. Based on the analysis results, the server calls various APIs to obtain inventory information and machine status in real time.
[1111] Based on the acquired information, the generative AI generates an appropriate response to the user. This generated response is temporarily stored as text data and then converted into audio data using a speech synthesis API. Finally, the generated audio data is output to the user.
[1112] The system components are as follows:
[1113] Means of receiving voice input: microphone or Bluetooth headset.
[1114] A means of converting voice input into text data: Speech recognition API.
[1115] A means of analyzing text data and determining user intent: a natural language processing engine (e.g., OpenAI's API).
[1116] Means of obtaining real-time information: Inventory management APIs and APIs for monitoring machine status.
[1117] Means for converting the generated response into audio data and outputting it to the user: Generative AI and speech synthesis APIs (e.g., Google Text-to-Speech).
[1118] As a concrete example in a factory, consider a scenario where a worker gives a voice command saying, "Tell me the current inventory count." In this case, the system converts the voice to text and sends that text data to a server. The server calls an inventory management API to obtain real-time inventory information and, based on the analysis results, uses generative AI to generate an appropriate response such as "The current inventory count is 100 units." This response is then converted back into voice data and output to the worker, providing real-time information.
[1119] Similarly, in the case of a voice command such as "Tell me the status of the machine," the system can call a machine monitoring API to obtain the latest status of the machine, and then generate and provide an appropriate response to the user based on that information.
[1120] By utilizing generative AI models, it is possible to generate appropriate responses even to complex questions. For example, the following example prompt can be used:
[1121] text
[1122] User: What is the current stock quantity?
[1123] response:
[1124] As a result, this system can improve work efficiency within the factory and enable the rapid and safe provision of information.
[1125] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1126] Step 1:
[1127] The user provides voice input through a microphone. Specifically, the user gives instructions such as, "Tell me the current stock quantity." This voice input serves as the initial trigger.
[1128] Input: User's voice instructions
[1129] Output: Audio data
[1130] Operation: Uses the microphone to capture audio.
[1131] Step 2:
[1132] The device converts the acquired audio data into text data using a speech recognition API.
[1133] Input: Audio data
[1134] Output: Text data
[1135] Operation: Calls a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text.
[1136] Step 3:
[1137] The terminal sends the converted text data to the server.
[1138] Input: Text data
[1139] Output: API request to the server
[1140] Operation: Sends text data to the server using an HTTP request.
[1141] Step 4:
[1142] The server receives text data and uses a natural language processing engine to analyze the user's intent.
[1143] Input: Text data
[1144] Output: Analysis results
[1145] Operation: Analyzes text data using a natural language processing engine (e.g., OpenAI) to extract the user's intent.
[1146] Step 5:
[1147] Based on the analysis results, the server calls APIs to obtain real-time inventory information and machine status.
[1148] Input: Analysis results
[1149] Output: Real-time data
[1150] Operation: Calls inventory management APIs and machine monitoring APIs to obtain necessary information.
[1151] Step 6:
[1152] Based on the real-time data it acquires, the server uses generative AI to generate appropriate responses for the user.
[1153] Input: Real-time data
[1154] Output: Response text
[1155] Operation: Uses generative AI (e.g., OpenAI) to generate response text based on real-time data.
[1156] Step 7:
[1157] The server calls a speech synthesis API to convert the generated response text into speech data.
[1158] Input: Response text
[1159] Output: Audio data
[1160] Operation: Converts text into speech data using a speech synthesis API (e.g., Google Text-to-Speech).
[1161] Step 8:
[1162] The terminal outputs the generated audio data to the user.
[1163] Input: Audio data
[1164] Output: Voice notification to the user
[1165] Operation: Plays audio data to the user using the speaker or headset connected to the device.
[1166] The system is designed to ensure that each step smoothly executes the entire process, from acquiring voice input and analyzing the data to obtaining real-time information and generating and converting appropriate responses into speech.
[1167] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1168] This invention is a system that allows motorcyclists to obtain real-time navigation information using voice commands while riding, enabling them to reach their destination safely and efficiently. In particular, this invention achieves more appropriate information provision and navigation by combining it with an emotion engine that recognizes the user's emotions.
[1169] Overall flow and process overview
[1170] 1. Acquisition of voice input:
[1171] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[1172] Example: "Check if the current road is congested."
[1173] 2. Speech-to-text conversion:
[1174] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[1175] 3. Recognition of emotions:
[1176] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[1177] Use an emotion recognition API (such as IBM Watson Tone Analyzer) to identify the user's emotions from their voice.
[1178] 4. Information Analysis:
[1179] The server receives text data and uses a natural language processing (NLP) engine to analyze the user's intent.
[1180] Based on text data and sentiment recognition results, the system determines the user's needs and emotional state.
[1181] 5. Obtaining real-time information:
[1182] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[1183] Example: Obtain current traffic and weather information.
[1184] 6. Generating the response:
[1185] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[1186] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[1187] 7. Conversion to and output of audio data:
[1188] The response text generated by the generative AI is converted into speech data using a speech synthesis API.
[1189] The device outputs audio data to the user.
[1190] Processing of specific examples
[1191] 1. Proposal of an alternative route
[1192] User: "Check if the current road is congested."
[1193] The device converts voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[1194] The server analyzes text data and emotion recognition results, and then calls a traffic information API to obtain real-time traffic information.
[1195] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[1196] The generative AI takes emotion recognition results into account and generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[1197] The device outputs audio data to the user.
[1198] 2. Providing information on recommended spots
[1199] User: "I'm hungry, can you recommend some good restaurants nearby?"
[1200] The device converts voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[1201] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[1202] Based on the emotion recognition results, the generative AI generates a response such as "There is a highly-rated Italian restaurant 1 kilometer away" for users in a relaxed state, and converts it into audio data.
[1203] The device outputs audio data to the user.
[1204] System Configuration
[1205] This system includes the following:
[1206] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[1207] A means of converting audio data into text data (speech recognition API).
[1208] A means of analyzing text data and determining the user's intent (natural language processing engine).
[1209] A means of recognizing a user's emotions (emotion recognition API).
[1210] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[1211] A means of generating responses based on acquired information and emotion recognition results (generative AI).
[1212] A means of converting a response into audio data and outputting it (speech synthesis API).
[1213] These features allow users to receive appropriate information and navigation based on their emotional state, enabling safe and efficient travel even while riding a motorcycle. This system, in particular, significantly improves the user experience and supports stress-free travel by combining it with an emotion engine.
[1214] The following describes the processing flow.
[1215] Step 1:
[1216] The user gives voice commands to the system via a helmet with a microphone or a Bluetooth headset.
[1217] Example: "Check if the current road is congested."
[1218] Step 2:
[1219] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[1220] This converts speech to text by calling a speech recognition API (for example, the Google Cloud Speech-to-Text API).
[1221] The converted text data is temporarily stored on the device.
[1222] Step 3:
[1223] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[1224] An emotion recognition API (for example, IBM Watson Tone Analyzer) is called to identify the user's emotions from their voice.
[1225] The analysis results (e.g., stress, anxiety, relaxation, etc.) are temporarily stored on the device.
[1226] Step 4:
[1227] The device sends text data and emotion recognition results to the server.
[1228] An HTTP request is sent to the server, including text data and sentiment recognition results.
[1229] Step 5:
[1230] The server receives text data and emotion recognition results.
[1231] The received data is passed to the natural language processing (NLP) engine.
[1232] Step 6:
[1233] The server uses a natural language processing engine to analyze text data and determine the user's intent.
[1234] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[1235] Step 7:
[1236] The server determines the need for real-time information acquisition based on the analysis results and emotion recognition results.
[1237] Consider the case where the analysis results request "checking traffic information" and the emotion recognition result indicates "stress."
[1238] Step 8:
[1239] The server calls a traffic information API to obtain real-time traffic information.
[1240] Example: Use the Google Maps API to retrieve current traffic information.
[1241] The acquired traffic information is temporarily stored.
[1242] Step 9:
[1243] The server integrates real-time traffic information acquired by the system with analysis results and sends it to the generative AI.
[1244] Based on the analysis results, traffic information, and emotion recognition results, the data is sent to a generative AI engine.
[1245] Step 10:
[1246] The generative AI generates a response to the user.
[1247] For example, considering traffic information and the user's stress level, it can generate responses such as, "There is currently traffic congestion, so we suggest a relaxing detour. Turn left at the next intersection and take the scenic route 3 kilometers ahead."
[1248] Step 11:
[1249] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[1250] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[1251] Step 12:
[1252] The device receives audio data and outputs it to the user.
[1253] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[1254] Step 13:
[1255] The user listens to voice guidance and follows the suggested detour route.
[1256] Choose a safe new route and head towards your destination.
[1257] (Example 2)
[1258] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1259] Conventional navigation systems made it difficult for motorcyclists to operate them using voice commands, posing a risk of reduced safety. Furthermore, they lacked the ability to provide information tailored to the user's emotional state, resulting in a limited user experience. Additionally, the difficulty in providing real-time traffic information and recommended spots meant that efficient travel was challenging.
[1260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1261] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data and determining the user's intent, means for recognizing the user's emotions from the text data, means for acquiring real-time information based on the analysis results and emotion recognition, means for generating a response based on the acquired information and emotion recognition, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire navigation information by voice operation even while riding a motorcycle, allowing for safe and efficient travel. Furthermore, it enables the provision of appropriate information according to the user's emotional state, significantly improving the user experience and realizing stress-free navigation.
[1262] 1. "Voice input" refers to instructions or data provided by the user via voice.
[1263] 2. "Text data" refers to digital information obtained by converting audio into text format.
[1264] 3. "Analysis" is the process of analyzing data to determine specific intentions or emotions.
[1265] 4. "User intent" refers to the actions a user wants to take or the information they want to obtain.
[1266] 5. "Emotion recognition" is a method of determining a user's emotional state from voice data.
[1267] 6. "Real-time information" refers to information that provides the latest data based on the current situation.
[1268] 7. A "response" is the answer that the system generates in response to user input.
[1269] 8. "Conversion to audio data" refers to the process of converting text data into an audio format.
[1270] 9. "Traffic information" refers to real-time traffic-related data such as current road conditions, congestion information, and accident information.
[1271] 10. A "detour route" is a proposed alternative route to avoid congestion or obstacles.
[1272] 11. "Recommended Spot Information" refers to information about places and facilities near the destination that are useful to the user.
[1273] 12. "Various information sources" refers to different information providers and systems used to collect data.
[1274] 13. "Generative AI" refers to artificial intelligence technology that generates responses to users based on analysis results and emotion recognition.
[1275] This invention provides a system that allows motorcycle riders to obtain real-time navigation information using voice commands and reach their destination safely and efficiently. This system is implemented by combining voice input, speech recognition, emotion recognition, natural language processing, real-time information acquisition, response generation, and speech synthesis.
[1276] Hardware and software configuration to be used
[1277] 1. User input means:
[1278] Users give voice commands using a helmet with a microphone or a Bluetooth headset.
[1279] 2. Voice recognition means:
[1280] The device receives voice input and uses a speech recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.
[1281] 3. Emotion recognition means:
[1282] The device sends the acquired text data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions.
[1283] 4. Natural language processing tools:
[1284] The server receives text data and sentiment recognition results, uses a natural language processing (NLP) engine (e.g., Google Natural Language API) to analyze the user's intent, and determines the user's needs and emotional state.
[1285] 5. Means of acquiring real-time information:
[1286] The server retrieves real-time data from APIs such as traffic information APIs (e.g., Google Maps API) and weather information APIs (e.g., OpenWeatherMap).
[1287] 6. Response generation means:
[1288] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[1289] 7. Speech synthesis means:
[1290] The response text generated by the generative AI is converted into speech data using a speech synthesis API (e.g., Amazon Polly).
[1291] The device outputs the generated audio data to the user.
[1292] Specific example behavior
[1293] Proposed detour route
[1294] 1. User: "Please check if the current road is congested."
[1295] The device converts voice input into text data and analyzes the emotion "irritated" through an emotion recognition API.
[1296] The server analyzes text data and sentiment recognition results, and uses a traffic information API to obtain real-time traffic congestion information.
[1297] The generative AI generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[1298] The device outputs audio data to the user.
[1299] Providing information on recommended spots
[1300] 1. User: "I'm hungry, can you recommend some good restaurants nearby?"
[1301] The device converts voice input into text data and analyzes the emotion of "being relaxed" through an emotion recognition API.
[1302] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[1303] The generative AI generates responses such as "There is a highly-rated Italian restaurant 1 kilometer away" and converts them into audio data.
[1304] The device outputs audio data to the user.
[1305] Example of a prompt
[1306] Prompt message for suggesting a detour route
[1307] Generate a response to the user's request, "Check if the current road is congested." The user's emotional state is "frustrated." Based on the current traffic information, suggest an alternative route.
[1308] Prompt message for providing recommended spot information
[1309] Generate a response to a user who asks, "I'm hungry, can you recommend a nearby restaurant?" The user's emotional state is "relaxed." Based on the latest restaurant information, recommend a suitable restaurant.
[1310] As described above, the present invention allows users to obtain real-time navigation information and recommended spot information by voice control even while riding a motorcycle. Furthermore, the introduction of an emotion engine enables the provision of appropriate information according to the user's emotional state, thereby providing a safe and comfortable navigation experience.
[1311] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1312] Step 1: Obtaining voice input
[1313] The user gives voice commands using a helmet with a microphone or a Bluetooth headset. The user can issue voice commands such as, "Check if this road is congested."
[1314] Input: User voice commands
[1315] Output: Audio data
[1316] Step 2: Convert speech to text
[1317] The device uses a speech recognition API (e.g., a speech recognition API) to convert the input speech data into text data.
[1318] Specifically, the speech recognition engine analyzes the audio data and generates corresponding text data.
[1319] Input: Audio data
[1320] Output: Text data (Example: "Check if the current road is congested")
[1321] Step 3: Recognizing Emotions
[1322] The device sends the acquired text data to an emotion recognition API (e.g., an emotion recognition API), which analyzes the user's emotions from the text. Specifically, the emotion recognition engine analyzes the text data and identifies emotional states such as "irritated."
[1323] Input: Text data
[1324] Output: Emotion recognition result (e.g., "irritated")
[1325] Step 4: Information Analysis
[1326] Based on the text data received by the server and the sentiment recognition results, a natural language processing (NLP) engine (e.g., a natural language processing engine) is used to analyze the user's intent. Specifically, the natural language processing engine analyzes the text data and identifies the user's instructions.
[1327] Input: Text data, sentiment recognition results
[1328] Output: User intent (e.g., "Check traffic information")
[1329] Step 5: Obtaining real-time information
[1330] Based on the user's intent, the server calls APIs such as traffic information APIs (e.g., Traffic Information API) and weather information APIs (e.g., Weather Information API) to obtain real-time data. Specifically, it sends requests to the corresponding API endpoints to obtain current traffic conditions and weather information.
[1331] Input: User intent
[1332] Output: Real-time information (e.g., "Traffic information for the current route")
[1333] Step 6: Generating the response
[1334] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results. Specifically, the generative AI considers the user's emotional state while generating an appropriate response (e.g., "Turning left at the next intersection and getting onto the highway 3 kilometers ahead will save you time").
[1335] Input: Analysis results, real-time information, emotion recognition results
[1336] Output: Response text
[1337] Step 7: Convert to audio data and output
[1338] The terminal converts the generated response text into audio data using a speech synthesis API (e.g., a speech synthesis API). Specifically, the speech synthesis engine analyzes the text data and generates an audio file. The generated audio data is then output to the user.
[1339] Input: Response text
[1340] Output: Audio data (Example: "There is traffic congestion on your current route, so you can save time by turning left at the next intersection and taking the highway 3 kilometers ahead.")
[1341] (Application Example 2)
[1342] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1343] Efficient and safe work navigation within factories is a critical issue directly linked to reducing worker stress and improving productivity. Conventional systems struggle to instantly acquire and provide real-time information on equipment status and work stations, and furthermore, they lack appropriate feedback tailored to the emotional state of workers, leading to the accumulation of stress and fatigue.
[1344] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data to determine the user's intent, means for recognizing the user's emotions, means for acquiring real-time information based on the analysis results, means for generating a response based on the acquired information, and means for converting the generated response into voice data and outputting it to the user. This enables efficient and safe work navigation within the factory, reducing worker stress and improving productivity.
[1345] "Means for receiving voice input" refers to a system element that acquires user voice commands through input devices such as microphones or head-mounted displays.
[1346] "Means for converting voice input into text data" refers to a system element that analyzes acquired voice and converts it into text data using speech recognition technology.
[1347] "Means for analyzing text data and determining user intent" refers to a system element that analyzes text data using natural language processing technology to understand and determine what the user wants.
[1348] "Means for recognizing user emotions" refers to system elements that include technologies for analyzing and recognizing a user's emotional state based on voice and text data.
[1349] "Means of acquiring real-time information" refers to elements of a system that acquires current conditions and the latest data in real time through networks and sensors.
[1350] "Means for generating a response based on acquired information" refers to a system element that includes technology for generating the optimal response to be provided to the user based on analysis results and information acquired in real time.
[1351] "Means for converting the generated response into audio data and outputting it to the user" refers to a system element that converts the generated text-based response into an audio format using speech synthesis technology and outputs it to the user.
[1352] As an example of carrying out this invention, we will take a work navigation system within a factory. The user wears a head-mounted display (HMD) and gives voice instructions to the system. The system uses the following hardware and software to provide real-time optimal work procedures and route guidance based on the user's instructions.
[1353] Hardware:
[1354] 1. Head-mounted display (HMD): It has a built-in microphone and speaker for voice input and output, and also provides visual information to the user.
[1355] 2. Server: Processes voice input from the user and performs speech recognition and response generation.
[1356] software:
[1357] 1. Speech Recognition API: For example, use Google Cloud Speech-to-Text. Convert voice input into text data.
[1358] 2. Emotion Recognition API: For example, using IBM Watson Tone Analyzer. Recognizes user emotions from text data.
[1359] 3. Natural Language Processing Engine (NLP Engine): For example, OpenAI GPT-3 is used. It analyzes text data and determines the user's intent.
[1360] 4. Real-time information acquisition method: Real-time information is acquired from sensor data within the factory and from an internal database.
[1361] 5. Generative AI: Generates appropriate responses based on the user's intent and emotion recognition results.
[1362] 6. Text-to-Speech API: For example, use Google Cloud Text-to-Speech. Convert the generated text response into speech data.
[1363] Specific processing steps:
[1364] The user inputs voice commands into the system through the HMD's built-in microphone. The HMD converts this voice input into text data, which is then sent to the server. A speech recognition API on the server converts this voice data back into text data, and then an emotion recognition API analyzes the user's emotions. The analysis results are passed to a natural language processing engine to determine the user's intent.
[1365] Next, the server acquires real-time equipment status and work station information through sensor data within the factory and an internal database. Based on the acquired real-time information and emotion recognition results, a generative AI generates an appropriate response. The generated response is converted into audio data using a speech synthesis API and output to the user through the HMD's built-in speaker. Visual information is also displayed on the HMD's screen.
[1366] Specific example:
[1367] Here's an example of a user checking the location of their next work station.
[1368] User: "Where is the next work station?"
[1369] 1. Voice recognition: "Where is the next work station?"
[1370] 2. Emotion Recognition: Calm tone (Emotion recognition result)
[1371] 3. Server analysis: Request ⇒ Search for next station information
[1372] 4. Real-time information acquisition: Obtain current location and work station status.
[1373] 5. Response generation: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[1374] 6. Speech Synthesis and Output: Output information via speech and display.
[1375] This system allows users to instantly grasp real-time factory information and the status of work stations through voice control, enabling them to work efficiently and safely.
[1376] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1377] Step 1:
[1378] The user inputs voice commands through the microphone on their head-mounted display (HMD). For example, they might say, "Where is the next work station?" This voice input becomes the system's initial input data.
[1379] Step 2:
[1380] The device (HMD) receives voice input and sends the voice data to a speech recognition API. The speech recognition API uses Google Cloud Speech-to-Text to convert the voice data into text data. The converted text data is output and used for the next process.
[1381] Step 3:
[1382] The terminal sends the converted text data to the emotion recognition API. The emotion recognition API uses IBM Watson Tone Analyzer to recognize the user's emotions from the text data. The resulting emotion recognition result is output, and this data is used for the next processing step.
[1383] Step 4:
[1384] The server receives text data and sentiment recognition results, and analyzes the data using a natural language processing engine (e.g., OpenAI GPT-3). The analysis determines the user's intent, and the result is output.
[1385] Step 5:
[1386] The server retrieves necessary real-time information according to the user's intent. It retrieves necessary information from sensor data within the factory and from internal databases, for example, obtaining information on the current equipment status and the next work station. This real-time information is output, and that data is used for the next processing.
[1387] Step 6:
[1388] Based on emotion recognition results and real-time information, the server uses generative AI to generate appropriate responses. For example, it might output the generated response text: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[1389] Step 7:
[1390] The server sends the generated response text to a text-to-speech API (e.g., Google Cloud Text-to-Speech). The text-to-speech API converts the response text into speech data, and that speech data is output.
[1391] Step 8:
[1392] The terminal outputs the generated audio data to the user via the HMD's speakers. It also displays the generated response text on the HMD's display, providing visual information simultaneously.
[1393] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1394] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1395] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1396] [Fourth Embodiment]
[1397] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1398] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1399] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1400] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1401] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1403] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1404] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1405] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1406] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1407] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1408] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1409] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1410] This invention relates to a navigation system for motorcycle riders, enabling them to obtain real-time information through voice control without using their hands, and to travel safely and efficiently. The processing and operation of the system's program are described in detail below.
[1411] Overall flow and process overview
[1412] 1. Acquisition of voice input:
[1413] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[1414] Example: "Check if the current road is congested."
[1415] 2. Speech-to-text conversion:
[1416] The device converts the received voice input into text data using a speech recognition API.
[1417] 3. Information Analysis:
[1418] The server receives text data and performs natural language processing to analyze the user's intent.
[1419] Based on the analysis results, appropriate real-time information is obtained for keywords such as "traffic congestion."
[1420] 4. Obtaining real-time information:
[1421] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[1422] The acquired traffic and weather information is integrated with the analysis results.
[1423] 5. Generating the response:
[1424] Generative AI generates responses to the user based on analysis results and acquired real-time information.
[1425] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[1426] 6. Conversion to and output of audio data:
[1427] The generative AI converts the generated response text into speech data using a speech synthesis API.
[1428] The device outputs audio data to the user.
[1429] Processing of specific examples
[1430] 1. Proposal of an alternative route
[1431] User: "Check if the current road is congested."
[1432] The terminal converts voice input into text data and sends it to the server.
[1433] The server analyzes the text data and calls a traffic information API to obtain real-time traffic information.
[1434] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[1435] The generative AI generates a response such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts it into audio data.
[1436] The device outputs audio data to the user.
[1437] 2. Providing information on recommended spots
[1438] User: "I'm hungry, can you recommend some good restaurants nearby?"
[1439] The terminal converts voice input into text data and sends it to the server.
[1440] The server analyzes the text data and calls a restaurant information API to retrieve the latest information.
[1441] The server sends the acquired restaurant information to a generative AI, which then generates a response suggesting the most suitable restaurant.
[1442] The generative AI generates the response, "There is a highly-rated Italian restaurant 1 kilometer away," and converts it into audio data.
[1443] The device outputs audio data to the user.
[1444] System Configuration
[1445] This system includes the following:
[1446] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[1447] A means of converting audio data into text data (speech recognition API).
[1448] A means of analyzing text data and determining the user's intent (natural language processing engine).
[1449] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[1450] A means of generating a response based on acquired information (generative AI).
[1451] A means of converting a response into audio data and outputting it (speech synthesis API).
[1452] These features allow users to travel safely and efficiently through voice control. This system is particularly effective when riding a motorcycle, as it can provide timely and appropriate information even when both hands are occupied.
[1453] The following describes the processing flow.
[1454] Step 1:
[1455] The user gives voice commands through a helmet with a microphone or a Bluetooth headset.
[1456] Example: "Check if the current road is congested."
[1457] Step 2:
[1458] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[1459] It calls a speech recognition API (for example, the Google Cloud Speech-to-Text API) to convert speech to text.
[1460] The converted data is temporarily stored on the device.
[1461] Step 3:
[1462] The terminal sends text data to the server.
[1463] Send an HTTP request to the server, including text data.
[1464] Step 4:
[1465] The server receives text data.
[1466] The natural language processing (NLP) engine is invoked to analyze the received text data.
[1467] Step 5:
[1468] The server analyzes the text data using a natural language processing engine to determine the user's intent.
[1469] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[1470] Step 6:
[1471] The server determines the need to acquire real-time information based on the analysis results.
[1472] The analysis results indicate a request to "check traffic congestion information."
[1473] Step 7:
[1474] The server calls a traffic information API to obtain real-time traffic information.
[1475] Example: Use the Google Maps API to retrieve current traffic information.
[1476] The acquired traffic information is temporarily stored.
[1477] Step 8:
[1478] The server integrates real-time traffic information with analysis results and sends it to the generative AI.
[1479] Based on the analysis results and traffic information, the data is sent to a generative AI engine.
[1480] Step 9:
[1481] The generative AI generates a response to the user.
[1482] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[1483] Generate response text.
[1484] Step 10:
[1485] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[1486] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[1487] Step 11:
[1488] The device receives audio data and outputs it to the user.
[1489] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[1490] Step 12:
[1491] The user listens to voice guidance and follows the suggested detour route.
[1492] Choose a safe new route and head towards your destination.
[1493] (Example 1)
[1494] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1495] It is not easy for motorcyclists to safely and efficiently obtain real-time traffic information and recommended spots without using both hands while riding. Conventional navigation systems require screen operation, which can compromise safety while driving. Therefore, there is a need for a more intuitive and safe information acquisition system that uses voice control.
[1496] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1497] In this invention, the server includes means for the user to give instructions by voice, means for converting the given voice instructions into text format, means for analyzing the text and determining the user's intent, means for acquiring real-time information based on the determined intent, a generative model for generating a detailed response based on the real-time information, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire necessary information in real time through voice operation and to move safely and efficiently.
[1498] "Voice input" refers to the act of capturing the user's voice into the system via a microphone or headset.
[1499] "Means of converting to text data" refers to technology or equipment for converting voice input into strings of characters.
[1500] "Means for determining user intent" refers to technologies or devices that analyze the content of received text data to understand the content of user requests or questions.
[1501] "Means for obtaining real-time information" refers to technologies or devices that obtain the latest data, such as current traffic conditions and weather, from appropriate sources based on analysis results.
[1502] "Means for generating responses" refers to technologies or devices that create appropriate answers or suggestions for users based on acquired real-time information.
[1503] "Means for converting into audio data and outputting to the user" refers to a technology or device that converts generated text data into audio and outputs it to the user in a format that is audible.
[1504] A "generative model" is a machine learning model that generates appropriate responses in natural language based on input prompts.
[1505] This invention relates to a voice-operated navigation system for use while riding a motorcycle, enabling the user to obtain real-time information without using both hands and to travel safely and efficiently. The following describes specific embodiments of this invention.
[1506] Hardware and software to be used
[1507] This system utilizes the following hardware and software.
[1508] 1. Hardware:
[1509] Helmet with microphone or Bluetooth headset: A device for capturing the user's voice.
[1510] Smartphone or GPS device: A device for processing audio data, acquiring real-time information, and playing back generated audio output.
[1511] 2. Software:
[1512] Speech recognition APIs (e.g., Google Cloud Speech-to-Text API): Convert voice input into text data.
[1513] Natural language processing engine (e.g., IBM Watson NLP): Analyzes text data to determine the user's intent.
[1514] Traffic information APIs and weather information APIs (e.g., Google Maps API, Weather API): Obtain real-time information.
[1515] Generative models (e.g., GPT-3): These models generate responses to the user based on the information they acquire.
[1516] Speech synthesis API (e.g., Amazon Polly): Converts the generated response into speech data.
[1517] Example of an operation flow
[1518] 1. Proposal of an alternative route
[1519] When a user uses voice input to say, "Check if the current road is congested," the system operates as follows:
[1520] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[1521] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[1522] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to check traffic information."
[1523] Server: Based on the determination result, it retrieves real-time traffic information from the Google Maps API.
[1524] Server: Passes the acquired traffic information to the generation model (GPT-3) and generates alternative route suggestions.
[1525] Generative model: Generates the response text "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[1526] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[1527] Terminal: Outputs the converted audio data to the user.
[1528] 2. Providing information on recommended spots
[1529] When a user uses voice input to say, "I'm hungry, can you recommend a good restaurant nearby?", the following process takes place.
[1530] User: Gives voice commands via a helmet with a microphone or a Bluetooth headset.
[1531] Terminal: Receives voice input and converts it into text data using the Google Cloud Speech-to-Text API.
[1532] Server: Receives the converted text data and analyzes the user's intent using IBM Watson NLP. It determines the intent is "to find recommended restaurants."
[1533] Server: Based on the determination result, retrieves information about nearby restaurants from the corresponding database (e.g., Yelp API).
[1534] Server: Passes the acquired restaurant information to the generative model (GPT-3) to generate suggestions for the best restaurants.
[1535] Generative model: Generates the response text "There is a highly-rated Italian restaurant 1 kilometer away."
[1536] Server: Sends response text to the terminal and converts it into audio data using Amazon Polly.
[1537] Terminal: Outputs the converted audio data to the user.
[1538] Effects of implementation
[1539] This allows users to obtain real-time traffic information and recommended spots via voice control without using their hands while riding a motorcycle, enabling safe and efficient travel. Furthermore, the intuitive operation improves safety while driving.
[1540] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1541] Step 1:
[1542] The user wears a helmet with a microphone or a Bluetooth headset and provides voice input. For example, they might say, "Check if this road is congested." This input is captured as analog voice data. The terminal receives this voice data and begins processing it in the appropriate format.
[1543] Step 2:
[1544] The device sends the received audio data to the Google Cloud Speech-to-Text API. The audio data is sent as input, and the speech recognition API converts it into text data. The output is text data that says, "Check if the current road is congested." This text data is prepared for further processing.
[1545] Step 3:
[1546] The terminal sends the acquired text data to the server. The server passes the text data to IBM Watson NLP, which then starts the process of analyzing the user's intent. The input is text data, and the output is the analysis result indicating the user's intent. For example, the analysis result might be "check traffic information."
[1547] Step 4:
[1548] Based on the analysis results, the server retrieves real-time traffic information from the Google Maps API. The current location and surrounding area information are input as part of the API request. The returned output includes data on current traffic conditions and congestion information. Specifically, API requests and responses are processed.
[1549] Step 5:
[1550] The server sends the acquired traffic information as a prompt to the generative model (GPT-3). The input to the generative model is a prompt statement that combines the user's intent with real-time information. The output is the response text to the user. For example, the generated text might say, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time."
[1551] Step 6:
[1552] The server sends the generated response text to the device and initiates the process of converting it into speech data using Amazon Polly. The input is the generated text, and the output is speech data. The speech synthesis API converts the text to speech, which the device receives.
[1553] Step 7:
[1554] The device plays the converted audio data to the user. The user receives voice guidance via a Bluetooth headset or helmet with a microphone, such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time." The specific action performed is the playback of the voice guidance.
[1555] (Application Example 1)
[1556] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1557] The objective of this invention is to efficiently acquire real-time inventory information and machine status within a factory using voice control, and to quickly provide optimal work instructions. Currently, factory workers have to manually check multiple data sets, which is inefficient and time-consuming.
[1558] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1559] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, and means for analyzing the text data to determine the user's intent. This makes it possible to acquire real-time inventory information and machine status within the factory and generate optimal work instructions based on this information.
[1560] "Means for receiving voice input" refers to a combination of hardware and software for collecting voice commands from the user.
[1561] "Means for converting voice input into text data" refers to technology that recognizes collected voice input and converts it into corresponding text data.
[1562] "Means for analyzing text data and determining user intent" refers to algorithms or engines that analyze text data to understand the intent behind the instructions or information requested by the user.
[1563] "Means for obtaining real-time information based on analysis results" refers to a mechanism that obtains real-time information from relevant databases and information sources according to the analyzed user's intent.
[1564] "Means for generating responses based on acquired information" refers to technology that generates appropriate responses to provide to users based on acquired real-time information.
[1565] "A means of converting generated responses into audio data and outputting it to the user" refers to a technology that converts generated text responses into audio and conveys that audio to the user.
[1566] "A means of acquiring real-time inventory information and machine status within a factory and generating optimal work instructions based on this" refers to a system that collects the latest data on inventory information and machine operating status within a factory and provides optimal instructions to workers via voice based on that data.
[1567] This invention is a system that uses voice input to acquire inventory information and machine status in real time within a factory and provides optimal work instructions to the user via voice. This system is realized by combining technologies such as speech recognition, natural language processing, real-time data acquisition, generative AI, and speech synthesis.
[1568] Specifically, when a user provides voice input through a microphone, the system converts that voice input into text data. A speech recognition API is used for this conversion. Next, the converted text data is sent to a server and analyzed by a natural language processing engine. Based on the analysis results, the server calls various APIs to obtain inventory information and machine status in real time.
[1569] Based on the acquired information, the generative AI generates an appropriate response to the user. This generated response is temporarily stored as text data and then converted into audio data using a speech synthesis API. Finally, the generated audio data is output to the user.
[1570] The system components are as follows:
[1571] Means of receiving voice input: microphone or Bluetooth headset.
[1572] A means of converting voice input into text data: Speech recognition API.
[1573] A means of analyzing text data and determining user intent: a natural language processing engine (e.g., OpenAI's API).
[1574] Means of obtaining real-time information: Inventory management APIs and APIs for monitoring machine status.
[1575] Means for converting the generated response into audio data and outputting it to the user: Generative AI and speech synthesis APIs (e.g., Google Text-to-Speech).
[1576] As a concrete example in a factory, consider a scenario where a worker gives a voice command saying, "Tell me the current inventory count." In this case, the system converts the voice to text and sends that text data to a server. The server calls an inventory management API to obtain real-time inventory information and, based on the analysis results, uses generative AI to generate an appropriate response such as "The current inventory count is 100 units." This response is then converted back into voice data and output to the worker, providing real-time information.
[1577] Similarly, in the case of a voice command such as "Tell me the status of the machine," the system can call a machine monitoring API to obtain the latest status of the machine, and then generate and provide an appropriate response to the user based on that information.
[1578] By utilizing generative AI models, it is possible to generate appropriate responses even to complex questions. For example, the following example prompt can be used:
[1579] text
[1580] User: What is the current stock quantity?
[1581] response:
[1582] As a result, this system can improve work efficiency within the factory and enable the rapid and safe provision of information.
[1583] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1584] Step 1:
[1585] The user provides voice input through a microphone. Specifically, the user gives instructions such as, "Tell me the current stock quantity." This voice input serves as the initial trigger.
[1586] Input: User's voice instructions
[1587] Output: Audio data
[1588] Operation: Uses the microphone to capture audio.
[1589] Step 2:
[1590] The device converts the acquired audio data into text data using a speech recognition API.
[1591] Input: Audio data
[1592] Output: Text data
[1593] Operation: Calls a speech recognition API (e.g., Google Speech-to-Text) to convert audio data into text.
[1594] Step 3:
[1595] The terminal sends the converted text data to the server.
[1596] Input: Text data
[1597] Output: API request to the server
[1598] Operation: Sends text data to the server using an HTTP request.
[1599] Step 4:
[1600] The server receives text data and uses a natural language processing engine to analyze the user's intent.
[1601] Input: Text data
[1602] Output: Analysis results
[1603] Operation: Analyzes text data using a natural language processing engine (e.g., OpenAI) to extract the user's intent.
[1604] Step 5:
[1605] Based on the analysis results, the server calls APIs to obtain real-time inventory information and machine status.
[1606] Input: Analysis results
[1607] Output: Real-time data
[1608] Operation: Calls inventory management APIs and machine monitoring APIs to obtain necessary information.
[1609] Step 6:
[1610] Based on the real-time data it acquires, the server uses generative AI to generate appropriate responses for the user.
[1611] Input: Real-time data
[1612] Output: Response text
[1613] Operation: Uses generative AI (e.g., OpenAI) to generate response text based on real-time data.
[1614] Step 7:
[1615] The server calls a speech synthesis API to convert the generated response text into speech data.
[1616] Input: Response text
[1617] Output: Audio data
[1618] Operation: Converts text into speech data using a speech synthesis API (e.g., Google Text-to-Speech).
[1619] Step 8:
[1620] The terminal outputs the generated audio data to the user.
[1621] Input: Audio data
[1622] Output: Voice notification to the user
[1623] Operation: Plays audio data to the user using the speaker or headset connected to the device.
[1624] The system is designed to ensure that each step smoothly executes the entire process, from acquiring voice input and analyzing the data to obtaining real-time information and generating and converting appropriate responses into speech.
[1625] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1626] This invention is a system that allows motorcyclists to obtain real-time navigation information using voice commands while riding, enabling them to reach their destination safely and efficiently. In particular, this invention achieves more appropriate information provision and navigation by combining it with an emotion engine that recognizes the user's emotions.
[1627] Overall flow and process overview
[1628] 1. Acquisition of voice input:
[1629] Users give voice commands to the system using a helmet with a microphone or a Bluetooth headset.
[1630] Example: "Check if the current road is congested."
[1631] 2. Speech-to-text conversion:
[1632] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[1633] 3. Recognition of emotions:
[1634] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[1635] Use an emotion recognition API (such as IBM Watson Tone Analyzer) to identify the user's emotions from their voice.
[1636] 4. Information Analysis:
[1637] The server receives text data and uses a natural language processing (NLP) engine to analyze the user's intent.
[1638] Based on text data and sentiment recognition results, the system determines the user's needs and emotional state.
[1639] 5. Obtaining real-time information:
[1640] The server retrieves real-time data from APIs such as traffic information APIs and weather information APIs.
[1641] Example: Obtain current traffic and weather information.
[1642] 6. Generating the response:
[1643] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[1644] Example: "Due to current traffic congestion, we suggest an alternative route. Turning left at the next intersection and entering the highway 3 kilometers ahead will save you time."
[1645] 7. Conversion to and output of audio data:
[1646] The response text generated by the generative AI is converted into speech data using a speech synthesis API.
[1647] The device outputs audio data to the user.
[1648] Processing of specific examples
[1649] 1. Proposal of an alternative route
[1650] User: "Check if the current road is congested."
[1651] The device converts voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[1652] The server analyzes text data and emotion recognition results, and then calls a traffic information API to obtain real-time traffic information.
[1653] The server sends analysis results based on traffic congestion information to a generative AI, which then calculates an alternative route.
[1654] The generative AI takes emotion recognition results into account and generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[1655] The device outputs audio data to the user.
[1656] 2. Providing information on recommended spots
[1657] User: "I'm hungry, can you recommend some good restaurants nearby?"
[1658] The device converts voice input into text data and analyzes the user's emotional state through an emotion recognition API.
[1659] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[1660] Based on the emotion recognition results, the generative AI generates a response such as "There is a highly-rated Italian restaurant 1 kilometer away" for users in a relaxed state, and converts it into audio data.
[1661] The device outputs audio data to the user.
[1662] System Configuration
[1663] This system includes the following:
[1664] Means of acquiring user voice (e.g., helmets with microphones or Bluetooth headsets).
[1665] A means of converting audio data into text data (speech recognition API).
[1666] A means of analyzing text data and determining the user's intent (natural language processing engine).
[1667] A means of recognizing a user's emotions (emotion recognition API).
[1668] Means of obtaining real-time information (traffic information APIs and weather information APIs).
[1669] A means of generating responses based on acquired information and emotion recognition results (generative AI).
[1670] A means of converting a response into audio data and outputting it (speech synthesis API).
[1671] These features allow users to receive appropriate information and navigation based on their emotional state, enabling safe and efficient travel even while riding a motorcycle. This system, in particular, significantly improves the user experience and supports stress-free travel by combining it with an emotion engine.
[1672] The following describes the processing flow.
[1673] Step 1:
[1674] The user gives voice commands to the system via a helmet with a microphone or a Bluetooth headset.
[1675] Example: "Check if the current road is congested."
[1676] Step 2:
[1677] The device receives voice input and uses a speech recognition API to convert the speech into text data.
[1678] This converts speech to text by calling a speech recognition API (for example, the Google Cloud Speech-to-Text API).
[1679] The converted text data is temporarily stored on the device.
[1680] Step 3:
[1681] The device sends the acquired audio data to an emotion recognition API to analyze the user's emotions.
[1682] An emotion recognition API (for example, IBM Watson Tone Analyzer) is called to identify the user's emotions from their voice.
[1683] The analysis results (e.g., stress, anxiety, relaxation, etc.) are temporarily stored on the device.
[1684] Step 4:
[1685] The device sends text data and emotion recognition results to the server.
[1686] An HTTP request is sent to the server, including text data and sentiment recognition results.
[1687] Step 5:
[1688] The server receives text data and emotion recognition results.
[1689] The received data is passed to the natural language processing (NLP) engine.
[1690] Step 6:
[1691] The server uses a natural language processing engine to analyze text data and determine the user's intent.
[1692] The system analyzes keywords and context to identify user requests (e.g., checking traffic information).
[1693] Step 7:
[1694] The server determines the need for real-time information acquisition based on the analysis results and emotion recognition results.
[1695] Consider the case where the analysis results request "checking traffic information" and the emotion recognition result indicates "stress."
[1696] Step 8:
[1697] The server calls a traffic information API to obtain real-time traffic information.
[1698] Example: Use the Google Maps API to retrieve current traffic information.
[1699] The acquired traffic information is temporarily stored.
[1700] Step 9:
[1701] The server integrates real-time traffic information acquired by the system with analysis results and sends it to the generative AI.
[1702] Based on the analysis results, traffic information, and emotion recognition results, the data is sent to a generative AI engine.
[1703] Step 10:
[1704] The generative AI generates a response to the user.
[1705] For example, considering traffic information and the user's stress level, it can generate responses such as, "There is currently traffic congestion, so we suggest a relaxing detour. Turn left at the next intersection and take the scenic route 3 kilometers ahead."
[1706] Step 11:
[1707] The speech synthesis API is called to convert the response text generated by the generative AI into speech data.
[1708] Example: Use the Google Cloud Text-to-Speech API to convert text into speech data.
[1709] Step 12:
[1710] The device receives audio data and outputs it to the user.
[1711] Voice guidance is played through the helmet's speakers or a Bluetooth headset.
[1712] Step 13:
[1713] The user listens to voice guidance and follows the suggested detour route.
[1714] Choose a safe new route and head towards your destination.
[1715] (Example 2)
[1716] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1717] Conventional navigation systems made it difficult for motorcyclists to operate them using voice commands, posing a risk of reduced safety. Furthermore, they lacked the ability to provide information tailored to the user's emotional state, resulting in a limited user experience. Additionally, the difficulty in providing real-time traffic information and recommended spots meant that efficient travel was challenging.
[1718] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1719] In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data and determining the user's intent, means for recognizing the user's emotions from the text data, means for acquiring real-time information based on the analysis results and emotion recognition, means for generating a response based on the acquired information and emotion recognition, and means for converting the generated response into voice data and outputting it to the user. This enables the user to acquire navigation information by voice operation even while riding a motorcycle, allowing for safe and efficient travel. Furthermore, it enables the provision of appropriate information according to the user's emotional state, significantly improving the user experience and realizing stress-free navigation.
[1720] 1. "Voice input" refers to instructions or data provided by the user via voice.
[1721] 2. "Text data" refers to digital information obtained by converting audio into text format.
[1722] 3. "Analysis" is the process of analyzing data to determine specific intentions or emotions.
[1723] 4. "User intent" refers to the actions a user wants to take or the information they want to obtain.
[1724] 5. "Emotion recognition" is a method of determining a user's emotional state from voice data.
[1725] 6. "Real-time information" refers to information that provides the latest data based on the current situation.
[1726] 7. A "response" is the answer that the system generates in response to user input.
[1727] 8. "Conversion to audio data" refers to the process of converting text data into an audio format.
[1728] 9. "Traffic information" refers to real-time traffic-related data such as current road conditions, congestion information, and accident information.
[1729] 10. A "detour route" is a proposed alternative route to avoid congestion or obstacles.
[1730] 11. "Recommended Spot Information" refers to information about places and facilities near the destination that are useful to the user.
[1731] 12. "Various information sources" refers to different information providers and systems used to collect data.
[1732] 13. "Generative AI" refers to artificial intelligence technology that generates responses to users based on analysis results and emotion recognition.
[1733] This invention provides a system that allows motorcycle riders to obtain real-time navigation information using voice commands and reach their destination safely and efficiently. This system is implemented by combining voice input, speech recognition, emotion recognition, natural language processing, real-time information acquisition, response generation, and speech synthesis.
[1734] Hardware and software configuration to be used
[1735] 1. User input means:
[1736] Users give voice commands using a helmet with a microphone or a Bluetooth headset.
[1737] 2. Voice recognition means:
[1738] The device receives voice input and uses a speech recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.
[1739] 3. Emotion recognition means:
[1740] The device sends the acquired text data to an emotion recognition API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions.
[1741] 4. Natural language processing tools:
[1742] The server receives text data and sentiment recognition results, uses a natural language processing (NLP) engine (e.g., Google Natural Language API) to analyze the user's intent, and determines the user's needs and emotional state.
[1743] 5. Means of acquiring real-time information:
[1744] The server retrieves real-time data from APIs such as traffic information APIs (e.g., Google Maps API) and weather information APIs (e.g., OpenWeatherMap).
[1745] 6. Response generation means:
[1746] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results.
[1747] 7. Speech synthesis means:
[1748] The response text generated by the generative AI is converted into speech data using a speech synthesis API (e.g., Amazon Polly).
[1749] The device outputs the generated audio data to the user.
[1750] Specific example behavior
[1751] Proposed detour route
[1752] 1. User: "Please check if the current road is congested."
[1753] The device converts voice input into text data and analyzes the emotion "irritated" through an emotion recognition API.
[1754] The server analyzes text data and sentiment recognition results, and uses a traffic information API to obtain real-time traffic congestion information.
[1755] The generative AI generates responses such as, "Turn left at the next intersection and get onto the highway 3 kilometers ahead to save time," and converts them into audio data.
[1756] The device outputs audio data to the user.
[1757] Providing information on recommended spots
[1758] 1. User: "I'm hungry, can you recommend some good restaurants nearby?"
[1759] The device converts voice input into text data and analyzes the emotion of "being relaxed" through an emotion recognition API.
[1760] The server analyzes the text data and emotion recognition results, then calls a restaurant information API to retrieve the latest information.
[1761] The generative AI generates responses such as "There is a highly-rated Italian restaurant 1 kilometer away" and converts them into audio data.
[1762] The device outputs audio data to the user.
[1763] Example of a prompt
[1764] Prompt message for suggesting a detour route
[1765] Generate a response to the user's request, "Check if the current road is congested." The user's emotional state is "frustrated." Based on the current traffic information, suggest an alternative route.
[1766] Prompt message for providing recommended spot information
[1767] Generate a response to a user who asks, "I'm hungry, can you recommend a nearby restaurant?" The user's emotional state is "relaxed." Based on the latest restaurant information, recommend a suitable restaurant.
[1768] As described above, the present invention allows users to obtain real-time navigation information and recommended spot information by voice control even while riding a motorcycle. Furthermore, the introduction of an emotion engine enables the provision of appropriate information according to the user's emotional state, thereby providing a safe and comfortable navigation experience.
[1769] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1770] Step 1: Obtaining voice input
[1771] The user gives voice commands using a helmet with a microphone or a Bluetooth headset. The user can issue voice commands such as, "Check if this road is congested."
[1772] Input: User voice commands
[1773] Output: Audio data
[1774] Step 2: Convert speech to text
[1775] The device uses a speech recognition API (e.g., a speech recognition API) to convert the input speech data into text data.
[1776] Specifically, the speech recognition engine analyzes the audio data and generates corresponding text data.
[1777] Input: Audio data
[1778] Output: Text data (Example: "Check if the current road is congested")
[1779] Step 3: Recognizing Emotions
[1780] The device sends the acquired text data to an emotion recognition API (e.g., an emotion recognition API), which analyzes the user's emotions from the text. Specifically, the emotion recognition engine analyzes the text data and identifies emotional states such as "irritated."
[1781] Input: Text data
[1782] Output: Emotion recognition result (e.g., "irritated")
[1783] Step 4: Information Analysis
[1784] Based on the text data received by the server and the sentiment recognition results, a natural language processing (NLP) engine (e.g., a natural language processing engine) is used to analyze the user's intent. Specifically, the natural language processing engine analyzes the text data and identifies the user's instructions.
[1785] Input: Text data, sentiment recognition results
[1786] Output: User intent (e.g., "Check traffic information")
[1787] Step 5: Obtaining real-time information
[1788] Based on the user's intent, the server calls APIs such as traffic information APIs (e.g., Traffic Information API) and weather information APIs (e.g., Weather Information API) to obtain real-time data. Specifically, it sends requests to the corresponding API endpoints to obtain current traffic conditions and weather information.
[1789] Input: User intent
[1790] Output: Real-time information (e.g., "Traffic information for the current route")
[1791] Step 6: Generating the response
[1792] The generative AI generates a response to the user based on the analysis results, acquired real-time information, and emotion recognition results. Specifically, the generative AI considers the user's emotional state while generating an appropriate response (e.g., "Turning left at the next intersection and getting onto the highway 3 kilometers ahead will save you time").
[1793] Input: Analysis results, real-time information, emotion recognition results
[1794] Output: Response text
[1795] Step 7: Convert to audio data and output
[1796] The terminal converts the generated response text into audio data using a speech synthesis API (e.g., a speech synthesis API). Specifically, the speech synthesis engine analyzes the text data and generates an audio file. The generated audio data is then output to the user.
[1797] Input: Response text
[1798] Output: Audio data (Example: "There is traffic congestion on your current route, so you can save time by turning left at the next intersection and taking the highway 3 kilometers ahead.")
[1799] (Application Example 2)
[1800] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1801] Efficient and safe work navigation within factories is a critical issue directly linked to reducing worker stress and improving productivity. Conventional systems struggle to instantly acquire and provide real-time information on equipment status and work stations, and furthermore, they lack appropriate feedback tailored to the emotional state of workers, leading to the accumulation of stress and fatigue.
[1802] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting voice input into text data, means for analyzing the text data to determine the user's intent, means for recognizing the user's emotions, means for acquiring real-time information based on the analysis results, means for generating a response based on the acquired information, and means for converting the generated response into voice data and outputting it to the user. This enables efficient and safe work navigation within the factory, reducing worker stress and improving productivity.
[1803] "Means for receiving voice input" refers to a system element that acquires user voice commands through input devices such as microphones or head-mounted displays.
[1804] "Means for converting voice input into text data" refers to a system element that analyzes acquired voice and converts it into text data using speech recognition technology.
[1805] "Means for analyzing text data and determining user intent" refers to a system element that analyzes text data using natural language processing technology to understand and determine what the user wants.
[1806] "Means for recognizing user emotions" refers to system elements that include technologies for analyzing and recognizing a user's emotional state based on voice and text data.
[1807] "Means of acquiring real-time information" refers to elements of a system that acquires current conditions and the latest data in real time through networks and sensors.
[1808] "Means for generating a response based on acquired information" refers to a system element that includes technology for generating the optimal response to be provided to the user based on analysis results and information acquired in real time.
[1809] "Means for converting the generated response into audio data and outputting it to the user" refers to a system element that converts the generated text-based response into an audio format using speech synthesis technology and outputs it to the user.
[1810] As an example of carrying out this invention, we will take a work navigation system within a factory. The user wears a head-mounted display (HMD) and gives voice instructions to the system. The system uses the following hardware and software to provide real-time optimal work procedures and route guidance based on the user's instructions.
[1811] Hardware:
[1812] 1. Head-mounted display (HMD): It has a built-in microphone and speaker for voice input and output, and also provides visual information to the user.
[1813] 2. Server: Processes voice input from the user and performs speech recognition and response generation.
[1814] software:
[1815] 1. Speech Recognition API: For example, use Google Cloud Speech-to-Text. Convert voice input into text data.
[1816] 2. Emotion Recognition API: For example, using IBM Watson Tone Analyzer. Recognizes user emotions from text data.
[1817] 3. Natural Language Processing Engine (NLP Engine): For example, OpenAI GPT-3 is used. It analyzes text data and determines the user's intent.
[1818] 4. Real-time information acquisition method: Real-time information is acquired from sensor data within the factory and from an internal database.
[1819] 5. Generative AI: Generates appropriate responses based on the user's intent and emotion recognition results.
[1820] 6. Text-to-Speech API: For example, use Google Cloud Text-to-Speech. Convert the generated text response into speech data.
[1821] Specific processing steps:
[1822] The user inputs voice commands into the system through the HMD's built-in microphone. The HMD converts this voice input into text data, which is then sent to the server. A speech recognition API on the server converts this voice data back into text data, and then an emotion recognition API analyzes the user's emotions. The analysis results are passed to a natural language processing engine to determine the user's intent.
[1823] Next, the server acquires real-time equipment status and work station information through sensor data within the factory and an internal database. Based on the acquired real-time information and emotion recognition results, a generative AI generates an appropriate response. The generated response is converted into audio data using a speech synthesis API and output to the user through the HMD's built-in speaker. Visual information is also displayed on the HMD's screen.
[1824] Specific example:
[1825] Here's an example of a user checking the location of their next work station.
[1826] User: "Where is the next work station?"
[1827] 1. Voice recognition: "Where is the next work station?"
[1828] 2. Emotion Recognition: Calm tone (Emotion recognition result)
[1829] 3. Server analysis: Request ⇒ Search for next station information
[1830] 4. Real-time information acquisition: Obtain current location and work station status.
[1831] 5. Response generation: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[1832] 6. Speech Synthesis and Output: Output information via speech and display.
[1833] This system allows users to instantly grasp real-time factory information and the status of work stations through voice control, enabling them to work efficiently and safely.
[1834] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1835] Step 1:
[1836] The user inputs voice commands through the microphone on their head-mounted display (HMD). For example, they might say, "Where is the next work station?" This voice input becomes the system's initial input data.
[1837] Step 2:
[1838] The device (HMD) receives voice input and sends the voice data to a speech recognition API. The speech recognition API uses Google Cloud Speech-to-Text to convert the voice data into text data. The converted text data is output and used for the next process.
[1839] Step 3:
[1840] The terminal sends the converted text data to the emotion recognition API. The emotion recognition API uses IBM Watson Tone Analyzer to recognize the user's emotions from the text data. The resulting emotion recognition result is output, and this data is used for the next processing step.
[1841] Step 4:
[1842] The server receives text data and sentiment recognition results, and analyzes the data using a natural language processing engine (e.g., OpenAI GPT-3). The analysis determines the user's intent, and the result is output.
[1843] Step 5:
[1844] The server retrieves necessary real-time information according to the user's intent. It retrieves necessary information from sensor data within the factory and from internal databases, for example, obtaining information on the current equipment status and the next work station. This real-time information is output, and that data is used for the next processing.
[1845] Step 6:
[1846] Based on emotion recognition results and real-time information, the server uses generative AI to generate appropriate responses. For example, it might output the generated response text: "The next work station is 5 meters away. The equipment is functioning normally, and the next task is to install a part."
[1847] Step 7:
[1848] The server sends the generated response text to a text-to-speech API (e.g., Google Cloud Text-to-Speech). The text-to-speech API converts the response text into speech data, and that speech data is output.
[1849] Step 8:
[1850] The terminal outputs the generated audio data to the user via the HMD's speakers. It also displays the generated response text on the HMD's display, providing visual information simultaneously.
[1851] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1852] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1853] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1854] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1855] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1856] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1857] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1858] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1859] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1860] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1861] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1862] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1863] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1864] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1865] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1866] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1867] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1868] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1869] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1870] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1871] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1872] The following is further disclosed regarding the embodiments described above.
[1873] (Claim 1)
[1874] A means of receiving voice input,
[1875] A means of converting voice input into text data,
[1876] A means of analyzing text data and determining the user's intent,
[1877] A means of obtaining real-time information based on the analysis results,
[1878] Means for generating a response based on acquired information,
[1879] A means of converting the generated response into audio data and outputting it to the user,
[1880] A system that includes this.
[1881] (Claim 2)
[1882] The system according to claim 1, which acquires real-time traffic information and proposes an alternative route based on the traffic congestion information.
[1883] (Claim 3)
[1884] The system according to claim 1, which acquires data from various sources and notifies the user of the most suitable spot information by voice based on the analysis results, in order to provide information on recommended spots around the destination.
[1885] "Example 1"
[1886] (Claim 1)
[1887] A means of receiving voice input,
[1888] A means of converting voice input into text data,
[1889] A means of analyzing text data and determining the user's intent,
[1890] A means of obtaining real-time information based on the analysis results,
[1891] Means for generating a response based on acquired information,
[1892] A means of converting the generated response into audio data and outputting it to the user,
[1893] A system that includes this.
[1894] (Claim 2)
[1895] The system according to claim 1, which acquires real-time traffic information and proposes an alternative route based on the traffic congestion information.
[1896] (Claim 3)
[1897] The system according to claim 1, which acquires data from various sources and notifies the user of the most suitable spot information by voice based on the analysis results, in order to provide information on recommended spots around the destination.
[1898] (Claim 4)
[1899] A means for the user to give instructions by voice,
[1900] A means of converting the instructed audio into text format,
[1901] A means of analyzing text and determining the user's intent,
[1902] A means of obtaining real-time information based on the identified intent,
[1903] A generative model that generates detailed responses based on real-time information,
[1904] A means of converting the generated response into audio data and outputting it to the user,
[1905] A system that includes this.
[1906] (Claim 5)
[1907] The system according to claim 4, which generates a response including real-time information obtained using a generative model and outputs it as audio data.
[1908] "Application Example 1"
[1909] (Claim 1)
[1910] A means of receiving voice input,
[1911] A means of converting voice input into text data,
[1912] A means of analyzing text data and determining the user's intent,
[1913] A means of obtaining real-time information based on the analysis results,
[1914] Means for generating a response based on acquired information,
[1915] A means of converting the generated response into audio data and outputting it to the user,
[1916] A means of receiving voice commands from the user, acquiring real-time inventory information and machine status within the factory, and generating the optimal response,
[1917] A system that includes this.
[1918] (Claim 2)
[1919] The system according to claim 1, which acquires real-time traffic information and proposes an alternative route based on the traffic congestion information.
[1920] (Claim 3)
[1921] The system according to claim 1, which acquires data from various sources and notifies the user of the most suitable spot information by voice based on the analysis results, in order to provide information on recommended spots around the destination.
[1922] (Claim 4)
[1923] The system according to claim 1, which acquires real-time inventory information and machine status within a factory and generates optimal work instructions based on this information.
[1924] "Example 2 of combining an emotion engine"
[1925] (Claim 1)
[1926] A means of receiving voice input,
[1927] A means of converting voice input into text data,
[1928] A means of analyzing text data and determining the user's intent,
[1929] A means of recognizing user emotions from text data,
[1930] A means of acquiring real-time information based on analysis results and emotion recognition,
[1931] Means for generating a response based on acquired information and emotion recognition,
[1932] A means of converting the generated response into audio data and outputting it to the user,
[1933] A system that includes this.
[1934] (Claim 2)
[1935] The system according to claim 1, which acquires real-time traffic information and proposes an alternative route based on the traffic congestion information.
[1936] (Claim 3)
[1937] The system according to claim 1, which acquires data from various sources to provide recommended spots around a destination, and notifies the user of the most suitable spot information by voice based on the analysis results and emotion recognition.
[1938] "Application example 2 when combining with an emotional engine"
[1939] (Claim 1)
[1940] A means of receiving voice input,
[1941] A means of converting voice input into text data,
[1942] A means of analyzing text data and determining the user's intent,
[1943] Means of recognizing user emotions,
[1944] A means of obtaining real-time information based on the analysis results,
[1945] Means for generating a response based on acquired information,
[1946] A means of converting the generated response into audio data and outputting it to the user,
[1947] A system that includes this.
[1948] (Claim 2)
[1949] The system according to claim 1, which acquires real-time factory information and proposes optimal work procedures and routes based on the status of equipment and information from work stations.
[1950] (Claim 3)
[1951] The system according to claim 1, which generates a response based on the emotion recognition result and notifies the user by voice. [Explanation of Symbols]
[1952] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving voice input, A means of converting voice input into text data, A means of analyzing text data and determining the user's intent, A means of obtaining real-time information based on the analysis results, Means for generating a response based on acquired information, A means of converting the generated response into audio data and outputting it to the user, A system that includes this.
2. The system according to claim 1, which acquires real-time traffic information and proposes an alternative route based on the traffic congestion information.
3. The system according to claim 1, which acquires data from various sources and notifies the user of the most suitable spot information by voice based on the analysis results, in order to provide information on recommended spots around the destination.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A