System

A voice-activated system for motorcycles allows hands-free operation of navigation, calls, and weather checks, enhancing safety and convenience by converting voice input to text, analyzing, and providing audio feedback.

JP2026022449APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123966
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Riding a motorcycle poses challenges in safely and efficiently performing operations like navigation, answering calls, and checking weather conditions due to the difficulty in using physical controls, which reduces rider safety and convenience.

Method used

A system that receives voice input, converts it to text, analyzes the text to determine appropriate actions, sends requests to a server, and provides audio feedback, enabling hands-free operation for functions such as navigation, music playback, phone call answering, and weather information retrieval.

Benefits of technology

Enables safe and efficient hands-free operation of various functions while riding a motorcycle, improving safety and convenience by allowing riders to perform tasks without taking their hands off the handlebars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022449000001_ABST
    Figure 2026022449000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for receiving speech input; means for converting the received speech input to text; means for parsing the converted text to determine an appropriate action; means for sending a request to a server based on the determined action; and means for receiving a response from the server and providing feedback to a user as speech.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] While riding a motorcycle, physical controls were difficult and dangerous, making it difficult to check navigation and answer phone calls, and there were limited ways to easily obtain information to quickly respond to changing weather conditions. These conditions reduced rider safety and convenience. [Means for solving the problem]

[0005] The present invention provides a system including a means for receiving voice input, a means for converting the received voice input into text, a means for analyzing the converted text to determine an appropriate action, a means for sending a request to a server based on the determined action, and a means for receiving a response from the server and providing audio feedback to the user. This allows the rider to safely and efficiently perform various operations hands-free. Furthermore, by including a means for analyzing the generated text and determining at least one action from among navigation, music playback, phone call answering, weather information, and emergency communication, the system achieves even more multifunctional operation. Furthermore, by including a means for utilizing a speech recognition service to convert the voice input into text, sending an HTTP request to a server based on the analysis result, and outputting the response from the server via a voice synthesizer, the system enables real-time information provision and operation.

[0006] "Voice input" is a means of capturing voice information uttered by a user using a microphone or the like of a device.

[0007] "Text conversion" is a method of converting captured voice input into textual information using voice recognition technology.

[0008] "Analysis" is the process of understanding the content of the converted text information and determining the appropriate action to take.

[0009] "Appropriate actions" are operations such as navigation, music playback, answering calls, checking weather information, and emergency communication, which are determined based on the analysis results.

[0010] A "server" is a central processing unit that receives requests and provides information or performs operations according to the requested action.

[0011] A "request" is a command or request sent to a server based on the parsed text information.

[0012] A "response" is information or a message that a server generates in response to a request and returns to a terminal.

[0013] "Voice feedback" is a means of converting a response from a server into a voice output and providing it to the user.

[0014] A "voice recognition service" is software or cloud-based service that converts voice input into text.

[0015] A "speech synthesizer" is hardware or software that converts text information into speech and outputs it.

[0016] "Communication means" refers to a network interface and communication protocol that enable data communication between a terminal and a server. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention relates to a system that allows a rider to perform various operations hands-free while riding a motorcycle. Specifically, it can realize functions such as navigation, music playback, phone call answering, weather information confirmation, and emergency communication using voice input. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0039] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as an audio file. The audio file is then converted into text data using a voice recognition service. The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Tokyo," for example, requires a navigation function.

[0040] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0041] The response from the server is received by the device and fed back to the user as a voice message that is easy for the rider to understand. This voice feedback is output by converting the text message into voice using a voice synthesizer built into the device. This allows the rider to obtain necessary information and perform operations using only voice, without having to take their hands off the handlebars.

[0042] For example, if a user says "Navigate to Tokyo" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Tokyo." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Tokyo and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Tokyo."

[0043] This system is designed to be safe and efficient to operate not only while driving a car, but also while driving a motorcycle, etc. By combining functions such as voice recognition, analysis, communication, and voice feedback, a wide range of operations can be performed hands-free, greatly improving the safety and convenience of riders.

[0044] The processing flow will be explained below.

[0045] Step 1:

[0046] The user issues a voice command such as "Navigate to Tokyo," and the voice is captured by a microphone attached to the device.

[0047] Step 2:

[0048] The device records the captured audio as a digital audio file, at which point the audio data is temporarily stored.

[0049] Step 3:

[0050] The device sends the recorded voice data to a voice recognition service, converts the voice data into text data, and receives the text data returned by the voice recognition service, for example, using a cloud-based voice recognition API.

[0051] Step 4:

[0052] The terminal analyzes the received text data and extracts the command content. For example, in the command "Navigate to Tokyo," the analysis means determines that "Tokyo" is the destination.

[0053] Step 5:

[0054] The terminal determines an appropriate action based on the analysis result. In this case, since "navigate" is requested, it determines to process it as a navigation action.

[0055] Step 6:

[0056] Based on the determined action, the device sends an HTTP request to the server, including necessary parameters such as destination information.

[0057] Step 7:

[0058] The server receives the HTTP request and performs the appropriate action based on the request, for example, generating route information to the destination.

[0059] Step 8:

[0060] The server sends the generated information back to the device as a response, which includes detailed route information for navigation.

[0061] Step 9:

[0062] The terminal receives the response from the server and analyzes the content in the form of a text message.

[0063] Step 10:

[0064] The device's speech synthesizer converts the text message into speech and provides feedback to the rider, such as a voice message saying, "Navigation to Tokyo is about to begin."

[0065] Step 11:

[0066] By receiving voice feedback from the device, users can obtain navigation information without taking their hands off the steering wheel.

[0067] Through the above processing steps, the present invention enables a rider to safely and efficiently use voice commands while riding a motorcycle.

[0068] Example 1

[0069] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0070] Currently, using devices such as smartphones to check navigation, play music, or make calls while riding a motorcycle is extremely dangerous. There is a demand for hands-free systems that allow these operations to be performed without using the hands while driving. There is also a need to develop systems that simultaneously improve safety and convenience.

[0071] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0072] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and determining an appropriate action, means for sending a request to an information processing device based on the determined action, means for receiving a response from the information processing device and feeding it back to the user as voice, means for using a voice recognition service, means for analyzing text using a natural language processing library, means for sending an HTTP request, and means for converting into voice using a voice synthesis device.

[0073] This allows users to navigate, play music, answer calls, check weather information, make emergency calls, and perform other operations using only voice input while riding a motorcycle, without using their hands.

[0074] "Voice input" refers to the means by which a user gives instructions to a system using voice.

[0075] "Speech Recognition Service" means external or internal software or functionality used to convert received speech into text data.

[0076] "Text conversion" refers to the process of analyzing speech input and converting it into corresponding text data.

[0077] A "natural language processing library" is a software library for analyzing text data and understanding meaning and actions.

[0078] An "HTTP request" is a communication method using a protocol for a client to request data from a server.

[0079] An "information processing device" refers to a computer or server that processes data based on a received request and generates a response.

[0080] A "voice synthesis device" is a device or software that converts text data into voice data and outputs it.

[0081] "Communication means" refers to the technology or protocol for transferring data, including wired and wireless communication methods.

[0082] "Feedback" refers to the system returning processing results or information to the user, and in this context it primarily refers to audio feedback.

[0083] "Location information provision" is a function that provides information about the user's specified destination or current location.

[0084] "Media content control" is a function for operating media content such as music playback and video playback.

[0085] "Communication response" is a function that responds to communication-related commands, such as incoming and outgoing phone calls.

[0086] "Weather information acquisition" refers to a function of acquiring weather information such as the current weather and forecast and providing it to the user.

[0087] "Safety communication" is a function for quickly communicating necessary information in an emergency.

[0088] The present invention relates to a system that allows riders to perform various operations hands-free while riding a motorcycle. This system uses voice input to realize functions such as navigation, music playback, answering calls, checking weather information, and emergency communication. The user can operate the terminal using voice input, and the terminal provides the necessary information and services through communication with a server.

[0089] The present invention consists of the following main components:

[0090] 1. Voice input means: A means for users to input commands by voice. The device is equipped with a high-sensitivity microphone that accurately captures the user's voice. Specific hardware used for this is a Bluetooth-enabled headset or smartphone.

[0091] 2. Speech recognition means: This is a means for converting received voice input into text, and uses speech recognition services such as IBM Watson and Google Cloud Speech-to-Text. These services are cloud-based and enable highly accurate speech recognition.

[0092] 3. Text analysis: This is a means of analyzing the converted text data and determining the appropriate action. The device uses natural language processing libraries such as spaCy and NLTK to analyze the text and understand the user's intent.

[0093] 4. Communication method: This is a method for sending an HTTP request to the server based on the determined action. The device uses this method to communicate the analysis results to the server. Communication is secure using HTTPS.

[0094] 5. Response processing means: This is the means for receiving responses from the server and providing audio feedback to the user. After receiving data from the server, the device uses Amazon Polly or Google Text-to-Speech to convert the text data into audio and provide it to the user.

[0095] For example, if a user says "Navigate to Tokyo" while driving, the following process will occur:

[0096] The device captures the user's voice and sends the voice data to a cloud service.

[0097] The voice recognition service converts the voice data into text data such as "Navigate to Tokyo."

[0098] The terminal analyzes this text data and determines that it is a navigation request.

[0099] The device creates an HTTP request and sends the destination information to the server.

[0100] The server generates route information to Tokyo using the Google Maps API or similar and sends that information back to the device.

[0101] The device analyzes the route information it receives and uses a voice synthesizer to provide voice feedback to the user, saying, "Navigation to Tokyo is about to begin."

[0102] An example prompt is:

[0103] "Describe a system that enables navigation functions based on voice input. For example, capture voice input of "Navigate to Tokyo" and demonstrate the flow of speech recognition, natural language processing, sending an HTTP request, processing on the server, and voice feedback."

[0104] This system allows riders to perform many operations while riding their motorcycle using only voice commands, without using their hands, greatly improving safety and convenience.The system can also be used in other driving situations and with a variety of voice commands, making it suitable for a wide range of uses.

[0105] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0106] Process flow steps and explanations

[0107] Step 1:

[0108] The terminal captures the user's voice.

[0109] Input: User's voice command (e.g., "Navigate to Tokyo")

[0110] How it works: The device's high-sensitivity microphone picks up the user's voice.

[0111] Output: Audio data

[0112] Step 2:

[0113] The device sends the captured audio to a speech recognition service.

[0114] Input: Audio data

[0115] How it works: Your device sends voice data to a cloud-based speech recognition service (e.g., IBM Watson, Google Cloud Speech-to-Text).

[0116] Output: Text data (e.g. "Navigate to Tokyo")

[0117] Step 3:

[0118] The device analyzes the received text data using a natural language processing library.

[0119] Input: Text data (e.g., "Navigate to Tokyo")

[0120] How it works: The terminal uses natural language processing libraries such as spaCy or NLTK to parse the text data and determine the intent of the command.

[0121] Output: Analysis results (e.g., navigation requests)

[0122] Step 4:

[0123] The terminal constructs an HTTP request based on the analysis results and sends it to the server.

[0124] Input: Analysis result (e.g., navigation request, destination "Tokyo")

[0125] How it works: The device constructs an HTTP request and sends it to the server, including the analysis results and additional parameters (e.g., current location, departure time).

[0126] Output: HTTP request to the server

[0127] Step 5:

[0128] The server processes the received HTTP request and retrieves the required data.

[0129] Input: HTTP request (e.g., destination "Tokyo")

[0130] How it works: The server parses the request and retrieves route information using external services such as the Google Maps API or OpenStreetMap.

[0131] Output: Route information

[0132] Step 6:

[0133] The server generates a response containing the route information and sends it back to the terminal.

[0134] Input: Route information

[0135] Operation: The server generates an HTTP response based on the obtained route information and sends it to the terminal.

[0136] Output: HTTP response to the device

[0137] Step 7:

[0138] The terminal receives the response from the server and converts it into speech using a speech synthesizer.

[0139] Input: HTTP response (e.g. route information)

[0140] How it works: The device uses Amazon Polly or Google Text-to-Speech to convert text data (route guidance) into voice data.

[0141] Output: Audio data

[0142] Step 8:

[0143] The terminal feeds back the voice data to the user.

[0144] Input: Voice data (e.g., "Start navigation to Tokyo")

[0145] How it works: The device provides audio feedback to the user through a Bluetooth headset or the built-in speaker.

[0146] Output: Audio feedback to the user

[0147] This is the flow of processing in the program for this system. At each step, specific data processing or data calculation is performed based on appropriate input data, and output data is generated as a result.

[0148] (Application example 1)

[0149] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0150] Modern delivery services require riders to access information and perform operations safely and efficiently while driving. However, many current systems require riders to manually operate smartphones and other devices, which reduces safety. Therefore, there is a need to develop a system that allows riders to navigate, check notifications, respond by voice, report situations, and make emergency calls while driving hands-free.

[0151] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0152] In this invention, the server includes means for receiving voice input, means for converting voice to text, means for analyzing the text and determining an action, means for providing functions of navigation, notification confirmation, voice response, status report, and emergency contact, and means for providing voice feedback, thereby enabling delivery riders to perform necessary operations and obtain information safely and efficiently in a hands-free manner while driving.

[0153] A "means for receiving audio input" is a microphone or other sound capturing device that detects and converts sound into a digital signal.

[0154] "Means for converting speech to text" means speech recognition software or algorithms for converting speech to text data.

[0155] The "means for determining an action" is an analysis algorithm for analyzing the converted text data and identifying the operation intended by the user.

[0156] The "means for sending a request to a server" is a communication interface for sending information to a server via the Internet based on the determined action.

[0157] The "means for receiving a response from the server and providing feedback to the user as voice" is a system for outputting information received from the server as voice using voice synthesis technology.

[0158] "Navigation" is a function that provides route guidance to a destination.

[0159] "Notification Check" is a function that notifies the user of new notifications and messages.

[0160] "Voice response" is a function that responds to voice input with voice.

[0161] "Status Report" is a function that provides audio reports on the progress of deliveries and other tasks.

[0162] "Emergency contact" is a function that allows you to make contact quickly when a problem occurs.

[0163] "Communication means" refers to infrastructure such as internet connections and mobile networks for sending and receiving data.

[0164] A "speech synthesizer" is hardware or software that converts text data into speech.

[0165] This invention relates to a system that uses voice input to enable delivery service riders to perform various operations hands-free while driving. Specifically, the system can provide functions such as navigation, notification confirmation, voice response, status reports, and emergency contact using voice input. This system comprises a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0166] A user issues a voice command, such as "Navigate to Shinjuku." This voice is captured by a device (e.g., a smartphone or smart glasses) and recorded as an audio file. The audio file is then converted into text data using a speech recognition service (e.g., Azure Speech to Text API). The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Shinjuku" requires a navigation function.

[0167] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0168] The response from the server is received by the device and fed back to the user as an easy-to-understand voice message. This voice feedback is output by converting the text message into voice using the Azure Text to Speech API. This means that riders can obtain necessary information and perform operations using only voice without having to operate their mobile device while driving.

[0169] As a specific example, if a user says "Navigate to Shinjuku" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Shinjuku." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Shinjuku and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Shinjuku."

[0170] Other prompts include "Check notifications," "Respond to customer, on my way," "Report status," and "Urgent," helping delivery riders navigate safely and efficiently while driving.

[0171] The hardware used in this system configuration is end-user devices such as smartphones and smart glasses. The software uses services such as Azure Speech to Text API, Azure Text to Speech API, Google Maps API, and Twilio API. The combination of this hardware and software improves the user experience and ensures security.

[0172] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0173] Step 1:

[0174] The user issues a voice command, for example, "Navigate to Shinjuku." This voice input is the starting point of the process.

[0175] Step 2:

[0176] The device uses a microphone to capture the voice and collects the voice data as a digital signal. This is the process of acquiring voice input. The input is the user's voice, and the output is digital voice data.

[0177] Step 3:

[0178] The device sends voice data to a speech recognition service (for example, Azure Speech to Text API) that converts the voice data into text data. The input is digital voice data and the output is text data. The speech recognition service analyzes the input voice and generates corresponding text.

[0179] Step 4:

[0180] The terminal analyzes the converted text data. Specifically, it analyzes the text "Navigate to Shinjuku" and determines that this command requests a navigation function. The input is text data, and the output is the analysis result (action).

[0181] Step 5:

[0182] The terminal sends an HTTP request to the server based on the analysis results. This request includes the analyzed destination information and other necessary parameters. The input is the analysis results, and the output is an HTTP request to the server.

[0183] Step 6:

[0184] The server receives the HTTP request and performs processing according to the request. In the case of navigation, it generates route information to the destination. The input is the HTTP request, and the output is the server's response, such as route information.

[0185] Step 7:

[0186] The server returns the generated route information to the terminal. The input is the route information, and the output is an HTTP response.

[0187] Step 8:

[0188] The device receives the response from the server and converts this information into a voice message using a speech synthesizer (for example, Azure Text to Speech API). The input is the server response and the output is the voice message. The speech synthesizer converts the text message into speech.

[0189] Step 9:

[0190] The terminal provides a voice message as feedback to the user. For example, it may say, "Starting navigation to Shinjuku." The input is the voice message, and the output is the voice feedback to the user.

[0191] Through this process, delivery riders can receive safe navigation to their destination while driving.

[0192] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0193] This invention provides a system that allows a user to perform various operations hands-free while riding a motorcycle, by combining it with an emotion engine that recognizes the user's emotions, and provides more appropriate and personalized feedback and actions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, the emotion engine, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[0194] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as a digital audio file in real time. The device then sends the recorded voice data to a voice recognition service, which converts the voice data into text data. At the same time, the emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[0195] The converted text data and emotional information are analyzed by an analysis means within the device. For example, in the command "Navigate to Tokyo," "Tokyo" is determined to be the destination, and the user's emotional state, such as whether they are nervous or calm, is recognized. Based on the results of this analysis, an appropriate action is determined. In this case, since the "navigate" function is requested, it is determined that the action should be processed as a navigation action.

[0196] Furthermore, the emotion engine's information is taken into account to adjust the expression and content of the feedback. For example, if the user is nervous, navigation instructions will be provided in a calmer voice tone, and conversely, if the user is calm, instructions will be provided in a normal tone.

[0197] Based on the determined action, the device sends an HTTP request to the server. The request includes necessary parameters such as destination information and emotional state. When the server receives the request, it processes it according to its content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0198] The response from the server is received by the device and fed back to the user. A speech synthesizer converts the text message into speech and provides it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[0199] For example, if a user says "Navigate to Tokyo" while riding, the voice is captured by the device, and the speech recognition system converts it into text data, while the emotion engine simultaneously recognizes the sense of tension. The analysis result is sent to the server as a navigation request, and the server provides route information to Tokyo. Finally, the device provides feedback to the user with an adjusted voice message, providing a safer and more comfortable riding experience.

[0200] This system can be operated safely and efficiently not only when driving a car but also when driving a motorcycle. In particular, by combining it with an emotion engine, personalized interactions are possible, greatly improving the safety and convenience of riders.

[0201] The processing flow will be explained below.

[0202] Step 1:

[0203] The user issues a voice command such as "Navigate to Tokyo," which is captured using a microphone on the device.

[0204] Step 2:

[0205] The device records the captured audio as a digital audio file, which is temporarily stored on the device's storage.

[0206] Step 3:

[0207] The device sends the recorded voice data to a voice recognition service, where it is converted into text data. At the same time, an emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[0208] Step 4:

[0209] The device receives the text data and emotion information and analyzes the command content and the user's emotion using an analysis means. Specifically, the device recognizes that the command "Navigate to Tokyo" means the destination "Tokyo" and that the user's emotion is, for example, nervous.

[0210] Step 5:

[0211] The terminal determines an appropriate action based on the analysis result. In this case, since the "navigate" function is requested, it determines to process it as a navigation action.

[0212] Step 6:

[0213] The device sends an HTTP request to the server according to the determined action, including necessary parameters such as destination information and user emotion information.

[0214] Step 7:

[0215] The server receives the HTTP request from the device and processes it accordingly. Specifically, it generates route information to the destination and formats and prepares the message.

[0216] Step 8:

[0217] The server sends the generated route information and an appropriate message back to the device as a response, which includes navigation information and an emotion-based message.

[0218] Step 9:

[0219] The terminal receives the response from the server and analyzes its contents to provide appropriate feedback to the user. In this case, a text message is analyzed.

[0220] Step 10:

[0221] The device's speech synthesizer converts the text message into speech, generating a voice message that reflects the results of the emotion engine. For example, the message might say, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[0222] Step 11:

[0223] Users can receive voice feedback from the device and obtain real-time navigation information without taking their hands off the steering wheel.

[0224] Through these steps, the system provides safe and efficient operation for the rider while riding the motorcycle, and in particular, by using an emotion engine, it enables personalized feedback according to the user's emotional state.

[0225] Example 2

[0226] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0227] When users operate hands-free while driving a motorcycle, they are faced with the challenge of being unable to operate the system safely and efficiently. Furthermore, conventional hands-free systems lack personalized feedback that takes into account the user's emotions, limiting the user experience. This increases stress and anxiety for users while driving, and creates the problem of insufficient safety and convenience.

[0228] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0229] In this invention, the server includes means for receiving a voice input, means for converting the received voice input into text, means for extracting emotional information from the converted text and voice information, means for analyzing the converted text and the extracted emotional information to determine an appropriate action, means for sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice based on the emotional information. This makes it possible to provide appropriate and personalized feedback and actions that take into account the emotional state of the user.

[0230] "Voice input" refers to commands or instructions spoken by a user received as digital data through a microphone or other audio collection device.

[0231] "Text" refers to a string of characters, such as alphabets or kanji, that has been converted from voice input, and is information in a format that can be analyzed on a device such as a computer.

[0232] "Emotional information" is information extracted by analyzing the user's emotional state (e.g., tension, calmness, anxiety, etc.) from voice data.

[0233] "Analysis" is a process for deriving appropriate actions based on the received text data and emotional information.

[0234] "Actions" refer to the operations or reactions the system performs based on the user's voice commands, such as navigation, music playback, and answering calls.

[0235] A "request" is an inquiry or request sent from a terminal to a server, and is executed in the form of an HTTP request or the like.

[0236] A "response" is a reply or answer sent from a server to a terminal, and is data that includes information or instructions corresponding to a request.

[0237] A "speech recognition service" is any cloud-based or on-premise technology or software that accepts voice data as input and converts it into text data.

[0238] A "speech synthesizer" is a device that converts text information into speech and provides speech output to the user.

[0239] This invention is a system that allows a user to perform various operations hands-free while riding a motorcycle, and provides personalized feedback and actions by combining it with an emotion engine that recognizes the user's emotions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for extracting emotion information from the text and voice data, a means for analyzing the text and emotion information and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[0240] First, when a user issues a voice command such as "navigate to a destination," the device captures this voice and records it as a digital audio file via the microphone. This is real-time voice capture.

[0241] Next, the terminal transmits the recorded voice data to a voice recognition service (e.g., voice recognition software) to convert the voice data into text data, and also uses an emotion engine (e.g., emotion analysis software) to analyze the user's emotions from the voice data and extract emotion information.

[0242] The converted text data and emotional information are then analyzed by an analysis unit within the device. For example, in the command "Navigate to destination," the device determines that "destination" is the destination and simultaneously recognizes the user's emotional state, such as whether they are nervous or calm.

[0243] Based on the analysis results, the system determines the appropriate action. In this case, it determines that the "navigation" function is requested. The expression and content of the feedback are adjusted taking into account the information from the emotion engine. If the user is nervous, the device will provide navigation guidance in a calm voice tone, and if the user is calm, it will provide guidance in a normal tone.

[0244] Furthermore, based on the determined action, the device sends an HTTP request to the server. This request includes necessary parameters such as destination information and emotional state. When the server receives this request, it performs processing according to the content and generates a corresponding response. For example, if it is a navigation request, it calculates and generates route information to the destination.

[0245] The response from the server is received by the device, which then uses a speech synthesizer (e.g., speech synthesis software) to convert the text message into speech and provide it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to your destination will begin."

[0246] As a concrete example, if a user says "Navigate to Tokyo" while driving, the device captures the voice and converts it into text data using a speech recognition service, while the emotion engine simultaneously recognizes the user's nervousness. Based on the analysis results, the device sends this as a navigation request to the server, which then provides route information to Tokyo. Using a speech synthesizer, the device then provides the user with an adjusted voice message such as "You seem nervous, but please remain calm while driving. Navigation to Tokyo will begin," providing a safer and more comfortable riding experience.

[0247] Examples of prompts for generative AI models include:

[0248] > "Design a system that allows a user to operate a motorcycle hands-free while driving. The system combines voice input with an emotion engine to provide appropriate feedback based on the user's emotional state. Specifically, if the user is nervous, the system will provide navigation information in a calm tone, and if the user is calm, the system will provide information in a normal tone."

[0249] The system's unique feature is that it significantly improves the user's safety and convenience while riding a motorcycle, and by utilizing an emotion engine, personalized interactions based on the user's emotions are possible.

[0250] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0251] Program processing flow

[0252] Step 1: Capture voice commands

[0253] The user issues the voice command "Navigate to Tokyo."

[0254] Input: Voice command

[0255] How it works: The device captures the user's voice commands through the microphone and records them as digital audio files in real time.

[0256] Output: Digital audio file

[0257] Step 2: Speech to text

[0258] The device transmits the recorded digital audio file to a speech recognition service.

[0259] Input: Digital audio file

[0260] How it works: A speech recognition service converts speech data into text data. For example, speech recognition software is used to generate the text "Navigate to Tokyo."

[0261] Output: Text data

[0262] Step 3: Extracting emotional information

[0263] The device sends the digital audio file to the emotion engine.

[0264] Input: Digital audio file

[0265] How it works: The emotion engine analyzes the user's emotions from the voice data and extracts emotional information. For example, it uses emotion analysis software to determine that the user is "nervous."

[0266] Output: Emotional information

[0267] Step 4: Analyzing text data and sentiment information

[0268] The analysis unit in the device analyzes the text data and emotional information.

[0269] Input: Text data, emotion information

[0270] Action: The analysis unit analyzes the text "Navigate to Tokyo" and the emotional information "I'm nervous" and determines the appropriate action. Specifically, it determines that "Tokyo" is the destination and that the "Navigate" function is required.

[0271] Output: Analysis results (actions, destination information, emotion information)

[0272] Step 5: Sending a request to the server

[0273] The terminal sends an HTTP request to the server based on the analysis results.

[0274] Input: Analysis results

[0275] How it works: The HTTP request contains destination and emotion information. By sending this request, the server is asked to provide navigation information.

[0276] Output: HTTP request

[0277] Step 6: Server-side processing

[0278] The server receives the HTTP request and performs the corresponding processing.

[0279] Input: HTTP request

[0280] Operation: The server analyzes the request (destination information, emotion information) and generates navigation information. For example, it uses route calculation software to calculate the route to Tokyo.

[0281] Output: Navigation information

[0282] Step 7: Receiving responses and coordinating feedback

[0283] The terminal receives the response from the server.

[0284] Input: Navigation information

[0285] How it works: The device uses a speech synthesizer to convert navigation information into a voice message to convey to the user. The device adjusts the tone of the voice message based on emotional information. For example, it generates a message like, "You seem nervous, but please stay calm and drive. We'll start navigating to Tokyo."

[0286] Output: Modified voice message

[0287] Step 8: Provide feedback to users

[0288] The terminal finally provides feedback to the user as a voice message.

[0289] Input: Modified voice message

[0290] Operation: The terminal outputs a voice message to the user through a voice synthesizer.

[0291] Output: The user receives the adjusted navigation instructions.

[0292] Specific examples

[0293] When a user says "Navigate to Tokyo," the speech is captured by the device and converted into text by the speech recognition service. At the same time, the speech data is analyzed by the emotion engine, which recognizes that the user is nervous. Based on the results, the analysis unit sends an HTTP request to the server as a navigation action, requesting directions to Tokyo. The server receives the request and provides route information to Tokyo, and the device provides the user with a voice message tailored based on the emotion information.

[0294] (Application example 2)

[0295] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0296] Existing technologies include systems that use voice commands to perform various operations, but these systems are unable to recognize the user's emotional state and provide appropriate feedback based on that. This results in a uniform user experience, which does not improve safety or convenience, especially in emergencies or stressful situations. Furthermore, there is a lack of a hands-free means of operating robots to improve work efficiency and safety in factories.

[0297] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and emotional information to determine an appropriate action, means for combining an emotion engine that recognizes the user's emotion and sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice. This makes it possible to provide personalized feedback and actions according to the user's emotional state, greatly improving work efficiency and safety in factories.

[0298] "Means for receiving voice input" refers to a function for capturing voice commands issued by a user through a device such as a microphone.

[0299] A "means for converting voice input to text" is a function that allows for analyzing captured voice data and converting it into text format.

[0300] "Means for analyzing the converted text and determining the appropriate action" refers to a function that understands and analyzes the voice command converted into text and selects the appropriate action based on the results.

[0301] The "emotion engine" is a function for recognizing and analyzing the user's emotional state from voice data.

[0302] The "means for sending a request to the server" is a function for sending a request to the server for an appropriate action based on the analyzed text data and emotion information.

[0303] The "means for receiving a response from the server and providing feedback to the user as audio" is a function for providing feedback to the user by outputting the data received from the server as audio.

[0304] "Navigation, music playback, call answering, weather information check, emergency communication, robot operation" are the various types of actions that the system can perform in response to a user's voice commands.

[0305] A "prompt sentence using a generative AI model" is an input sentence used by a generative AI technology to obtain a desired output.

[0306] This invention relates to a system that enables a worker working in a factory to operate a robot hands-free using smart glasses. The system includes means for receiving voice input, means for converting the voice input into text, means for analyzing the converted text and emotion information to determine an appropriate action, an emotion engine for recognizing the emotion of the user, means for sending a request to a server based on the determined action and emotion information, and means for receiving a response from the server and providing feedback to the user as voice.

[0307] As a concrete example of the system, a worker puts on smart glasses and issues a command by voice, such as "Transport this part." The smart glasses' microphone captures the voice command and generates a digital audio file in real time. The smart glasses then send the recorded voice data to a voice recognition service (e.g., Google Speech-to-Text) to convert the voice data into text data. At the same time, an emotion engine (e.g., EmotionRecognition API) analyzes the user's emotional state and extracts emotional information.

[0308] The converted text data and emotion information are further analyzed by the analysis means in the smart glasses. For example, for a command such as "transport this part," the emotion engine determines that there is a transport command and recognizes that the user is nervous. Based on this result, an appropriate action is determined. If the user is nervous, commands are sent to the robot to operate as smoothly and safely as possible.

[0309] Based on the determined action, the smart glasses send an HTTP request to the server. The request includes parameters related to the transport command and the emotional state. The server receives the request and generates specific operational commands for the robot, such as calculating and returning a smooth transport route to the destination.

[0310] The response from the server is received by the smart glasses. The content is read as a text message and audio feedback is provided to the user using a speech synthesizer (e.g., Google Text-to-Speech). Feedback that takes into account the user's state of tension is provided, such as "You seem nervous, but please calm down. We'll start transporting the parts."

[0311] Hardware and software used

[0312] Smart glasses (e.g., generic name)

[0313] microphone

[0314] Connected robot arm or mobile robot

[0315] Speech recognition services (e.g., Google Speech-to-Text)

[0316] Emotion recognition engine (e.g., EmotionRecognition API)

[0317] Speech synthesizers (e.g., Google Text-to-Speech)

[0318] Prompt Sentence Examples

[0319] Prompt: Recognizes user voice commands and their emotions

[0320] Input: Audio data text and its audio file

[0321] Output: Command converted to text and recognized emotion information (e.g., carrying, stressed)

[0322] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0323] Step 1:

[0324] A user issues a voice command through the smart glasses. The input includes the user's speech. A microphone on the smart glasses captures the voice command and records it as a digital audio file. The output is a digital audio file.

[0325] Step 2:

[0326] The device sends the captured voice data to a speech recognition service, which converts the voice data into text data. The input includes a digital audio file. The speech recognition service used (e.g., Google Speech-to-Text) analyzes the voice signal and generates corresponding text. The output is the converted text data.

[0327] Step 3:

[0328] The device simultaneously sends the voice data to an emotion recognition engine to analyze the user's emotional state. The input includes a digital audio file. The emotion recognition engine (e.g., EmotionRecognition API) analyzes the tone and rate of the voice to extract emotional information. The output is the user's emotional information.

[0329] Step 4:

[0330] The device analyzes the converted text data and emotional information to determine the appropriate action. The input includes text data and emotional information. The text analysis engine understands the command content and combines it with the emotional information to determine the appropriate action. The output is the determined action.

[0331] Step 5:

[0332] The device sends an HTTP request to the server based on the determined action and emotion information. The input includes the action and emotion information. The HTTP protocol is used to send the request with appropriate parameters to the server. The output is a request that is processed on the server side.

[0333] Step 6:

[0334] Based on the request received by the server, it generates specific operation commands for the robot. The input includes action and emotion information. The logic engine on the server calculates the optimal route and operation parameters based on this data and sends the command to the robot. The output is the operation command sent to the robot.

[0335] Step 7:

[0336] The robot starts working based on the operation command from the server. The operation command from the server is included as input. The robot executes the specified task according to the command. The output is the execution status of the task.

[0337] Step 8:

[0338] The server receives feedback from the robot and sends the feedback to the smart glasses. The input includes feedback information from the robot. The server generates this information as a voice message and uses a voice synthesizer (e.g., Google Text-to-Speech) to generate the voice feedback. The output is the voice message.

[0339] Step 9:

[0340] The smart glasses provide the generated audio feedback to the user. The input includes an audio message. The smart glasses provide audio feedback to the user such as, "You seem nervous, please stay calm. We'll start transporting the parts." The output is the audio feedback provided to the user.

[0341] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0342] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0343] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0344] [Second embodiment]

[0345] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0346] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0347] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0348] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0349] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0350] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0351] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0352] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0353] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0354] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0355] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0356] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0357] The present invention relates to a system that allows a rider to perform various operations hands-free while riding a motorcycle. Specifically, it can realize functions such as navigation, music playback, phone call answering, weather information confirmation, and emergency communication using voice input. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0358] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as an audio file. The audio file is then converted into text data using a voice recognition service. The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Tokyo," for example, requires a navigation function.

[0359] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0360] The response from the server is received by the device and fed back to the user as a voice message that is easy for the rider to understand. This voice feedback is output by converting the text message into voice using a voice synthesizer built into the device. This allows the rider to obtain necessary information and perform operations using only voice, without having to take their hands off the handlebars.

[0361] For example, if a user says "Navigate to Tokyo" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Tokyo." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Tokyo and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Tokyo."

[0362] This system is designed to be safe and efficient to operate not only while driving a car, but also while driving a motorcycle, etc. By combining functions such as voice recognition, analysis, communication, and voice feedback, a wide range of operations can be performed hands-free, greatly improving the safety and convenience of riders.

[0363] The processing flow will be explained below.

[0364] Step 1:

[0365] The user issues a voice command such as "Navigate to Tokyo," and the voice is captured by a microphone attached to the device.

[0366] Step 2:

[0367] The device records the captured audio as a digital audio file, at which point the audio data is temporarily stored.

[0368] Step 3:

[0369] The device sends the recorded voice data to a voice recognition service, converts the voice data into text data, and receives the text data returned by the voice recognition service, for example, using a cloud-based voice recognition API.

[0370] Step 4:

[0371] The terminal analyzes the received text data and extracts the command content. For example, in the command "Navigate to Tokyo," the analysis means determines that "Tokyo" is the destination.

[0372] Step 5:

[0373] The terminal determines an appropriate action based on the analysis result. In this case, since "navigate" is requested, it determines to process it as a navigation action.

[0374] Step 6:

[0375] Based on the determined action, the device sends an HTTP request to the server, including necessary parameters such as destination information.

[0376] Step 7:

[0377] The server receives the HTTP request and performs the appropriate action based on the request, for example, generating route information to the destination.

[0378] Step 8:

[0379] The server sends the generated information back to the device as a response, which includes detailed route information for navigation.

[0380] Step 9:

[0381] The terminal receives the response from the server and analyzes the content in the form of a text message.

[0382] Step 10:

[0383] The device's speech synthesizer converts the text message into speech and provides feedback to the rider, such as a voice message saying, "Navigation to Tokyo is about to begin."

[0384] Step 11:

[0385] By receiving voice feedback from the device, users can obtain navigation information without taking their hands off the steering wheel.

[0386] Through the above processing steps, the present invention enables a rider to safely and efficiently use voice commands while riding a motorcycle.

[0387] Example 1

[0388] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0389] Currently, using devices such as smartphones to check navigation, play music, or make calls while riding a motorcycle is extremely dangerous. There is a demand for hands-free systems that allow these operations to be performed without using the hands while driving. There is also a need to develop systems that simultaneously improve safety and convenience.

[0390] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0391] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and determining an appropriate action, means for sending a request to an information processing device based on the determined action, means for receiving a response from the information processing device and feeding it back to the user as voice, means for using a voice recognition service, means for analyzing text using a natural language processing library, means for sending an HTTP request, and means for converting into voice using a voice synthesis device.

[0392] This allows users to navigate, play music, answer calls, check weather information, make emergency calls, and perform other operations using only voice input while riding a motorcycle, without using their hands.

[0393] "Voice input" refers to the means by which a user gives instructions to a system using voice.

[0394] "Speech Recognition Service" means external or internal software or functionality used to convert received speech into text data.

[0395] "Text conversion" refers to the process of analyzing speech input and converting it into corresponding text data.

[0396] A "natural language processing library" is a software library for analyzing text data and understanding meaning and actions.

[0397] An "HTTP request" is a communication method using a protocol for a client to request data from a server.

[0398] An "information processing device" refers to a computer or server that processes data based on a received request and generates a response.

[0399] A "voice synthesis device" is a device or software that converts text data into voice data and outputs it.

[0400] "Communication means" refers to the technology or protocol for transferring data, including wired and wireless communication methods.

[0401] "Feedback" refers to the system returning processing results or information to the user, and in this context it primarily refers to audio feedback.

[0402] "Location information provision" is a function that provides information about the user's specified destination or current location.

[0403] "Media content control" is a function for operating media content such as music playback and video playback.

[0404] "Communication response" is a function that responds to communication-related commands, such as incoming and outgoing phone calls.

[0405] "Weather information acquisition" refers to a function of acquiring weather information such as the current weather and forecast and providing it to the user.

[0406] "Safety communication" is a function for quickly communicating necessary information in an emergency.

[0407] The present invention relates to a system that allows riders to perform various operations hands-free while riding a motorcycle. This system uses voice input to realize functions such as navigation, music playback, answering calls, checking weather information, and emergency communication. The user can operate the terminal using voice input, and the terminal provides the necessary information and services through communication with a server.

[0408] The present invention consists of the following main components:

[0409] 1. Voice input means: A means for users to input commands by voice. The device is equipped with a high-sensitivity microphone that accurately captures the user's voice. Specific hardware used for this is a Bluetooth-enabled headset or smartphone.

[0410] 2. Speech recognition means: This is a means for converting received voice input into text, and uses speech recognition services such as IBM Watson and Google Cloud Speech-to-Text. These services are cloud-based and enable highly accurate speech recognition.

[0411] 3. Text analysis: This is a means of analyzing the converted text data and determining the appropriate action. The device uses natural language processing libraries such as spaCy and NLTK to analyze the text and understand the user's intent.

[0412] 4. Communication method: This is a method for sending an HTTP request to the server based on the determined action. The device uses this method to communicate the analysis results to the server. Communication is secure using HTTPS.

[0413] 5. Response processing means: This is the means for receiving responses from the server and providing audio feedback to the user. After receiving data from the server, the device uses Amazon Polly or Google Text-to-Speech to convert the text data into audio and provide it to the user.

[0414] For example, if a user says "Navigate to Tokyo" while driving, the following process will occur:

[0415] The device captures the user's voice and sends the voice data to a cloud service.

[0416] The voice recognition service converts the voice data into text data such as "Navigate to Tokyo."

[0417] The terminal analyzes this text data and determines that it is a navigation request.

[0418] The device creates an HTTP request and sends the destination information to the server.

[0419] The server generates route information to Tokyo using the Google Maps API or similar and sends that information back to the device.

[0420] The device analyzes the route information it receives and uses a voice synthesizer to provide voice feedback to the user, saying, "Navigation to Tokyo is about to begin."

[0421] An example prompt is:

[0422] "Describe a system that enables navigation functions based on voice input. For example, capture voice input of "Navigate to Tokyo" and demonstrate the flow of speech recognition, natural language processing, sending an HTTP request, processing on the server, and voice feedback."

[0423] This system allows riders to perform many operations while riding their motorcycle using only voice commands, without using their hands, greatly improving safety and convenience.The system can also be used in other driving situations and with a variety of voice commands, making it suitable for a wide range of uses.

[0424] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0425] Process flow steps and explanations

[0426] Step 1:

[0427] The terminal captures the user's voice.

[0428] Input: User's voice command (e.g., "Navigate to Tokyo")

[0429] How it works: The device's high-sensitivity microphone picks up the user's voice.

[0430] Output: Audio data

[0431] Step 2:

[0432] The device sends the captured audio to a speech recognition service.

[0433] Input: Audio data

[0434] How it works: Your device sends voice data to a cloud-based speech recognition service (e.g., IBM Watson, Google Cloud Speech-to-Text).

[0435] Output: Text data (e.g. "Navigate to Tokyo")

[0436] Step 3:

[0437] The device analyzes the received text data using a natural language processing library.

[0438] Input: Text data (e.g., "Navigate to Tokyo")

[0439] How it works: The terminal uses natural language processing libraries such as spaCy or NLTK to parse the text data and determine the intent of the command.

[0440] Output: Analysis results (e.g., navigation requests)

[0441] Step 4:

[0442] The terminal constructs an HTTP request based on the analysis results and sends it to the server.

[0443] Input: Analysis result (e.g., navigation request, destination "Tokyo")

[0444] How it works: The device constructs an HTTP request and sends it to the server, including the analysis results and additional parameters (e.g., current location, departure time).

[0445] Output: HTTP request to the server

[0446] Step 5:

[0447] The server processes the received HTTP request and retrieves the required data.

[0448] Input: HTTP request (e.g., destination "Tokyo")

[0449] How it works: The server parses the request and retrieves route information using external services such as the Google Maps API or OpenStreetMap.

[0450] Output: Route information

[0451] Step 6:

[0452] The server generates a response containing the route information and sends it back to the terminal.

[0453] Input: Route information

[0454] Operation: The server generates an HTTP response based on the obtained route information and sends it to the terminal.

[0455] Output: HTTP response to the device

[0456] Step 7:

[0457] The terminal receives the response from the server and converts it into speech using a speech synthesizer.

[0458] Input: HTTP response (e.g. route information)

[0459] How it works: The device uses Amazon Polly or Google Text-to-Speech to convert text data (route guidance) into voice data.

[0460] Output: Audio data

[0461] Step 8:

[0462] The terminal feeds back the voice data to the user.

[0463] Input: Voice data (e.g., "Start navigation to Tokyo")

[0464] How it works: The device provides audio feedback to the user through a Bluetooth headset or the built-in speaker.

[0465] Output: Audio feedback to the user

[0466] This is the flow of processing in the program for this system. At each step, specific data processing or data calculation is performed based on appropriate input data, and output data is generated as a result.

[0467] (Application example 1)

[0468] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0469] Modern delivery services require riders to access information and perform operations safely and efficiently while driving. However, many current systems require riders to manually operate smartphones and other devices, which reduces safety. Therefore, there is a need to develop a system that allows riders to navigate, check notifications, respond by voice, report situations, and make emergency calls while driving hands-free.

[0470] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0471] In this invention, the server includes means for receiving voice input, means for converting voice to text, means for analyzing the text and determining an action, means for providing functions of navigation, notification confirmation, voice response, status report, and emergency contact, and means for providing voice feedback, thereby enabling delivery riders to perform necessary operations and obtain information safely and efficiently in a hands-free manner while driving.

[0472] A "means for receiving audio input" is a microphone or other sound capturing device that detects and converts sound into a digital signal.

[0473] "Means for converting speech to text" means speech recognition software or algorithms for converting speech to text data.

[0474] The "means for determining an action" is an analysis algorithm for analyzing the converted text data and identifying the operation intended by the user.

[0475] The "means for sending a request to a server" is a communication interface for sending information to a server via the Internet based on the determined action.

[0476] The "means for receiving a response from the server and providing feedback to the user as voice" is a system for outputting information received from the server as voice using voice synthesis technology.

[0477] "Navigation" is a function that provides route guidance to a destination.

[0478] "Notification Check" is a function that notifies the user of new notifications and messages.

[0479] "Voice response" is a function that responds to voice input with voice.

[0480] "Status Report" is a function that provides audio reports on the progress of deliveries and other tasks.

[0481] "Emergency contact" is a function that allows you to make contact quickly when a problem occurs.

[0482] "Communication means" refers to infrastructure such as internet connections and mobile networks for sending and receiving data.

[0483] A "speech synthesizer" is hardware or software that converts text data into speech.

[0484] This invention relates to a system that uses voice input to enable delivery service riders to perform various operations hands-free while driving. Specifically, the system can provide functions such as navigation, notification confirmation, voice response, status reports, and emergency contact using voice input. This system comprises a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0485] A user issues a voice command, such as "Navigate to Shinjuku." This voice is captured by a device (e.g., a smartphone or smart glasses) and recorded as an audio file. The audio file is then converted into text data using a speech recognition service (e.g., Azure Speech to Text API). The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Shinjuku" requires a navigation function.

[0486] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0487] The response from the server is received by the device and fed back to the user as an easy-to-understand voice message. This voice feedback is output by converting the text message into voice using the Azure Text to Speech API. This means that riders can obtain necessary information and perform operations using only voice without having to operate their mobile device while driving.

[0488] As a specific example, if a user says "Navigate to Shinjuku" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Shinjuku." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Shinjuku and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Shinjuku."

[0489] Other prompts include "Check notifications," "Respond to customer, on my way," "Report status," and "Urgent," helping delivery riders navigate safely and efficiently while driving.

[0490] The hardware used in this system configuration is end-user devices such as smartphones and smart glasses. The software uses services such as Azure Speech to Text API, Azure Text to Speech API, Google Maps API, and Twilio API. The combination of this hardware and software improves the user experience and ensures security.

[0491] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0492] Step 1:

[0493] The user issues a voice command, for example, "Navigate to Shinjuku." This voice input is the starting point of the process.

[0494] Step 2:

[0495] The device uses a microphone to capture the voice and collects the voice data as a digital signal. This is the process of acquiring voice input. The input is the user's voice, and the output is digital voice data.

[0496] Step 3:

[0497] The device sends voice data to a speech recognition service (for example, Azure Speech to Text API) that converts the voice data into text data. The input is digital voice data and the output is text data. The speech recognition service analyzes the input voice and generates corresponding text.

[0498] Step 4:

[0499] The terminal analyzes the converted text data. Specifically, it analyzes the text "Navigate to Shinjuku" and determines that this command requests a navigation function. The input is text data, and the output is the analysis result (action).

[0500] Step 5:

[0501] The terminal sends an HTTP request to the server based on the analysis results. This request includes the analyzed destination information and other necessary parameters. The input is the analysis results, and the output is an HTTP request to the server.

[0502] Step 6:

[0503] The server receives the HTTP request and performs processing according to the request. In the case of navigation, it generates route information to the destination. The input is the HTTP request, and the output is the server's response, such as route information.

[0504] Step 7:

[0505] The server returns the generated route information to the terminal. The input is the route information, and the output is an HTTP response.

[0506] Step 8:

[0507] The device receives the response from the server and converts this information into a voice message using a speech synthesizer (for example, Azure Text to Speech API). The input is the server response and the output is the voice message. The speech synthesizer converts the text message into speech.

[0508] Step 9:

[0509] The terminal provides a voice message as feedback to the user. For example, it may say, "Starting navigation to Shinjuku." The input is the voice message, and the output is the voice feedback to the user.

[0510] Through this process, delivery riders can receive safe navigation to their destination while driving.

[0511] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0512] This invention provides a system that allows a user to perform various operations hands-free while riding a motorcycle, by combining it with an emotion engine that recognizes the user's emotions, and provides more appropriate and personalized feedback and actions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, the emotion engine, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[0513] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as a digital audio file in real time. The device then sends the recorded voice data to a voice recognition service, which converts the voice data into text data. At the same time, the emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[0514] The converted text data and emotional information are analyzed by an analysis means within the device. For example, in the command "Navigate to Tokyo," "Tokyo" is determined to be the destination, and the user's emotional state, such as whether they are nervous or calm, is recognized. Based on the results of this analysis, an appropriate action is determined. In this case, since the "navigate" function is requested, it is determined that the action should be processed as a navigation action.

[0515] Furthermore, the emotion engine's information is taken into account to adjust the expression and content of the feedback. For example, if the user is nervous, navigation instructions will be provided in a calmer voice tone, and conversely, if the user is calm, instructions will be provided in a normal tone.

[0516] Based on the determined action, the device sends an HTTP request to the server. The request includes necessary parameters such as destination information and emotional state. When the server receives the request, it processes it according to its content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0517] The response from the server is received by the device and fed back to the user. A speech synthesizer converts the text message into speech and provides it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[0518] For example, if a user says "Navigate to Tokyo" while riding, the voice is captured by the device, and the speech recognition system converts it into text data, while the emotion engine simultaneously recognizes the sense of tension. The analysis result is sent to the server as a navigation request, and the server provides route information to Tokyo. Finally, the device provides feedback to the user with an adjusted voice message, providing a safer and more comfortable riding experience.

[0519] This system can be operated safely and efficiently not only when driving a car but also when driving a motorcycle. In particular, by combining it with an emotion engine, personalized interactions are possible, greatly improving the safety and convenience of riders.

[0520] The processing flow will be explained below.

[0521] Step 1:

[0522] The user issues a voice command such as "Navigate to Tokyo," which is captured using a microphone on the device.

[0523] Step 2:

[0524] The device records the captured audio as a digital audio file, which is temporarily stored on the device's storage.

[0525] Step 3:

[0526] The device sends the recorded voice data to a voice recognition service, where it is converted into text data. At the same time, an emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[0527] Step 4:

[0528] The device receives the text data and emotion information and analyzes the command content and the user's emotion using an analysis means. Specifically, the device recognizes that the command "Navigate to Tokyo" means the destination "Tokyo" and that the user's emotion is, for example, nervous.

[0529] Step 5:

[0530] The terminal determines an appropriate action based on the analysis result. In this case, since the "navigate" function is requested, it determines to process it as a navigation action.

[0531] Step 6:

[0532] The device sends an HTTP request to the server according to the determined action, including necessary parameters such as destination information and user emotion information.

[0533] Step 7:

[0534] The server receives the HTTP request from the device and processes it accordingly. Specifically, it generates route information to the destination and formats and prepares the message.

[0535] Step 8:

[0536] The server sends the generated route information and an appropriate message back to the device as a response, which includes navigation information and an emotion-based message.

[0537] Step 9:

[0538] The terminal receives the response from the server and analyzes its contents to provide appropriate feedback to the user. In this case, a text message is analyzed.

[0539] Step 10:

[0540] The device's speech synthesizer converts the text message into speech, generating a voice message that reflects the results of the emotion engine. For example, the message might say, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[0541] Step 11:

[0542] Users can receive voice feedback from the device and obtain real-time navigation information without taking their hands off the steering wheel.

[0543] Through these steps, the system provides safe and efficient operation for the rider while riding the motorcycle, and in particular, by using an emotion engine, it enables personalized feedback according to the user's emotional state.

[0544] Example 2

[0545] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0546] When users operate hands-free while driving a motorcycle, they are faced with the challenge of being unable to operate the system safely and efficiently. Furthermore, conventional hands-free systems lack personalized feedback that takes into account the user's emotions, limiting the user experience. This increases stress and anxiety for users while driving, and creates the problem of insufficient safety and convenience.

[0547] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0548] In this invention, the server includes means for receiving a voice input, means for converting the received voice input into text, means for extracting emotional information from the converted text and voice information, means for analyzing the converted text and the extracted emotional information to determine an appropriate action, means for sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice based on the emotional information. This makes it possible to provide appropriate and personalized feedback and actions that take into account the emotional state of the user.

[0549] "Voice input" refers to commands or instructions spoken by a user received as digital data through a microphone or other audio collection device.

[0550] "Text" refers to a string of characters, such as alphabets or kanji, that has been converted from voice input, and is information in a format that can be analyzed on a device such as a computer.

[0551] "Emotional information" is information extracted by analyzing the user's emotional state (e.g., tension, calmness, anxiety, etc.) from voice data.

[0552] "Analysis" is a process for deriving appropriate actions based on the received text data and emotional information.

[0553] "Actions" refer to the operations or reactions the system performs based on the user's voice commands, such as navigation, music playback, and answering calls.

[0554] A "request" is an inquiry or request sent from a terminal to a server, and is executed in the form of an HTTP request or the like.

[0555] A "response" is a reply or answer sent from a server to a terminal, and is data that includes information or instructions corresponding to a request.

[0556] A "speech recognition service" is any cloud-based or on-premise technology or software that accepts voice data as input and converts it into text data.

[0557] A "speech synthesizer" is a device that converts text information into speech and provides speech output to the user.

[0558] This invention is a system that allows a user to perform various operations hands-free while riding a motorcycle, and provides personalized feedback and actions by combining it with an emotion engine that recognizes the user's emotions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for extracting emotion information from the text and voice data, a means for analyzing the text and emotion information and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[0559] First, when a user issues a voice command such as "navigate to a destination," the device captures this voice and records it as a digital audio file via the microphone. This is real-time voice capture.

[0560] Next, the terminal transmits the recorded voice data to a voice recognition service (e.g., voice recognition software) to convert the voice data into text data, and also uses an emotion engine (e.g., emotion analysis software) to analyze the user's emotions from the voice data and extract emotion information.

[0561] The converted text data and emotional information are then analyzed by an analysis unit within the device. For example, in the command "Navigate to destination," the device determines that "destination" is the destination and simultaneously recognizes the user's emotional state, such as whether they are nervous or calm.

[0562] Based on the analysis results, the system determines the appropriate action. In this case, it determines that the "navigation" function is requested. The expression and content of the feedback are adjusted taking into account the information from the emotion engine. If the user is nervous, the device will provide navigation guidance in a calm voice tone, and if the user is calm, it will provide guidance in a normal tone.

[0563] Furthermore, based on the determined action, the device sends an HTTP request to the server. This request includes necessary parameters such as destination information and emotional state. When the server receives this request, it performs processing according to the content and generates a corresponding response. For example, if it is a navigation request, it calculates and generates route information to the destination.

[0564] The response from the server is received by the device, which then uses a speech synthesizer (e.g., speech synthesis software) to convert the text message into speech and provide it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to your destination will begin."

[0565] As a concrete example, if a user says "Navigate to Tokyo" while driving, the device captures the voice and converts it into text data using a speech recognition service, while the emotion engine simultaneously recognizes the user's nervousness. Based on the analysis results, the device sends this as a navigation request to the server, which then provides route information to Tokyo. Using a speech synthesizer, the device then provides the user with an adjusted voice message such as "You seem nervous, but please remain calm while driving. Navigation to Tokyo will begin," providing a safer and more comfortable riding experience.

[0566] Examples of prompts for generative AI models include:

[0567] > "Design a system that allows a user to operate a motorcycle hands-free while driving. The system combines voice input with an emotion engine to provide appropriate feedback based on the user's emotional state. Specifically, if the user is nervous, the system will provide navigation information in a calm tone, and if the user is calm, the system will provide information in a normal tone."

[0568] The system's unique feature is that it significantly improves the user's safety and convenience while riding a motorcycle, and by utilizing an emotion engine, personalized interactions based on the user's emotions are possible.

[0569] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0570] Program processing flow

[0571] Step 1: Capture voice commands

[0572] The user issues the voice command "Navigate to Tokyo."

[0573] Input: Voice command

[0574] How it works: The device captures the user's voice commands through the microphone and records them as digital audio files in real time.

[0575] Output: Digital audio file

[0576] Step 2: Speech to text

[0577] The device transmits the recorded digital audio file to a speech recognition service.

[0578] Input: Digital audio file

[0579] How it works: A speech recognition service converts speech data into text data. For example, speech recognition software is used to generate the text "Navigate to Tokyo."

[0580] Output: Text data

[0581] Step 3: Extracting emotional information

[0582] The device sends the digital audio file to the emotion engine.

[0583] Input: Digital audio file

[0584] How it works: The emotion engine analyzes the user's emotions from the voice data and extracts emotional information. For example, it uses emotion analysis software to determine that the user is "nervous."

[0585] Output: Emotional information

[0586] Step 4: Analyzing text data and sentiment information

[0587] The analysis unit in the device analyzes the text data and emotional information.

[0588] Input: Text data, emotion information

[0589] Action: The analysis unit analyzes the text "Navigate to Tokyo" and the emotional information "I'm nervous" and determines the appropriate action. Specifically, it determines that "Tokyo" is the destination and that the "Navigate" function is required.

[0590] Output: Analysis results (actions, destination information, emotion information)

[0591] Step 5: Sending a request to the server

[0592] The terminal sends an HTTP request to the server based on the analysis results.

[0593] Input: Analysis results

[0594] How it works: The HTTP request contains destination and emotion information. By sending this request, the server is asked to provide navigation information.

[0595] Output: HTTP request

[0596] Step 6: Server-side processing

[0597] The server receives the HTTP request and performs the corresponding processing.

[0598] Input: HTTP request

[0599] Operation: The server analyzes the request (destination information, emotion information) and generates navigation information. For example, it uses route calculation software to calculate the route to Tokyo.

[0600] Output: Navigation information

[0601] Step 7: Receiving responses and coordinating feedback

[0602] The terminal receives the response from the server.

[0603] Input: Navigation information

[0604] How it works: The device uses a speech synthesizer to convert navigation information into a voice message to convey to the user. The device adjusts the tone of the voice message based on emotional information. For example, it generates a message like, "You seem nervous, but please stay calm and drive. We'll start navigating to Tokyo."

[0605] Output: Modified voice message

[0606] Step 8: Provide feedback to users

[0607] The terminal finally provides feedback to the user as a voice message.

[0608] Input: Modified voice message

[0609] Operation: The terminal outputs a voice message to the user through a voice synthesizer.

[0610] Output: The user receives the adjusted navigation instructions.

[0611] Specific examples

[0612] When a user says "Navigate to Tokyo," the speech is captured by the device and converted into text by the speech recognition service. At the same time, the speech data is analyzed by the emotion engine, which recognizes that the user is nervous. Based on the results, the analysis unit sends an HTTP request to the server as a navigation action, requesting directions to Tokyo. The server receives the request and provides route information to Tokyo, and the device provides the user with a voice message tailored based on the emotion information.

[0613] (Application example 2)

[0614] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0615] Existing technologies include systems that use voice commands to perform various operations, but these systems are unable to recognize the user's emotional state and provide appropriate feedback based on that. This results in a uniform user experience, which does not improve safety or convenience, especially in emergencies or stressful situations. Furthermore, there is a lack of a hands-free means of operating robots to improve work efficiency and safety in factories.

[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and emotional information to determine an appropriate action, means for combining an emotion engine that recognizes the user's emotion and sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice. This makes it possible to provide personalized feedback and actions according to the user's emotional state, greatly improving work efficiency and safety in factories.

[0617] "Means for receiving voice input" refers to a function for capturing voice commands issued by a user through a device such as a microphone.

[0618] A "means for converting voice input to text" is a function that allows for analyzing captured voice data and converting it into text format.

[0619] "Means for analyzing the converted text and determining the appropriate action" refers to a function that understands and analyzes the voice command converted into text and selects the appropriate action based on the results.

[0620] The "emotion engine" is a function for recognizing and analyzing the user's emotional state from voice data.

[0621] The "means for sending a request to the server" is a function for sending a request to the server for an appropriate action based on the analyzed text data and emotion information.

[0622] The "means for receiving a response from the server and providing feedback to the user as audio" is a function for providing feedback to the user by outputting the data received from the server as audio.

[0623] "Navigation, music playback, call answering, weather information check, emergency communication, robot operation" are the various types of actions that the system can perform in response to a user's voice commands.

[0624] A "prompt sentence using a generative AI model" is an input sentence used by a generative AI technology to obtain a desired output.

[0625] This invention relates to a system that enables a worker working in a factory to operate a robot hands-free using smart glasses. The system includes means for receiving voice input, means for converting the voice input into text, means for analyzing the converted text and emotion information to determine an appropriate action, an emotion engine for recognizing the emotion of the user, means for sending a request to a server based on the determined action and emotion information, and means for receiving a response from the server and providing feedback to the user as voice.

[0626] As a concrete example of the system, a worker puts on smart glasses and issues a command by voice, such as "Transport this part." The smart glasses' microphone captures the voice command and generates a digital audio file in real time. The smart glasses then send the recorded voice data to a voice recognition service (e.g., Google Speech-to-Text) to convert the voice data into text data. At the same time, an emotion engine (e.g., EmotionRecognition API) analyzes the user's emotional state and extracts emotional information.

[0627] The converted text data and emotion information are further analyzed by the analysis means in the smart glasses. For example, for a command such as "transport this part," the emotion engine determines that there is a transport command and recognizes that the user is nervous. Based on this result, an appropriate action is determined. If the user is nervous, commands are sent to the robot to operate as smoothly and safely as possible.

[0628] Based on the determined action, the smart glasses send an HTTP request to the server. The request includes parameters related to the transport command and the emotional state. The server receives the request and generates specific operational commands for the robot, such as calculating and returning a smooth transport route to the destination.

[0629] The response from the server is received by the smart glasses. The content is read as a text message and audio feedback is provided to the user using a speech synthesizer (e.g., Google Text-to-Speech). Feedback that takes into account the user's state of tension is provided, such as "You seem nervous, but please calm down. We'll start transporting the parts."

[0630] Hardware and software used

[0631] Smart glasses (e.g., generic name)

[0632] microphone

[0633] Connected robot arm or mobile robot

[0634] Speech recognition services (e.g., Google Speech-to-Text)

[0635] Emotion recognition engine (e.g., EmotionRecognition API)

[0636] Speech synthesizers (e.g., Google Text-to-Speech)

[0637] Prompt Sentence Examples

[0638] Prompt: Recognizes user voice commands and their emotions

[0639] Input: Audio data text and its audio file

[0640] Output: Command converted to text and recognized emotion information (e.g., carrying, stressed)

[0641] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0642] Step 1:

[0643] A user issues a voice command through the smart glasses. The input includes the user's speech. A microphone on the smart glasses captures the voice command and records it as a digital audio file. The output is a digital audio file.

[0644] Step 2:

[0645] The device sends the captured voice data to a speech recognition service, which converts the voice data into text data. The input includes a digital audio file. The speech recognition service used (e.g., Google Speech-to-Text) analyzes the voice signal and generates corresponding text. The output is the converted text data.

[0646] Step 3:

[0647] The device simultaneously sends the voice data to an emotion recognition engine to analyze the user's emotional state. The input includes a digital audio file. The emotion recognition engine (e.g., EmotionRecognition API) analyzes the tone and rate of the voice to extract emotional information. The output is the user's emotional information.

[0648] Step 4:

[0649] The device analyzes the converted text data and emotional information to determine the appropriate action. The input includes text data and emotional information. The text analysis engine understands the command content and combines it with the emotional information to determine the appropriate action. The output is the determined action.

[0650] Step 5:

[0651] The device sends an HTTP request to the server based on the determined action and emotion information. The input includes the action and emotion information. The HTTP protocol is used to send the request with appropriate parameters to the server. The output is a request that is processed on the server side.

[0652] Step 6:

[0653] Based on the request received by the server, it generates specific operation commands for the robot. The input includes action and emotion information. The logic engine on the server calculates the optimal route and operation parameters based on this data and sends the command to the robot. The output is the operation command sent to the robot.

[0654] Step 7:

[0655] The robot starts working based on the operation command from the server. The operation command from the server is included as input. The robot executes the specified task according to the command. The output is the execution status of the task.

[0656] Step 8:

[0657] The server receives feedback from the robot and sends the feedback to the smart glasses. The input includes feedback information from the robot. The server generates this information as a voice message and uses a voice synthesizer (e.g., Google Text-to-Speech) to generate the voice feedback. The output is the voice message.

[0658] Step 9:

[0659] The smart glasses provide the generated audio feedback to the user. The input includes an audio message. The smart glasses provide audio feedback to the user such as, "You seem nervous, please stay calm. We'll start transporting the parts." The output is the audio feedback provided to the user.

[0660] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0661] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0662] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0663] [Third embodiment]

[0664] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0665] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0666] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0667] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0668] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0669] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0670] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0671] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0672] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0675] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0676] The present invention relates to a system that allows a rider to perform various operations hands-free while riding a motorcycle. Specifically, it can realize functions such as navigation, music playback, phone call answering, weather information confirmation, and emergency communication using voice input. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0677] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as an audio file. The audio file is then converted into text data using a voice recognition service. The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Tokyo," for example, requires a navigation function.

[0678] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0679] The response from the server is received by the device and fed back to the user as a voice message that is easy for the rider to understand. This voice feedback is output by converting the text message into voice using a voice synthesizer built into the device. This allows the rider to obtain necessary information and perform operations using only voice, without having to take their hands off the handlebars.

[0680] For example, if a user says "Navigate to Tokyo" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Tokyo." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Tokyo and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Tokyo."

[0681] This system is designed to be safe and efficient to operate not only while driving a car, but also while driving a motorcycle, etc. By combining functions such as voice recognition, analysis, communication, and voice feedback, a wide range of operations can be performed hands-free, greatly improving the safety and convenience of riders.

[0682] The processing flow will be explained below.

[0683] Step 1:

[0684] The user issues a voice command such as "Navigate to Tokyo," and the voice is captured by a microphone attached to the device.

[0685] Step 2:

[0686] The device records the captured audio as a digital audio file, at which point the audio data is temporarily stored.

[0687] Step 3:

[0688] The device sends the recorded voice data to a voice recognition service, converts the voice data into text data, and receives the text data returned by the voice recognition service, for example, using a cloud-based voice recognition API.

[0689] Step 4:

[0690] The terminal analyzes the received text data and extracts the command content. For example, in the command "Navigate to Tokyo," the analysis means determines that "Tokyo" is the destination.

[0691] Step 5:

[0692] The terminal determines an appropriate action based on the analysis result. In this case, since "navigate" is requested, it determines to process it as a navigation action.

[0693] Step 6:

[0694] Based on the determined action, the device sends an HTTP request to the server, including necessary parameters such as destination information.

[0695] Step 7:

[0696] The server receives the HTTP request and performs the appropriate action based on the request, for example, generating route information to the destination.

[0697] Step 8:

[0698] The server sends the generated information back to the device as a response, which includes detailed route information for navigation.

[0699] Step 9:

[0700] The terminal receives the response from the server and analyzes the content in the form of a text message.

[0701] Step 10:

[0702] The device's speech synthesizer converts the text message into speech and provides feedback to the rider, such as a voice message saying, "Navigation to Tokyo is about to begin."

[0703] Step 11:

[0704] By receiving voice feedback from the device, users can obtain navigation information without taking their hands off the steering wheel.

[0705] Through the above processing steps, the present invention enables a rider to safely and efficiently use voice commands while riding a motorcycle.

[0706] Example 1

[0707] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0708] Currently, using devices such as smartphones to check navigation, play music, or make calls while riding a motorcycle is extremely dangerous. There is a demand for hands-free systems that allow these operations to be performed without using the hands while driving. There is also a need to develop systems that simultaneously improve safety and convenience.

[0709] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0710] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and determining an appropriate action, means for sending a request to an information processing device based on the determined action, means for receiving a response from the information processing device and feeding it back to the user as voice, means for using a voice recognition service, means for analyzing text using a natural language processing library, means for sending an HTTP request, and means for converting into voice using a voice synthesis device.

[0711] This allows users to navigate, play music, answer calls, check weather information, make emergency calls, and perform other operations using only voice input while riding a motorcycle, without using their hands.

[0712] "Voice input" refers to the means by which a user gives instructions to a system using voice.

[0713] "Speech Recognition Service" means external or internal software or functionality used to convert received speech into text data.

[0714] "Text conversion" refers to the process of analyzing speech input and converting it into corresponding text data.

[0715] A "natural language processing library" is a software library for analyzing text data and understanding meaning and actions.

[0716] An "HTTP request" is a communication method using a protocol for a client to request data from a server.

[0717] An "information processing device" refers to a computer or server that processes data based on a received request and generates a response.

[0718] A "voice synthesis device" is a device or software that converts text data into voice data and outputs it.

[0719] "Communication means" refers to the technology or protocol for transferring data, including wired and wireless communication methods.

[0720] "Feedback" refers to the system returning processing results or information to the user, and in this context it primarily refers to audio feedback.

[0721] "Location information provision" is a function that provides information about the user's specified destination or current location.

[0722] "Media content control" is a function for operating media content such as music playback and video playback.

[0723] "Communication response" is a function that responds to communication-related commands, such as incoming and outgoing phone calls.

[0724] "Weather information acquisition" refers to a function of acquiring weather information such as the current weather and forecast and providing it to the user.

[0725] "Safety communication" is a function for quickly communicating necessary information in an emergency.

[0726] The present invention relates to a system that allows riders to perform various operations hands-free while riding a motorcycle. This system uses voice input to realize functions such as navigation, music playback, answering calls, checking weather information, and emergency communication. The user can operate the terminal using voice input, and the terminal provides the necessary information and services through communication with a server.

[0727] The present invention consists of the following main components:

[0728] 1. Voice input means: A means for users to input commands by voice. The device is equipped with a high-sensitivity microphone that accurately captures the user's voice. Specific hardware used for this is a Bluetooth-enabled headset or smartphone.

[0729] 2. Speech recognition means: This is a means for converting received voice input into text, and uses speech recognition services such as IBM Watson and Google Cloud Speech-to-Text. These services are cloud-based and enable highly accurate speech recognition.

[0730] 3. Text analysis: This is a means of analyzing the converted text data and determining the appropriate action. The device uses natural language processing libraries such as spaCy and NLTK to analyze the text and understand the user's intent.

[0731] 4. Communication method: This is a method for sending an HTTP request to the server based on the determined action. The device uses this method to communicate the analysis results to the server. Communication is secure using HTTPS.

[0732] 5. Response processing means: This is the means for receiving responses from the server and providing audio feedback to the user. After receiving data from the server, the device uses Amazon Polly or Google Text-to-Speech to convert the text data into audio and provide it to the user.

[0733] For example, if a user says "Navigate to Tokyo" while driving, the following process will occur:

[0734] The device captures the user's voice and sends the voice data to a cloud service.

[0735] The voice recognition service converts the voice data into text data such as "Navigate to Tokyo."

[0736] The terminal analyzes this text data and determines that it is a navigation request.

[0737] The device creates an HTTP request and sends the destination information to the server.

[0738] The server generates route information to Tokyo using the Google Maps API or similar and sends that information back to the device.

[0739] The device analyzes the route information it receives and uses a voice synthesizer to provide voice feedback to the user, saying, "Navigation to Tokyo is about to begin."

[0740] An example prompt is:

[0741] "Describe a system that enables navigation functions based on voice input. For example, capture voice input of "Navigate to Tokyo" and demonstrate the flow of speech recognition, natural language processing, sending an HTTP request, processing on the server, and voice feedback."

[0742] This system allows riders to perform many operations while riding their motorcycle using only voice commands, without using their hands, greatly improving safety and convenience.The system can also be used in other driving situations and with a variety of voice commands, making it suitable for a wide range of uses.

[0743] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0744] Process flow steps and explanations

[0745] Step 1:

[0746] The terminal captures the user's voice.

[0747] Input: User's voice command (e.g., "Navigate to Tokyo")

[0748] How it works: The device's high-sensitivity microphone picks up the user's voice.

[0749] Output: Audio data

[0750] Step 2:

[0751] The device sends the captured audio to a speech recognition service.

[0752] Input: Audio data

[0753] How it works: Your device sends voice data to a cloud-based speech recognition service (e.g., IBM Watson, Google Cloud Speech-to-Text).

[0754] Output: Text data (e.g. "Navigate to Tokyo")

[0755] Step 3:

[0756] The device analyzes the received text data using a natural language processing library.

[0757] Input: Text data (e.g., "Navigate to Tokyo")

[0758] How it works: The terminal uses natural language processing libraries such as spaCy or NLTK to parse the text data and determine the intent of the command.

[0759] Output: Analysis results (e.g., navigation requests)

[0760] Step 4:

[0761] The terminal constructs an HTTP request based on the analysis results and sends it to the server.

[0762] Input: Analysis result (e.g., navigation request, destination "Tokyo")

[0763] How it works: The device constructs an HTTP request and sends it to the server, including the analysis results and additional parameters (e.g., current location, departure time).

[0764] Output: HTTP request to the server

[0765] Step 5:

[0766] The server processes the received HTTP request and retrieves the required data.

[0767] Input: HTTP request (e.g., destination "Tokyo")

[0768] How it works: The server parses the request and retrieves route information using external services such as the Google Maps API or OpenStreetMap.

[0769] Output: Route information

[0770] Step 6:

[0771] The server generates a response containing the route information and sends it back to the terminal.

[0772] Input: Route information

[0773] Operation: The server generates an HTTP response based on the obtained route information and sends it to the terminal.

[0774] Output: HTTP response to the device

[0775] Step 7:

[0776] The terminal receives the response from the server and converts it into speech using a speech synthesizer.

[0777] Input: HTTP response (e.g. route information)

[0778] How it works: The device uses Amazon Polly or Google Text-to-Speech to convert text data (route guidance) into voice data.

[0779] Output: Audio data

[0780] Step 8:

[0781] The terminal feeds back the voice data to the user.

[0782] Input: Voice data (e.g., "Start navigation to Tokyo")

[0783] How it works: The device provides audio feedback to the user through a Bluetooth headset or the built-in speaker.

[0784] Output: Audio feedback to the user

[0785] This is the flow of processing in the program for this system. At each step, specific data processing or data calculation is performed based on appropriate input data, and output data is generated as a result.

[0786] (Application example 1)

[0787] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0788] Modern delivery services require riders to access information and perform operations safely and efficiently while driving. However, many current systems require riders to manually operate smartphones and other devices, which reduces safety. Therefore, there is a need to develop a system that allows riders to navigate, check notifications, respond by voice, report situations, and make emergency calls while driving hands-free.

[0789] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0790] In this invention, the server includes means for receiving voice input, means for converting voice to text, means for analyzing the text and determining an action, means for providing functions of navigation, notification confirmation, voice response, status report, and emergency contact, and means for providing voice feedback, thereby enabling delivery riders to perform necessary operations and obtain information safely and efficiently in a hands-free manner while driving.

[0791] A "means for receiving audio input" is a microphone or other sound capturing device that detects and converts sound into a digital signal.

[0792] "Means for converting speech to text" means speech recognition software or algorithms for converting speech to text data.

[0793] The "means for determining an action" is an analysis algorithm for analyzing the converted text data and identifying the operation intended by the user.

[0794] The "means for sending a request to a server" is a communication interface for sending information to a server via the Internet based on the determined action.

[0795] The "means for receiving a response from the server and providing feedback to the user as voice" is a system for outputting information received from the server as voice using voice synthesis technology.

[0796] "Navigation" is a function that provides route guidance to a destination.

[0797] "Notification Check" is a function that notifies the user of new notifications and messages.

[0798] "Voice response" is a function that responds to voice input with voice.

[0799] "Status Report" is a function that provides audio reports on the progress of deliveries and other tasks.

[0800] "Emergency contact" is a function that allows you to make contact quickly when a problem occurs.

[0801] "Communication means" refers to infrastructure such as internet connections and mobile networks for sending and receiving data.

[0802] A "speech synthesizer" is hardware or software that converts text data into speech.

[0803] This invention relates to a system that uses voice input to enable delivery service riders to perform various operations hands-free while driving. Specifically, the system can provide functions such as navigation, notification confirmation, voice response, status reports, and emergency contact using voice input. This system comprises a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0804] A user issues a voice command, such as "Navigate to Shinjuku." This voice is captured by a device (e.g., a smartphone or smart glasses) and recorded as an audio file. The audio file is then converted into text data using a speech recognition service (e.g., Azure Speech to Text API). The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Shinjuku" requires a navigation function.

[0805] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0806] The response from the server is received by the device and fed back to the user as an easy-to-understand voice message. This voice feedback is output by converting the text message into voice using the Azure Text to Speech API. This means that riders can obtain necessary information and perform operations using only voice without having to operate their mobile device while driving.

[0807] As a specific example, if a user says "Navigate to Shinjuku" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Shinjuku." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Shinjuku and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Shinjuku."

[0808] Other prompts include "Check notifications," "Respond to customer, on my way," "Report status," and "Urgent," helping delivery riders navigate safely and efficiently while driving.

[0809] The hardware used in this system configuration is end-user devices such as smartphones and smart glasses. The software uses services such as Azure Speech to Text API, Azure Text to Speech API, Google Maps API, and Twilio API. The combination of this hardware and software improves the user experience and ensures security.

[0810] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0811] Step 1:

[0812] The user issues a voice command, for example, "Navigate to Shinjuku." This voice input is the starting point of the process.

[0813] Step 2:

[0814] The device uses a microphone to capture the voice and collects the voice data as a digital signal. This is the process of acquiring voice input. The input is the user's voice, and the output is digital voice data.

[0815] Step 3:

[0816] The device sends voice data to a speech recognition service (for example, Azure Speech to Text API) that converts the voice data into text data. The input is digital voice data and the output is text data. The speech recognition service analyzes the input voice and generates corresponding text.

[0817] Step 4:

[0818] The terminal analyzes the converted text data. Specifically, it analyzes the text "Navigate to Shinjuku" and determines that this command requests a navigation function. The input is text data, and the output is the analysis result (action).

[0819] Step 5:

[0820] The terminal sends an HTTP request to the server based on the analysis results. This request includes the analyzed destination information and other necessary parameters. The input is the analysis results, and the output is an HTTP request to the server.

[0821] Step 6:

[0822] The server receives the HTTP request and performs processing according to the request. In the case of navigation, it generates route information to the destination. The input is the HTTP request, and the output is the server's response, such as route information.

[0823] Step 7:

[0824] The server returns the generated route information to the terminal. The input is the route information, and the output is an HTTP response.

[0825] Step 8:

[0826] The device receives the response from the server and converts this information into a voice message using a speech synthesizer (for example, Azure Text to Speech API). The input is the server response and the output is the voice message. The speech synthesizer converts the text message into speech.

[0827] Step 9:

[0828] The terminal provides a voice message as feedback to the user. For example, it may say, "Starting navigation to Shinjuku." The input is the voice message, and the output is the voice feedback to the user.

[0829] Through this process, delivery riders can receive safe navigation to their destination while driving.

[0830] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0831] This invention provides a system that allows a user to perform various operations hands-free while riding a motorcycle, by combining it with an emotion engine that recognizes the user's emotions, and provides more appropriate and personalized feedback and actions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, the emotion engine, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[0832] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as a digital audio file in real time. The device then sends the recorded voice data to a voice recognition service, which converts the voice data into text data. At the same time, the emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[0833] The converted text data and emotional information are analyzed by an analysis means within the device. For example, in the command "Navigate to Tokyo," "Tokyo" is determined to be the destination, and the user's emotional state, such as whether they are nervous or calm, is recognized. Based on the results of this analysis, an appropriate action is determined. In this case, since the "navigate" function is requested, it is determined that the action should be processed as a navigation action.

[0834] Furthermore, the emotion engine's information is taken into account to adjust the expression and content of the feedback. For example, if the user is nervous, navigation instructions will be provided in a calmer voice tone, and conversely, if the user is calm, instructions will be provided in a normal tone.

[0835] Based on the determined action, the device sends an HTTP request to the server. The request includes necessary parameters such as destination information and emotional state. When the server receives the request, it processes it according to its content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0836] The response from the server is received by the device and fed back to the user. A speech synthesizer converts the text message into speech and provides it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[0837] For example, if a user says "Navigate to Tokyo" while riding, the voice is captured by the device, and the speech recognition system converts it into text data, while the emotion engine simultaneously recognizes the sense of tension. The analysis result is sent to the server as a navigation request, and the server provides route information to Tokyo. Finally, the device provides feedback to the user with an adjusted voice message, providing a safer and more comfortable riding experience.

[0838] This system can be operated safely and efficiently not only when driving a car but also when driving a motorcycle. In particular, by combining it with an emotion engine, personalized interactions are possible, greatly improving the safety and convenience of riders.

[0839] The processing flow will be explained below.

[0840] Step 1:

[0841] The user issues a voice command such as "Navigate to Tokyo," which is captured using a microphone on the device.

[0842] Step 2:

[0843] The device records the captured audio as a digital audio file, which is temporarily stored on the device's storage.

[0844] Step 3:

[0845] The device sends the recorded voice data to a voice recognition service, where it is converted into text data. At the same time, an emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[0846] Step 4:

[0847] The device receives the text data and emotion information and analyzes the command content and the user's emotion using an analysis means. Specifically, the device recognizes that the command "Navigate to Tokyo" means the destination "Tokyo" and that the user's emotion is, for example, nervous.

[0848] Step 5:

[0849] The terminal determines an appropriate action based on the analysis result. In this case, since the "navigate" function is requested, it determines to process it as a navigation action.

[0850] Step 6:

[0851] The device sends an HTTP request to the server according to the determined action, including necessary parameters such as destination information and user emotion information.

[0852] Step 7:

[0853] The server receives the HTTP request from the device and processes it accordingly. Specifically, it generates route information to the destination and formats and prepares the message.

[0854] Step 8:

[0855] The server sends the generated route information and an appropriate message back to the device as a response, which includes navigation information and an emotion-based message.

[0856] Step 9:

[0857] The terminal receives the response from the server and analyzes its contents to provide appropriate feedback to the user. In this case, a text message is analyzed.

[0858] Step 10:

[0859] The device's speech synthesizer converts the text message into speech, generating a voice message that reflects the results of the emotion engine. For example, the message might say, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[0860] Step 11:

[0861] Users can receive voice feedback from the device and obtain real-time navigation information without taking their hands off the steering wheel.

[0862] Through these steps, the system provides safe and efficient operation for the rider while riding the motorcycle, and in particular, by using an emotion engine, it enables personalized feedback according to the user's emotional state.

[0863] Example 2

[0864] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0865] When users operate hands-free while driving a motorcycle, they are faced with the challenge of being unable to operate the system safely and efficiently. Furthermore, conventional hands-free systems lack personalized feedback that takes into account the user's emotions, limiting the user experience. This increases stress and anxiety for users while driving, and creates the problem of insufficient safety and convenience.

[0866] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0867] In this invention, the server includes means for receiving a voice input, means for converting the received voice input into text, means for extracting emotional information from the converted text and voice information, means for analyzing the converted text and the extracted emotional information to determine an appropriate action, means for sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice based on the emotional information. This makes it possible to provide appropriate and personalized feedback and actions that take into account the emotional state of the user.

[0868] "Voice input" refers to commands or instructions spoken by a user received as digital data through a microphone or other audio collection device.

[0869] "Text" refers to a string of characters, such as alphabets or kanji, that has been converted from voice input, and is information in a format that can be analyzed on a device such as a computer.

[0870] "Emotional information" is information extracted by analyzing the user's emotional state (e.g., tension, calmness, anxiety, etc.) from voice data.

[0871] "Analysis" is a process for deriving appropriate actions based on the received text data and emotional information.

[0872] "Actions" refer to the operations or reactions the system performs based on the user's voice commands, such as navigation, music playback, and answering calls.

[0873] A "request" is an inquiry or request sent from a terminal to a server, and is executed in the form of an HTTP request or the like.

[0874] A "response" is a reply or answer sent from a server to a terminal, and is data that includes information or instructions corresponding to a request.

[0875] A "speech recognition service" is any cloud-based or on-premise technology or software that accepts voice data as input and converts it into text data.

[0876] A "speech synthesizer" is a device that converts text information into speech and provides speech output to the user.

[0877] This invention is a system that allows a user to perform various operations hands-free while riding a motorcycle, and provides personalized feedback and actions by combining it with an emotion engine that recognizes the user's emotions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for extracting emotion information from the text and voice data, a means for analyzing the text and emotion information and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[0878] First, when a user issues a voice command such as "navigate to a destination," the device captures this voice and records it as a digital audio file via the microphone. This is real-time voice capture.

[0879] Next, the terminal transmits the recorded voice data to a voice recognition service (e.g., voice recognition software) to convert the voice data into text data, and also uses an emotion engine (e.g., emotion analysis software) to analyze the user's emotions from the voice data and extract emotion information.

[0880] The converted text data and emotional information are then analyzed by an analysis unit within the device. For example, in the command "Navigate to destination," the device determines that "destination" is the destination and simultaneously recognizes the user's emotional state, such as whether they are nervous or calm.

[0881] Based on the analysis results, the system determines the appropriate action. In this case, it determines that the "navigation" function is requested. The expression and content of the feedback are adjusted taking into account the information from the emotion engine. If the user is nervous, the device will provide navigation guidance in a calm voice tone, and if the user is calm, it will provide guidance in a normal tone.

[0882] Furthermore, based on the determined action, the device sends an HTTP request to the server. This request includes necessary parameters such as destination information and emotional state. When the server receives this request, it performs processing according to the content and generates a corresponding response. For example, if it is a navigation request, it calculates and generates route information to the destination.

[0883] The response from the server is received by the device, which then uses a speech synthesizer (e.g., speech synthesis software) to convert the text message into speech and provide it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to your destination will begin."

[0884] As a concrete example, if a user says "Navigate to Tokyo" while driving, the device captures the voice and converts it into text data using a speech recognition service, while the emotion engine simultaneously recognizes the user's nervousness. Based on the analysis results, the device sends this as a navigation request to the server, which then provides route information to Tokyo. Using a speech synthesizer, the device then provides the user with an adjusted voice message such as "You seem nervous, but please remain calm while driving. Navigation to Tokyo will begin," providing a safer and more comfortable riding experience.

[0885] Examples of prompts for generative AI models include:

[0886] > "Design a system that allows a user to operate a motorcycle hands-free while driving. The system combines voice input with an emotion engine to provide appropriate feedback based on the user's emotional state. Specifically, if the user is nervous, the system will provide navigation information in a calm tone, and if the user is calm, the system will provide information in a normal tone."

[0887] The system's unique feature is that it significantly improves the user's safety and convenience while riding a motorcycle, and by utilizing an emotion engine, personalized interactions based on the user's emotions are possible.

[0888] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0889] Program processing flow

[0890] Step 1: Capture voice commands

[0891] The user issues the voice command "Navigate to Tokyo."

[0892] Input: Voice command

[0893] How it works: The device captures the user's voice commands through the microphone and records them as digital audio files in real time.

[0894] Output: Digital audio file

[0895] Step 2: Speech to text

[0896] The device transmits the recorded digital audio file to a speech recognition service.

[0897] Input: Digital audio file

[0898] How it works: A speech recognition service converts speech data into text data. For example, speech recognition software is used to generate the text "Navigate to Tokyo."

[0899] Output: Text data

[0900] Step 3: Extracting emotional information

[0901] The device sends the digital audio file to the emotion engine.

[0902] Input: Digital audio file

[0903] How it works: The emotion engine analyzes the user's emotions from the voice data and extracts emotional information. For example, it uses emotion analysis software to determine that the user is "nervous."

[0904] Output: Emotional information

[0905] Step 4: Analyzing text data and sentiment information

[0906] The analysis unit in the device analyzes the text data and emotional information.

[0907] Input: Text data, emotion information

[0908] Action: The analysis unit analyzes the text "Navigate to Tokyo" and the emotional information "I'm nervous" and determines the appropriate action. Specifically, it determines that "Tokyo" is the destination and that the "Navigate" function is required.

[0909] Output: Analysis results (actions, destination information, emotion information)

[0910] Step 5: Sending a request to the server

[0911] The terminal sends an HTTP request to the server based on the analysis results.

[0912] Input: Analysis results

[0913] How it works: The HTTP request contains destination and emotion information. By sending this request, the server is asked to provide navigation information.

[0914] Output: HTTP request

[0915] Step 6: Server-side processing

[0916] The server receives the HTTP request and performs the corresponding processing.

[0917] Input: HTTP request

[0918] Operation: The server analyzes the request (destination information, emotion information) and generates navigation information. For example, it uses route calculation software to calculate the route to Tokyo.

[0919] Output: Navigation information

[0920] Step 7: Receiving responses and coordinating feedback

[0921] The terminal receives the response from the server.

[0922] Input: Navigation information

[0923] How it works: The device uses a speech synthesizer to convert navigation information into a voice message to convey to the user. The device adjusts the tone of the voice message based on emotional information. For example, it generates a message like, "You seem nervous, but please stay calm and drive. We'll start navigating to Tokyo."

[0924] Output: Modified voice message

[0925] Step 8: Provide feedback to users

[0926] The terminal finally provides feedback to the user as a voice message.

[0927] Input: Modified voice message

[0928] Operation: The terminal outputs a voice message to the user through a voice synthesizer.

[0929] Output: The user receives the adjusted navigation instructions.

[0930] Specific examples

[0931] When a user says "Navigate to Tokyo," the speech is captured by the device and converted into text by the speech recognition service. At the same time, the speech data is analyzed by the emotion engine, which recognizes that the user is nervous. Based on the results, the analysis unit sends an HTTP request to the server as a navigation action, requesting directions to Tokyo. The server receives the request and provides route information to Tokyo, and the device provides the user with a voice message tailored based on the emotion information.

[0932] (Application example 2)

[0933] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0934] Existing technologies include systems that use voice commands to perform various operations, but these systems are unable to recognize the user's emotional state and provide appropriate feedback based on that. This results in a uniform user experience, which does not improve safety or convenience, especially in emergencies or stressful situations. Furthermore, there is a lack of a hands-free means of operating robots to improve work efficiency and safety in factories.

[0935] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and emotional information to determine an appropriate action, means for combining an emotion engine that recognizes the user's emotion and sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice. This makes it possible to provide personalized feedback and actions according to the user's emotional state, greatly improving work efficiency and safety in factories.

[0936] "Means for receiving voice input" refers to a function for capturing voice commands issued by a user through a device such as a microphone.

[0937] A "means for converting voice input to text" is a function that allows for analyzing captured voice data and converting it into text format.

[0938] "Means for analyzing the converted text and determining the appropriate action" refers to a function that understands and analyzes the voice command converted into text and selects the appropriate action based on the results.

[0939] The "emotion engine" is a function for recognizing and analyzing the user's emotional state from voice data.

[0940] The "means for sending a request to the server" is a function for sending a request to the server for an appropriate action based on the analyzed text data and emotion information.

[0941] The "means for receiving a response from the server and providing feedback to the user as audio" is a function for providing feedback to the user by outputting the data received from the server as audio.

[0942] "Navigation, music playback, call answering, weather information check, emergency communication, robot operation" are the various types of actions that the system can perform in response to a user's voice commands.

[0943] A "prompt sentence using a generative AI model" is an input sentence used by a generative AI technology to obtain a desired output.

[0944] This invention relates to a system that enables a worker working in a factory to operate a robot hands-free using smart glasses. The system includes means for receiving voice input, means for converting the voice input into text, means for analyzing the converted text and emotion information to determine an appropriate action, an emotion engine for recognizing the emotion of the user, means for sending a request to a server based on the determined action and emotion information, and means for receiving a response from the server and providing feedback to the user as voice.

[0945] As a concrete example of the system, a worker puts on smart glasses and issues a command by voice, such as "Transport this part." The smart glasses' microphone captures the voice command and generates a digital audio file in real time. The smart glasses then send the recorded voice data to a voice recognition service (e.g., Google Speech-to-Text) to convert the voice data into text data. At the same time, an emotion engine (e.g., EmotionRecognition API) analyzes the user's emotional state and extracts emotional information.

[0946] The converted text data and emotion information are further analyzed by the analysis means in the smart glasses. For example, for a command such as "transport this part," the emotion engine determines that there is a transport command and recognizes that the user is nervous. Based on this result, an appropriate action is determined. If the user is nervous, commands are sent to the robot to operate as smoothly and safely as possible.

[0947] Based on the determined action, the smart glasses send an HTTP request to the server. The request includes parameters related to the transport command and the emotional state. The server receives the request and generates specific operational commands for the robot, such as calculating and returning a smooth transport route to the destination.

[0948] The response from the server is received by the smart glasses. The content is read as a text message and audio feedback is provided to the user using a speech synthesizer (e.g., Google Text-to-Speech). Feedback that takes into account the user's state of tension is provided, such as "You seem nervous, but please calm down. We'll start transporting the parts."

[0949] Hardware and software used

[0950] Smart glasses (e.g., generic name)

[0951] microphone

[0952] Connected robot arm or mobile robot

[0953] Speech recognition services (e.g., Google Speech-to-Text)

[0954] Emotion recognition engine (e.g., EmotionRecognition API)

[0955] Speech synthesizers (e.g., Google Text-to-Speech)

[0956] Prompt Sentence Examples

[0957] Prompt: Recognizes user voice commands and their emotions

[0958] Input: Audio data text and its audio file

[0959] Output: Command converted to text and recognized emotion information (e.g., carrying, stressed)

[0960] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0961] Step 1:

[0962] A user issues a voice command through the smart glasses. The input includes the user's speech. A microphone on the smart glasses captures the voice command and records it as a digital audio file. The output is a digital audio file.

[0963] Step 2:

[0964] The device sends the captured voice data to a speech recognition service, which converts the voice data into text data. The input includes a digital audio file. The speech recognition service used (e.g., Google Speech-to-Text) analyzes the voice signal and generates corresponding text. The output is the converted text data.

[0965] Step 3:

[0966] The device simultaneously sends the voice data to an emotion recognition engine to analyze the user's emotional state. The input includes a digital audio file. The emotion recognition engine (e.g., EmotionRecognition API) analyzes the tone and rate of the voice to extract emotional information. The output is the user's emotional information.

[0967] Step 4:

[0968] The device analyzes the converted text data and emotional information to determine the appropriate action. The input includes text data and emotional information. The text analysis engine understands the command content and combines it with the emotional information to determine the appropriate action. The output is the determined action.

[0969] Step 5:

[0970] The device sends an HTTP request to the server based on the determined action and emotion information. The input includes the action and emotion information. The HTTP protocol is used to send the request with appropriate parameters to the server. The output is a request that is processed on the server side.

[0971] Step 6:

[0972] Based on the request received by the server, it generates specific operation commands for the robot. The input includes action and emotion information. The logic engine on the server calculates the optimal route and operation parameters based on this data and sends the command to the robot. The output is the operation command sent to the robot.

[0973] Step 7:

[0974] The robot starts working based on the operation command from the server. The operation command from the server is included as input. The robot executes the specified task according to the command. The output is the execution status of the task.

[0975] Step 8:

[0976] The server receives feedback from the robot and sends the feedback to the smart glasses. The input includes feedback information from the robot. The server generates this information as a voice message and uses a voice synthesizer (e.g., Google Text-to-Speech) to generate the voice feedback. The output is the voice message.

[0977] Step 9:

[0978] The smart glasses provide the generated audio feedback to the user. The input includes an audio message. The smart glasses provide audio feedback to the user such as, "You seem nervous, please stay calm. We'll start transporting the parts." The output is the audio feedback provided to the user.

[0979] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0980] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0981] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0982] [Fourth embodiment]

[0983] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0984] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0985] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0986] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0987] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0988] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0989] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0990] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0991] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0992] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0993] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0994] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0995] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0996] The present invention relates to a system that allows a rider to perform various operations hands-free while riding a motorcycle. Specifically, it can realize functions such as navigation, music playback, phone call answering, weather information confirmation, and emergency communication using voice input. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[0997] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as an audio file. The audio file is then converted into text data using a voice recognition service. The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Tokyo," for example, requires a navigation function.

[0998] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[0999] The response from the server is received by the device and fed back to the user as a voice message that is easy for the rider to understand. This voice feedback is output by converting the text message into voice using a voice synthesizer built into the device. This allows the rider to obtain necessary information and perform operations using only voice, without having to take their hands off the handlebars.

[1000] For example, if a user says "Navigate to Tokyo" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Tokyo." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Tokyo and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Tokyo."

[1001] This system is designed to be safe and efficient to operate not only while driving a car, but also while driving a motorcycle, etc. By combining functions such as voice recognition, analysis, communication, and voice feedback, a wide range of operations can be performed hands-free, greatly improving the safety and convenience of riders.

[1002] The processing flow will be explained below.

[1003] Step 1:

[1004] The user issues a voice command such as "Navigate to Tokyo," and the voice is captured by a microphone attached to the device.

[1005] Step 2:

[1006] The device records the captured audio as a digital audio file, at which point the audio data is temporarily stored.

[1007] Step 3:

[1008] The device sends the recorded voice data to a voice recognition service, converts the voice data into text data, and receives the text data returned by the voice recognition service, for example, using a cloud-based voice recognition API.

[1009] Step 4:

[1010] The terminal analyzes the received text data and extracts the command content. For example, in the command "Navigate to Tokyo," the analysis means determines that "Tokyo" is the destination.

[1011] Step 5:

[1012] The terminal determines an appropriate action based on the analysis result. In this case, since "navigate" is requested, it determines to process it as a navigation action.

[1013] Step 6:

[1014] Based on the determined action, the device sends an HTTP request to the server, including necessary parameters such as destination information.

[1015] Step 7:

[1016] The server receives the HTTP request and performs the appropriate action based on the request, for example, generating route information to the destination.

[1017] Step 8:

[1018] The server sends the generated information back to the device as a response, which includes detailed route information for navigation.

[1019] Step 9:

[1020] The terminal receives the response from the server and analyzes the content in the form of a text message.

[1021] Step 10:

[1022] The device's speech synthesizer converts the text message into speech and provides feedback to the rider, such as a voice message saying, "Navigation to Tokyo is about to begin."

[1023] Step 11:

[1024] By receiving voice feedback from the device, users can obtain navigation information without taking their hands off the steering wheel.

[1025] Through the above processing steps, the present invention enables a rider to safely and efficiently use voice commands while riding a motorcycle.

[1026] Example 1

[1027] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1028] Currently, using devices such as smartphones to check navigation, play music, or make calls while riding a motorcycle is extremely dangerous. There is a demand for hands-free systems that allow these operations to be performed without using the hands while driving. There is also a need to develop systems that simultaneously improve safety and convenience.

[1029] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1030] In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and determining an appropriate action, means for sending a request to an information processing device based on the determined action, means for receiving a response from the information processing device and feeding it back to the user as voice, means for using a voice recognition service, means for analyzing text using a natural language processing library, means for sending an HTTP request, and means for converting into voice using a voice synthesis device.

[1031] This allows users to navigate, play music, answer calls, check weather information, make emergency calls, and perform other operations using only voice input while riding a motorcycle, without using their hands.

[1032] "Voice input" refers to the means by which a user gives instructions to a system using voice.

[1033] "Speech Recognition Service" means external or internal software or functionality used to convert received speech into text data.

[1034] "Text conversion" refers to the process of analyzing speech input and converting it into corresponding text data.

[1035] A "natural language processing library" is a software library for analyzing text data and understanding meaning and actions.

[1036] An "HTTP request" is a communication method using a protocol for a client to request data from a server.

[1037] An "information processing device" refers to a computer or server that processes data based on a received request and generates a response.

[1038] A "voice synthesis device" is a device or software that converts text data into voice data and outputs it.

[1039] "Communication means" refers to the technology or protocol for transferring data, including wired and wireless communication methods.

[1040] "Feedback" refers to the system returning processing results or information to the user, and in this context it primarily refers to audio feedback.

[1041] "Location information provision" is a function that provides information about the user's specified destination or current location.

[1042] "Media content control" is a function for operating media content such as music playback and video playback.

[1043] "Communication response" is a function that responds to communication-related commands, such as incoming and outgoing phone calls.

[1044] "Weather information acquisition" refers to a function of acquiring weather information such as the current weather and forecast and providing it to the user.

[1045] "Safety communication" is a function for quickly communicating necessary information in an emergency.

[1046] The present invention relates to a system that allows riders to perform various operations hands-free while riding a motorcycle. This system uses voice input to realize functions such as navigation, music playback, answering calls, checking weather information, and emergency communication. The user can operate the terminal using voice input, and the terminal provides the necessary information and services through communication with a server.

[1047] The present invention consists of the following main components:

[1048] 1. Voice input means: A means for users to input commands by voice. The device is equipped with a high-sensitivity microphone that accurately captures the user's voice. Specific hardware used for this is a Bluetooth-enabled headset or smartphone.

[1049] 2. Speech recognition means: This is a means for converting received voice input into text, and uses speech recognition services such as IBM Watson and Google Cloud Speech-to-Text. These services are cloud-based and enable highly accurate speech recognition.

[1050] 3. Text analysis: This is a means of analyzing the converted text data and determining the appropriate action. The device uses natural language processing libraries such as spaCy and NLTK to analyze the text and understand the user's intent.

[1051] 4. Communication method: This is a method for sending an HTTP request to the server based on the determined action. The device uses this method to communicate the analysis results to the server. Communication is secure using HTTPS.

[1052] 5. Response processing means: This is the means for receiving responses from the server and providing audio feedback to the user. After receiving data from the server, the device uses Amazon Polly or Google Text-to-Speech to convert the text data into audio and provide it to the user.

[1053] For example, if a user says "Navigate to Tokyo" while driving, the following process will occur:

[1054] The device captures the user's voice and sends the voice data to a cloud service.

[1055] The voice recognition service converts the voice data into text data such as "Navigate to Tokyo."

[1056] The terminal analyzes this text data and determines that it is a navigation request.

[1057] The device creates an HTTP request and sends the destination information to the server.

[1058] The server generates route information to Tokyo using the Google Maps API or similar and sends that information back to the device.

[1059] The device analyzes the route information it receives and uses a voice synthesizer to provide voice feedback to the user, saying, "Navigation to Tokyo is about to begin."

[1060] An example prompt is:

[1061] "Describe a system that enables navigation functions based on voice input. For example, capture voice input of "Navigate to Tokyo" and demonstrate the flow of speech recognition, natural language processing, sending an HTTP request, processing on the server, and voice feedback."

[1062] This system allows riders to perform many operations while riding their motorcycle using only voice commands, without using their hands, greatly improving safety and convenience.The system can also be used in other driving situations and with a variety of voice commands, making it suitable for a wide range of uses.

[1063] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1064] Process flow steps and explanations

[1065] Step 1:

[1066] The terminal captures the user's voice.

[1067] Input: User's voice command (e.g., "Navigate to Tokyo")

[1068] How it works: The device's high-sensitivity microphone picks up the user's voice.

[1069] Output: Audio data

[1070] Step 2:

[1071] The device sends the captured audio to a speech recognition service.

[1072] Input: Audio data

[1073] How it works: Your device sends voice data to a cloud-based speech recognition service (e.g., IBM Watson, Google Cloud Speech-to-Text).

[1074] Output: Text data (e.g. "Navigate to Tokyo")

[1075] Step 3:

[1076] The device analyzes the received text data using a natural language processing library.

[1077] Input: Text data (e.g., "Navigate to Tokyo")

[1078] How it works: The terminal uses natural language processing libraries such as spaCy or NLTK to parse the text data and determine the intent of the command.

[1079] Output: Analysis results (e.g., navigation requests)

[1080] Step 4:

[1081] The terminal constructs an HTTP request based on the analysis results and sends it to the server.

[1082] Input: Analysis result (e.g., navigation request, destination "Tokyo")

[1083] How it works: The device constructs an HTTP request and sends it to the server, including the analysis results and additional parameters (e.g., current location, departure time).

[1084] Output: HTTP request to the server

[1085] Step 5:

[1086] The server processes the received HTTP request and retrieves the required data.

[1087] Input: HTTP request (e.g., destination "Tokyo")

[1088] How it works: The server parses the request and retrieves route information using external services such as the Google Maps API or OpenStreetMap.

[1089] Output: Route information

[1090] Step 6:

[1091] The server generates a response containing the route information and sends it back to the terminal.

[1092] Input: Route information

[1093] Operation: The server generates an HTTP response based on the obtained route information and sends it to the terminal.

[1094] Output: HTTP response to the device

[1095] Step 7:

[1096] The terminal receives the response from the server and converts it into speech using a speech synthesizer.

[1097] Input: HTTP response (e.g. route information)

[1098] How it works: The device uses Amazon Polly or Google Text-to-Speech to convert text data (route guidance) into voice data.

[1099] Output: Audio data

[1100] Step 8:

[1101] The terminal feeds back the voice data to the user.

[1102] Input: Voice data (e.g., "Start navigation to Tokyo")

[1103] How it works: The device provides audio feedback to the user through a Bluetooth headset or the built-in speaker.

[1104] Output: Audio feedback to the user

[1105] This is the flow of processing in the program for this system. At each step, specific data processing or data calculation is performed based on appropriate input data, and output data is generated as a result.

[1106] (Application example 1)

[1107] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1108] Modern delivery services require riders to access information and perform operations safely and efficiently while driving. However, many current systems require riders to manually operate smartphones and other devices, which reduces safety. Therefore, there is a need to develop a system that allows riders to navigate, check notifications, respond by voice, report situations, and make emergency calls while driving hands-free.

[1109] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1110] In this invention, the server includes means for receiving voice input, means for converting voice to text, means for analyzing the text and determining an action, means for providing functions of navigation, notification confirmation, voice response, status report, and emergency contact, and means for providing voice feedback, thereby enabling delivery riders to perform necessary operations and obtain information safely and efficiently in a hands-free manner while driving.

[1111] A "means for receiving audio input" is a microphone or other sound capturing device that detects and converts sound into a digital signal.

[1112] "Means for converting speech to text" means speech recognition software or algorithms for converting speech to text data.

[1113] The "means for determining an action" is an analysis algorithm for analyzing the converted text data and identifying the operation intended by the user.

[1114] The "means for sending a request to a server" is a communication interface for sending information to a server via the Internet based on the determined action.

[1115] The "means for receiving a response from the server and providing feedback to the user as voice" is a system for outputting information received from the server as voice using voice synthesis technology.

[1116] "Navigation" is a function that provides route guidance to a destination.

[1117] "Notification Check" is a function that notifies the user of new notifications and messages.

[1118] "Voice response" is a function that responds to voice input with voice.

[1119] "Status Report" is a function that provides audio reports on the progress of deliveries and other tasks.

[1120] "Emergency contact" is a function that allows you to make contact quickly when a problem occurs.

[1121] "Communication means" refers to infrastructure such as internet connections and mobile networks for sending and receiving data.

[1122] A "speech synthesizer" is hardware or software that converts text data into speech.

[1123] This invention relates to a system that uses voice input to enable delivery service riders to perform various operations hands-free while driving. Specifically, the system can provide functions such as navigation, notification confirmation, voice response, status reports, and emergency contact using voice input. This system comprises a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for providing feedback to the user of the response from the server as voice.

[1124] A user issues a voice command, such as "Navigate to Shinjuku." This voice is captured by a device (e.g., a smartphone or smart glasses) and recorded as an audio file. The audio file is then converted into text data using a speech recognition service (e.g., Azure Speech to Text API). The converted text data is analyzed by an analysis means within the device, and it is determined that the command "Navigate to Shinjuku" requires a navigation function.

[1125] Based on the analysis results, the device sends an HTTP request to the server. The request includes the analyzed destination information and other necessary parameters. When the server receives the request, it processes it according to the content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[1126] The response from the server is received by the device and fed back to the user as an easy-to-understand voice message. This voice feedback is output by converting the text message into voice using the Azure Text to Speech API. This means that riders can obtain necessary information and perform operations using only voice without having to operate their mobile device while driving.

[1127] As a specific example, if a user says "Navigate to Shinjuku" while driving, the speech is captured by the device, and the speech recognition system converts it into text "Navigate to Shinjuku." The device analyzes this text and sends it to the server as a navigation request. The server generates route information to Shinjuku and sends it back to the device. Finally, the device provides voice feedback saying, "Starting navigation to Shinjuku."

[1128] Other prompts include "Check notifications," "Respond to customer, on my way," "Report status," and "Urgent," helping delivery riders navigate safely and efficiently while driving.

[1129] The hardware used in this system configuration is end-user devices such as smartphones and smart glasses. The software uses services such as Azure Speech to Text API, Azure Text to Speech API, Google Maps API, and Twilio API. The combination of this hardware and software improves the user experience and ensures security.

[1130] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1131] Step 1:

[1132] The user issues a voice command, for example, "Navigate to Shinjuku." This voice input is the starting point of the process.

[1133] Step 2:

[1134] The device uses a microphone to capture the voice and collects the voice data as a digital signal. This is the process of acquiring voice input. The input is the user's voice, and the output is digital voice data.

[1135] Step 3:

[1136] The device sends voice data to a speech recognition service (for example, Azure Speech to Text API) that converts the voice data into text data. The input is digital voice data and the output is text data. The speech recognition service analyzes the input voice and generates corresponding text.

[1137] Step 4:

[1138] The terminal analyzes the converted text data. Specifically, it analyzes the text "Navigate to Shinjuku" and determines that this command requests a navigation function. The input is text data, and the output is the analysis result (action).

[1139] Step 5:

[1140] The terminal sends an HTTP request to the server based on the analysis results. This request includes the analyzed destination information and other necessary parameters. The input is the analysis results, and the output is an HTTP request to the server.

[1141] Step 6:

[1142] The server receives the HTTP request and performs processing according to the request. In the case of navigation, it generates route information to the destination. The input is the HTTP request, and the output is the server's response, such as route information.

[1143] Step 7:

[1144] The server returns the generated route information to the terminal. The input is the route information, and the output is an HTTP response.

[1145] Step 8:

[1146] The device receives the response from the server and converts this information into a voice message using a speech synthesizer (for example, Azure Text to Speech API). The input is the server response and the output is the voice message. The speech synthesizer converts the text message into speech.

[1147] Step 9:

[1148] The terminal provides a voice message as feedback to the user. For example, it may say, "Starting navigation to Shinjuku." The input is the voice message, and the output is the voice feedback to the user.

[1149] Through this process, delivery riders can receive safe navigation to their destination while driving.

[1150] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1151] This invention provides a system that allows a user to perform various operations hands-free while riding a motorcycle, by combining it with an emotion engine that recognizes the user's emotions, and provides more appropriate and personalized feedback and actions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for analyzing the text and determining an appropriate action, the emotion engine, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[1152] First, the user issues a voice command, such as "Navigate to Tokyo." This voice is captured by the device and recorded as a digital audio file in real time. The device then sends the recorded voice data to a voice recognition service, which converts the voice data into text data. At the same time, the emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[1153] The converted text data and emotional information are analyzed by an analysis means within the device. For example, in the command "Navigate to Tokyo," "Tokyo" is determined to be the destination, and the user's emotional state, such as whether they are nervous or calm, is recognized. Based on the results of this analysis, an appropriate action is determined. In this case, since the "navigate" function is requested, it is determined that the action should be processed as a navigation action.

[1154] Furthermore, the emotion engine's information is taken into account to adjust the expression and content of the feedback. For example, if the user is nervous, navigation instructions will be provided in a calmer voice tone, and conversely, if the user is calm, instructions will be provided in a normal tone.

[1155] Based on the determined action, the device sends an HTTP request to the server. The request includes necessary parameters such as destination information and emotional state. When the server receives the request, it processes it according to its content and generates a corresponding response. For example, in the case of navigation, it returns route information to the destination.

[1156] The response from the server is received by the device and fed back to the user. A speech synthesizer converts the text message into speech and provides it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[1157] For example, if a user says "Navigate to Tokyo" while riding, the voice is captured by the device, and the speech recognition system converts it into text data, while the emotion engine simultaneously recognizes the sense of tension. The analysis result is sent to the server as a navigation request, and the server provides route information to Tokyo. Finally, the device provides feedback to the user with an adjusted voice message, providing a safer and more comfortable riding experience.

[1158] This system can be operated safely and efficiently not only when driving a car but also when driving a motorcycle. In particular, by combining it with an emotion engine, personalized interactions are possible, greatly improving the safety and convenience of riders.

[1159] The processing flow will be explained below.

[1160] Step 1:

[1161] The user issues a voice command such as "Navigate to Tokyo," which is captured using a microphone on the device.

[1162] Step 2:

[1163] The device records the captured audio as a digital audio file, which is temporarily stored on the device's storage.

[1164] Step 3:

[1165] The device sends the recorded voice data to a voice recognition service, where it is converted into text data. At the same time, an emotion engine analyzes the user's emotions from the voice data and extracts emotional information.

[1166] Step 4:

[1167] The device receives the text data and emotion information and analyzes the command content and the user's emotion using an analysis means. Specifically, the device recognizes that the command "Navigate to Tokyo" means the destination "Tokyo" and that the user's emotion is, for example, nervous.

[1168] Step 5:

[1169] The terminal determines an appropriate action based on the analysis result. In this case, since the "navigate" function is requested, it determines to process it as a navigation action.

[1170] Step 6:

[1171] The device sends an HTTP request to the server according to the determined action, including necessary parameters such as destination information and user emotion information.

[1172] Step 7:

[1173] The server receives the HTTP request from the device and processes it accordingly. Specifically, it generates route information to the destination and formats and prepares the message.

[1174] Step 8:

[1175] The server sends the generated route information and an appropriate message back to the device as a response, which includes navigation information and an emotion-based message.

[1176] Step 9:

[1177] The terminal receives the response from the server and analyzes its contents to provide appropriate feedback to the user. In this case, a text message is analyzed.

[1178] Step 10:

[1179] The device's speech synthesizer converts the text message into speech, generating a voice message that reflects the results of the emotion engine. For example, the message might say, "You seem nervous, but please stay calm and drive. Navigation to Tokyo will begin."

[1180] Step 11:

[1181] Users can receive voice feedback from the device and obtain real-time navigation information without taking their hands off the steering wheel.

[1182] Through these steps, the system provides safe and efficient operation for the rider while riding the motorcycle, and in particular, by using an emotion engine, it enables personalized feedback according to the user's emotional state.

[1183] Example 2

[1184] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1185] When users operate hands-free while driving a motorcycle, they are faced with the challenge of being unable to operate the system safely and efficiently. Furthermore, conventional hands-free systems lack personalized feedback that takes into account the user's emotions, limiting the user experience. This increases stress and anxiety for users while driving, and creates the problem of insufficient safety and convenience.

[1186] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1187] In this invention, the server includes means for receiving a voice input, means for converting the received voice input into text, means for extracting emotional information from the converted text and voice information, means for analyzing the converted text and the extracted emotional information to determine an appropriate action, means for sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice based on the emotional information. This makes it possible to provide appropriate and personalized feedback and actions that take into account the emotional state of the user.

[1188] "Voice input" refers to commands or instructions spoken by a user received as digital data through a microphone or other audio collection device.

[1189] "Text" refers to a string of characters, such as alphabets or kanji, that has been converted from voice input, and is information in a format that can be analyzed on a device such as a computer.

[1190] "Emotional information" is information extracted by analyzing the user's emotional state (e.g., tension, calmness, anxiety, etc.) from voice data.

[1191] "Analysis" is a process for deriving appropriate actions based on the received text data and emotional information.

[1192] "Actions" refer to the operations or reactions the system performs based on the user's voice commands, such as navigation, music playback, and answering calls.

[1193] A "request" is an inquiry or request sent from a terminal to a server, and is executed in the form of an HTTP request or the like.

[1194] A "response" is a reply or answer sent from a server to a terminal, and is data that includes information or instructions corresponding to a request.

[1195] A "speech recognition service" is any cloud-based or on-premise technology or software that accepts voice data as input and converts it into text data.

[1196] A "speech synthesizer" is a device that converts text information into speech and provides speech output to the user.

[1197] This invention is a system that allows a user to perform various operations hands-free while riding a motorcycle, and provides personalized feedback and actions by combining it with an emotion engine that recognizes the user's emotions. This system is composed of a terminal for receiving voice input, a voice recognition means for converting the voice input into text, a means for extracting emotion information from the text and voice data, a means for analyzing the text and emotion information and determining an appropriate action, a communication means for sending a request to a server based on the determined action, and a means for feeding back the response from the server to the user as voice.

[1198] First, when a user issues a voice command such as "navigate to a destination," the device captures this voice and records it as a digital audio file via the microphone. This is real-time voice capture.

[1199] Next, the terminal transmits the recorded voice data to a voice recognition service (e.g., voice recognition software) to convert the voice data into text data, and also uses an emotion engine (e.g., emotion analysis software) to analyze the user's emotions from the voice data and extract emotion information.

[1200] The converted text data and emotional information are then analyzed by an analysis unit within the device. For example, in the command "Navigate to destination," the device determines that "destination" is the destination and simultaneously recognizes the user's emotional state, such as whether they are nervous or calm.

[1201] Based on the analysis results, the system determines the appropriate action. In this case, it determines that the "navigation" function is requested. The expression and content of the feedback are adjusted taking into account the information from the emotion engine. If the user is nervous, the device will provide navigation guidance in a calm voice tone, and if the user is calm, it will provide guidance in a normal tone.

[1202] Furthermore, based on the determined action, the device sends an HTTP request to the server. This request includes necessary parameters such as destination information and emotional state. When the server receives this request, it performs processing according to the content and generates a corresponding response. For example, if it is a navigation request, it calculates and generates route information to the destination.

[1203] The response from the server is received by the device, which then uses a speech synthesizer (e.g., speech synthesis software) to convert the text message into speech and provide it to the user. This feedback reflects the results of the emotion engine, and may be a message such as, "You seem nervous, but please stay calm and drive. Navigation to your destination will begin."

[1204] As a concrete example, if a user says "Navigate to Tokyo" while driving, the device captures the voice and converts it into text data using a speech recognition service, while the emotion engine simultaneously recognizes the user's nervousness. Based on the analysis results, the device sends this as a navigation request to the server, which then provides route information to Tokyo. Using a speech synthesizer, the device then provides the user with an adjusted voice message such as "You seem nervous, but please remain calm while driving. Navigation to Tokyo will begin," providing a safer and more comfortable riding experience.

[1205] Examples of prompts for generative AI models include:

[1206] > "Design a system that allows a user to operate a motorcycle hands-free while driving. The system combines voice input with an emotion engine to provide appropriate feedback based on the user's emotional state. Specifically, if the user is nervous, the system will provide navigation information in a calm tone, and if the user is calm, the system will provide information in a normal tone."

[1207] The system's unique feature is that it significantly improves the user's safety and convenience while riding a motorcycle, and by utilizing an emotion engine, personalized interactions based on the user's emotions are possible.

[1208] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1209] Program processing flow

[1210] Step 1: Capture voice commands

[1211] The user issues the voice command "Navigate to Tokyo."

[1212] Input: Voice command

[1213] How it works: The device captures the user's voice commands through the microphone and records them as digital audio files in real time.

[1214] Output: Digital audio file

[1215] Step 2: Speech to text

[1216] The device transmits the recorded digital audio file to a speech recognition service.

[1217] Input: Digital audio file

[1218] How it works: A speech recognition service converts speech data into text data. For example, speech recognition software is used to generate the text "Navigate to Tokyo."

[1219] Output: Text data

[1220] Step 3: Extracting emotional information

[1221] The device sends the digital audio file to the emotion engine.

[1222] Input: Digital audio file

[1223] How it works: The emotion engine analyzes the user's emotions from the voice data and extracts emotional information. For example, it uses emotion analysis software to determine that the user is "nervous."

[1224] Output: Emotional information

[1225] Step 4: Analyzing text data and sentiment information

[1226] The analysis unit in the device analyzes the text data and emotional information.

[1227] Input: Text data, emotion information

[1228] Action: The analysis unit analyzes the text "Navigate to Tokyo" and the emotional information "I'm nervous" and determines the appropriate action. Specifically, it determines that "Tokyo" is the destination and that the "Navigate" function is required.

[1229] Output: Analysis results (actions, destination information, emotion information)

[1230] Step 5: Sending a request to the server

[1231] The terminal sends an HTTP request to the server based on the analysis results.

[1232] Input: Analysis results

[1233] How it works: The HTTP request contains destination and emotion information. By sending this request, the server is asked to provide navigation information.

[1234] Output: HTTP request

[1235] Step 6: Server-side processing

[1236] The server receives the HTTP request and performs the corresponding processing.

[1237] Input: HTTP request

[1238] Operation: The server analyzes the request (destination information, emotion information) and generates navigation information. For example, it uses route calculation software to calculate the route to Tokyo.

[1239] Output: Navigation information

[1240] Step 7: Receiving responses and coordinating feedback

[1241] The terminal receives the response from the server.

[1242] Input: Navigation information

[1243] How it works: The device uses a speech synthesizer to convert navigation information into a voice message to convey to the user. The device adjusts the tone of the voice message based on emotional information. For example, it generates a message like, "You seem nervous, but please stay calm and drive. We'll start navigating to Tokyo."

[1244] Output: Modified voice message

[1245] Step 8: Provide feedback to users

[1246] The terminal finally provides feedback to the user as a voice message.

[1247] Input: Modified voice message

[1248] Operation: The terminal outputs a voice message to the user through a voice synthesizer.

[1249] Output: The user receives the adjusted navigation instructions.

[1250] Specific examples

[1251] When a user says "Navigate to Tokyo," the speech is captured by the device and converted into text by the speech recognition service. At the same time, the speech data is analyzed by the emotion engine, which recognizes that the user is nervous. Based on the results, the analysis unit sends an HTTP request to the server as a navigation action, requesting directions to Tokyo. The server receives the request and provides route information to Tokyo, and the device provides the user with a voice message tailored based on the emotion information.

[1252] (Application example 2)

[1253] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1254] Existing technologies include systems that use voice commands to perform various operations, but these systems are unable to recognize the user's emotional state and provide appropriate feedback based on that. This results in a uniform user experience, which does not improve safety or convenience, especially in emergencies or stressful situations. Furthermore, there is a lack of a hands-free means of operating robots to improve work efficiency and safety in factories.

[1255] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input, means for converting the received voice input into text, means for analyzing the converted text and emotional information to determine an appropriate action, means for combining an emotion engine that recognizes the user's emotion and sending a request to the server based on the determined action and emotional information, and means for receiving a response from the server and providing feedback to the user as voice. This makes it possible to provide personalized feedback and actions according to the user's emotional state, greatly improving work efficiency and safety in factories.

[1256] "Means for receiving voice input" refers to a function for capturing voice commands issued by a user through a device such as a microphone.

[1257] A "means for converting voice input to text" is a function that allows for analyzing captured voice data and converting it into text format.

[1258] "Means for analyzing the converted text and determining the appropriate action" refers to a function that understands and analyzes the voice command converted into text and selects the appropriate action based on the results.

[1259] The "emotion engine" is a function for recognizing and analyzing the user's emotional state from voice data.

[1260] The "means for sending a request to the server" is a function for sending a request to the server for an appropriate action based on the analyzed text data and emotion information.

[1261] The "means for receiving a response from the server and providing feedback to the user as audio" is a function for providing feedback to the user by outputting the data received from the server as audio.

[1262] "Navigation, music playback, call answering, weather information check, emergency communication, robot operation" are the various types of actions that the system can perform in response to a user's voice commands.

[1263] A "prompt sentence using a generative AI model" is an input sentence used by a generative AI technology to obtain a desired output.

[1264] This invention relates to a system that enables a worker working in a factory to operate a robot hands-free using smart glasses. The system includes means for receiving voice input, means for converting the voice input into text, means for analyzing the converted text and emotion information to determine an appropriate action, an emotion engine for recognizing the emotion of the user, means for sending a request to a server based on the determined action and emotion information, and means for receiving a response from the server and providing feedback to the user as voice.

[1265] As a concrete example of the system, a worker puts on smart glasses and issues a command by voice, such as "Transport this part." The smart glasses' microphone captures the voice command and generates a digital audio file in real time. The smart glasses then send the recorded voice data to a voice recognition service (e.g., Google Speech-to-Text) to convert the voice data into text data. At the same time, an emotion engine (e.g., EmotionRecognition API) analyzes the user's emotional state and extracts emotional information.

[1266] The converted text data and emotion information are further analyzed by the analysis means in the smart glasses. For example, for a command such as "transport this part," the emotion engine determines that there is a transport command and recognizes that the user is nervous. Based on this result, an appropriate action is determined. If the user is nervous, commands are sent to the robot to operate as smoothly and safely as possible.

[1267] Based on the determined action, the smart glasses send an HTTP request to the server. The request includes parameters related to the transport command and the emotional state. The server receives the request and generates specific operational commands for the robot, such as calculating and returning a smooth transport route to the destination.

[1268] The response from the server is received by the smart glasses. The content is read as a text message and audio feedback is provided to the user using a speech synthesizer (e.g., Google Text-to-Speech). Feedback that takes into account the user's state of tension is provided, such as "You seem nervous, but please calm down. We'll start transporting the parts."

[1269] Hardware and software used

[1270] Smart glasses (e.g., generic name)

[1271] microphone

[1272] Connected robot arm or mobile robot

[1273] Speech recognition services (e.g., Google Speech-to-Text)

[1274] Emotion recognition engine (e.g., EmotionRecognition API)

[1275] Speech synthesizers (e.g., Google Text-to-Speech)

[1276] Prompt Sentence Examples

[1277] Prompt: Recognizes user voice commands and their emotions

[1278] Input: Audio data text and its audio file

[1279] Output: Command converted to text and recognized emotion information (e.g., carrying, stressed)

[1280] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1281] Step 1:

[1282] A user issues a voice command through the smart glasses. The input includes the user's speech. A microphone on the smart glasses captures the voice command and records it as a digital audio file. The output is a digital audio file.

[1283] Step 2:

[1284] The device sends the captured voice data to a speech recognition service, which converts the voice data into text data. The input includes a digital audio file. The speech recognition service used (e.g., Google Speech-to-Text) analyzes the voice signal and generates corresponding text. The output is the converted text data.

[1285] Step 3:

[1286] The device simultaneously sends the voice data to an emotion recognition engine to analyze the user's emotional state. The input includes a digital audio file. The emotion recognition engine (e.g., EmotionRecognition API) analyzes the tone and rate of the voice to extract emotional information. The output is the user's emotional information.

[1287] Step 4:

[1288] The device analyzes the converted text data and emotional information to determine the appropriate action. The input includes text data and emotional information. The text analysis engine understands the command content and combines it with the emotional information to determine the appropriate action. The output is the determined action.

[1289] Step 5:

[1290] The device sends an HTTP request to the server based on the determined action and emotion information. The input includes the action and emotion information. The HTTP protocol is used to send the request with appropriate parameters to the server. The output is a request that is processed on the server side.

[1291] Step 6:

[1292] Based on the request received by the server, it generates specific operation commands for the robot. The input includes action and emotion information. The logic engine on the server calculates the optimal route and operation parameters based on this data and sends the command to the robot. The output is the operation command sent to the robot.

[1293] Step 7:

[1294] The robot starts working based on the operation command from the server. The operation command from the server is included as input. The robot executes the specified task according to the command. The output is the execution status of the task.

[1295] Step 8:

[1296] The server receives feedback from the robot and sends the feedback to the smart glasses. The input includes feedback information from the robot. The server generates this information as a voice message and uses a voice synthesizer (e.g., Google Text-to-Speech) to generate the voice feedback. The output is the voice message.

[1297] Step 9:

[1298] The smart glasses provide the generated audio feedback to the user. The input includes an audio message. The smart glasses provide audio feedback to the user such as, "You seem nervous, please stay calm. We'll start transporting the parts." The output is the audio feedback provided to the user.

[1299] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1300] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1301] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1302] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1303] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1304] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1305] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1306] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1307] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1308] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1309] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1310] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1311] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1312] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1313] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1314] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1315] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1316] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1317] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1318] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1319] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1320] The following is further disclosed regarding the above embodiment.

[1321] (Claim 1)

[1322] means for receiving audio input;

[1323] means for converting received voice input into text;

[1324] a means for analyzing the converted text to determine an appropriate action;

[1325] means for sending a request to a server based on the determined action;

[1326] a means for receiving a response from the server and providing feedback to the user as audio;

[1327] A system including:

[1328] (Claim 2)

[1329] means for analyzing the generated text and determining at least one action from among navigation, music playback, answering a call, checking weather information, and emergency communication;

[1330] a communication means for receiving a response from the server;

[1331] a means for providing audio feedback to the user;

[1332] The system of claim 1 further comprising:

[1333] (Claim 3)

[1334] a means for utilizing a speech recognition service to convert voice input into text;

[1335] a means for sending an HTTP request to a server based on the analysis results;

[1336] means for receiving a response from the server and using a speech synthesizer to convert the text message into a speech output;

[1337] The system of claim 1 further comprising:

[1338] "Example 1"

[1339] (Claim 1)

[1340] means for receiving audio input;

[1341] means for converting received voice input into text;

[1342] a means for analyzing the converted text to determine an appropriate action;

[1343] means for transmitting a request to the information processing device based on the determined action;

[1344] means for receiving a response from the information processing device and providing feedback to the user in the form of voice;

[1345] a means for using a speech recognition service;

[1346] A means of analyzing the text using a natural language processing library;

[1347] A means of sending an HTTP request;

[1348] means for converting the speech into speech using a speech synthesizer;

[1349] A system including:

[1350] (Claim 2)

[1351] means for analyzing the generated text and determining at least one action from among providing location information, controlling media content, responding to communication, obtaining weather information, and safety communication;

[1352] a communication means for receiving a response from the information processing device;

[1353] a means for providing audio feedback to the user;

[1354] The system of claim 1 further comprising:

[1355] (Claim 3)

[1356] a means for utilizing a speech recognition service to convert voice input into text;

[1357] means for transmitting an HTTP request to an information processing device based on the analysis result;

[1358] means for receiving a response from the information processing device and using a speech synthesizer to convert the text message into a speech output;

[1359] The system of claim 1 further comprising:

[1360] "Application Example 1"

[1361] (Claim 1)

[1362] means for receiving audio input;

[1363] means for converting received voice input into text;

[1364] a means for analyzing the converted text to determine an appropriate action;

[1365] means for sending a request to a server based on the determined action;

[1366] a means for receiving a response from the server and providing feedback to the user as audio;

[1367] In particular, a means for providing navigation, notification confirmation, voice response, situation reporting, and emergency contact functions to enable delivery service riders to obtain information safely and efficiently while driving;

[1368] A system including:

[1369] (Claim 2)

[1370] a means for analyzing the generated text and determining at least one action from among navigation, music playback, answering a call, checking weather information, emergency communication, checking delivery service notifications, voice response, status reporting, and emergency contact;

[1371] a communication means for receiving a response from the server;

[1372] a means for providing audio feedback to the user;

[1373] The system of claim 1 further comprising:

[1374] (Claim 3)

[1375] a means for utilizing a speech recognition service to convert voice input into text;

[1376] a means for sending an HTTP request to a server based on the analysis results;

[1377] means for receiving a response from the server and using a speech synthesizer to convert the text message into a speech output;

[1378] an audio guide means for providing information specific to delivery work while driving;

[1379] The system of claim 1 further comprising:

[1380] "Example 2: Combining Emotion Engines"

[1381] (Claim 1)

[1382] means for receiving audio input;

[1383] means for converting received voice input into text;

[1384] means for extracting emotion information from the converted text and speech information;

[1385] means for analyzing the converted text and the extracted sentiment information to determine an appropriate action;

[1386] means for sending a request to a server based on the determined action and emotion information;

[1387] a means for receiving a response from the server and providing feedback to the user as voice based on the emotion information;

[1388] A system including:

[1389] (Claim 2)

[1390] a means for analyzing the generated text and emotion information and determining at least one action from among navigation, music playback, answering a call, checking weather information, and emergency communication;

[1391] a communication means for receiving a response from the server;

[1392] a means for providing feedback to the user as a voice based on the emotion information;

[1393] The system of claim 1 further comprising:

[1394] (Claim 3)

[1395] a means for utilizing a speech recognition service to convert voice input into text;

[1396] means for sending an HTTP request to a server based on the analysis result and the emotion information;

[1397] means for receiving a response from the server and using a speech synthesizer to convert the text message into a speech output based on the emotion information;

[1398] The system of claim 1 further comprising:

[1399] "Application example 2 when combining emotion engines"

[1400] (Claim 1)

[1401] means for receiving audio input;

[1402] means for converting received voice input into text;

[1403] a means for analyzing the converted text to determine an appropriate action;

[1404] means for combining an emotion engine that recognizes the user's emotion and sending a request to a server based on the determined action and emotion information;

[1405] a means for receiving a response from the server and providing feedback to the user as audio;

[1406] A system including:

[1407] (Claim 2)

[1408] a means for analyzing the generated text and emotion information and determining at least one action from among navigation, music playback, call answering, weather information check, emergency communication, and robot operation;

[1409] a communication means for receiving a response from the server;

[1410] a means for providing audio feedback to the user;

[1411] The system of claim 1 further comprising:

[1412] (Claim 3)

[1413] a means for utilizing a speech recognition service to convert voice input into text;

[1414] means for sending an HTTP request to a server based on the analysis result and the emotion information;

[1415] means for receiving a response from the server and using a speech synthesizer to convert the text message into speech output based on a prompt utilizing the generative AI model;

[1416] The system of claim 1 further comprising: [Explanation of symbols]

[1417] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving audio input; means for converting received voice input into text; a means for analyzing the converted text to determine an appropriate action; means for sending a request to a server based on the determined action; a means for receiving a response from the server and providing feedback to the user as audio; A system including:

2. means for analyzing the generated text and determining at least one action from among navigation, music playback, answering a call, checking weather information, and emergency communication; a communication means for receiving a response from the server; a means for providing audio feedback to the user; The system of claim 1 further comprising:

3. a means for utilizing a speech recognition service to convert voice input into text; a means for sending an HTTP request to a server based on the analysis results; means for receiving a response from the server and using a speech synthesizer to convert the text message into a speech output; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A