System

The car navigation system addresses unclear route guidance by allowing voice input and feedback, ensuring safe and convenient driving with improved navigation accuracy through user data analysis.

JP2026033968APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137089
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Conventional car navigation systems fail to provide clear route guidance, leading to confusion at intersections and forks, which diverts user attention and compromises driving safety, and lack the ability to learn from user confusion patterns to improve navigation accuracy.

Method used

A car navigation system that allows voice input for setting destinations and asking questions, converting voice to text, analyzing with a server, generating guidance, and providing voice responses, while collecting user data to update the navigation algorithm for improved clarity.

Benefits of technology

Enables safe and convenient driving by providing clear voice guidance and continuously improving navigation accuracy through user feedback, enhancing user experience and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026033968000001_ABST
    Figure 2026033968000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting voice for a user to set a destination and receive route guidance; means for converting the voice input into text data; means for analyzing the text data and generating route guidance information; means for transmitting the route guidance information as text data to a terminal; means for converting the text data into voice and providing the voice to the user; means for collecting and analyzing question data of the user; and means for updating a route guidance algorithm based on an analysis result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional car navigation systems make it difficult for users to obtain appropriate guidance when route directions are unclear, resulting in problems with safe driving and convenience. For example, users often get lost at complex intersections and forks, which requires them to divert their attention while driving, compromising driving safety. Furthermore, even though there are patterns in which many users get confused in the same places, this information has not been used to improve the system. [Means for solving the problem]

[0005] The present invention provides a car navigation system that allows users to set a destination and ask questions about directions via voice input. Specifically, the system includes a means for converting the voice input into text data and sending the text data to a server for analysis, and a means for the server to analyze the text data and generate appropriate route guidance information. The system also includes a means for sending the generated route guidance information to a terminal, which then converts the information into voice and provides it to the user. This system allows users to receive route guidance without taking their hands off the wheel while driving, improving driving safety. Furthermore, by analyzing question data collected from all users, the system can identify areas where many users are confused and update the route guidance algorithm based on that information, thereby providing even easier-to-understand route guidance.

[0006] "User" means a driver or passenger who uses the system to receive navigation services.

[0007] "Voice input" refers to voice data uttered by a user through a voice input device such as a microphone.

[0008] A "destination" is the final destination that a user aims to reach.

[0009] "Directions" refers to the act of providing route information and instructions necessary for a user to reach a destination.

[0010] A "terminal" is a device that functions as a car navigation system and provides a voice interface with the user.

[0011] The "server" is a remote computer system that analyzes the received text data, generates route guidance information, and transmits it to the terminal.

[0012] "Text data" refers to data obtained by converting voice input into text information.

[0013] "Natural language processing" is a technology that analyzes voice or text data as human language and understands its meaning and intent.

[0014] A "map information database" is a database that stores geographical information and is referenced for route guidance.

[0015] "Speech synthesis" is a technology that converts text data into speech that humans can understand.

[0016] "Question data" is text data of a question posed by a user to the system.

[0017] A "route guidance algorithm" is a computer program method that generates optimal routes and instructions based on map information.

[0018] "Update" refers to the system making improvements or changes based on new information or analytical results. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0041] User Actions

[0042] The user sets the destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0043] Terminal handling

[0044] The device converts the user's voice input into text data using a speech recognition engine. For example, the speech "Should I turn right at the next fork?" is converted into text data such as "Should I turn right at the next fork?" The converted text data is designed to be sent to a server. The device then receives the text data from the server, converts it into speech using a speech synthesis engine, and provides it to the user.

[0045] Server Processing

[0046] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm whether or not to turn right at the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[0047] Feedback and Learning

[0048] The device converts the text data sent from the server into speech and provides the user with a voice response saying, "That's right." In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users have problems. Based on this information, the server updates its route guidance algorithm and provides even easier-to-understand route guidance.

[0049] Specific examples

[0050] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server references a map information database, confirms that the next fork will be a right turn, and generates the appropriate answer, "That's right." The generated answer is then sent back to the device as text data, converted into audio on the device, and provided to the user. The user can confirm that the route guidance is accurate by hearing the audio answer, "That's right."

[0051] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[0052] The processing flow will be explained below.

[0053] Specific processing steps of the program

[0054] 1. User Actions

[0055] Step 1:

[0056] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[0057] 2. Terminal Processing

[0058] Step 2:

[0059] The device receives the user's voice input and captures the audio signal with the microphone.

[0060] Step 3:

[0061] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[0062] Step 4:

[0063] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[0064] Step 5:

[0065] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[0066] Step 6:

[0067] The device again receives the user's voice input, capturing the audio signal with the microphone.

[0068] Step 7:

[0069] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[0070] Step 8:

[0071] The device sends text data to the server as an HTTP request.

[0072] 3. Server Processing

[0073] Step 9:

[0074] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[0075] Step 10:

[0076] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[0077] Step 11:

[0078] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[0079] Step 12:

[0080] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[0081] Step 13:

[0082] The server generates text data and sends it to the terminal as an HTTP response.

[0083] 4. Device reprocessing and user feedback

[0084] Step 14:

[0085] The device receives the text data from the server. The device receives the text data "That's right."

[0086] Step 15:

[0087] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[0088] Step 16:

[0089] The device generates a voice message and plays it over the speaker, telling the user "That's right."

[0090] 5. Questionnaire data collection and analysis

[0091] Step 17:

[0092] The server stores the user's question data. The question "Should I turn right at the next fork?" and the answer "Yes, that's right" are stored in the database.

[0093] Step 18:

[0094] The server periodically analyzes the stored question data to identify areas where many users have experienced problems.

[0095] Step 19:

[0096] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[0097] Through this processing step, users can ask questions by voice and receive appropriate directions, while the server analyzes the question data to improve the accuracy of the system.

[0098] Example 1

[0099] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0100] Conventional car navigation systems often provide inappropriate directions when receiving voice input and are unable to accurately understand the user's intent. Furthermore, users are unable to ask questions by voice if they are unable to understand the directions, which is inconvenient. Furthermore, the lack of a feedback function that allows many users to identify areas where they have problems and improve the navigation algorithm makes it difficult to improve the system's accuracy and user experience.

[0101] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0102] In this invention, the server includes means for converting text data received by the terminal from the server into speech and providing it to the user through an output device in the vehicle, means for analyzing the user's intentions using a natural language processing engine, and means for referencing a map information database to obtain information on the user's current location and the next intersection. This makes it possible to provide appropriate route guidance and answers to destination settings and questions entered through voice input, and to improve the accuracy of the entire system and the user experience through the feedback function.

[0103] 1. "User" means a driver or passenger who uses a car navigation system in a vehicle.

[0104] 2. "Voice input" refers to the act of the user speaking their destination or questions into the car navigation system through a microphone.

[0105] 3. "Text data" means data in the form of a string of characters converted from voice input by a voice recognition engine.

[0106] 4. "Terminal" means an electronic device, including input and output devices, installed in a vehicle in which a car navigation system is installed.

[0107] 5. "Server" means a remote computer system for receiving and analyzing text data.

[0108] 6. "Speech recognition engine" means a software or hardware mechanism that converts voice input into text data.

[0109] 7. A "natural language processing engine" is an artificial intelligence technology that analyzes text data and understands user intent.

[0110] 8. "Map information database" means a database containing road information and geographic information that the server references to generate route guidance information.

[0111] 9. "Route Guidance Information" means information regarding the direction and route generated by the server based on the user's destination and current location.

[0112] 10. "Speech synthesis engine" means a software or hardware mechanism that converts text data into speech.

[0113] 11. "Feedback" is the process of collecting and analyzing user query data to improve the system's navigation algorithms.

[0114] 12. A "guidance algorithm" is a set of procedures or calculation methods for providing users with appropriate directions and routes.

[0115] 13. "Output device" means a device, such as a speaker, that provides voice prompts or other information to the user in the vehicle.

[0116] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0117] First, the user sets their destination by voice into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0118] The device converts the user's voice input into text data using a speech recognition engine (e.g., an ambiguous speech recognition engine or a third-party speech recognition engine). After the voice is converted into text data, the text data is sent to a server via the Internet. The server then analyzes the received text data using a natural language processing engine (e.g., a generative AI model or a third-party natural language processing engine) to understand the intent of the user's question.

[0119] If the analysis of the text data determines that the user wants to confirm whether they should turn right at the next intersection, the server references a map information database (for example, a general map information database API or a map information database API from another company). The server obtains route guidance information based on the user's current location and the information about the next intersection, and generates text data for the appropriate answer, such as "That's right." This text data is then sent back to the device.

[0120] The terminal converts the text data sent from the server into speech using a speech synthesis engine (for example, a rough speech synthesis engine or a speech synthesis engine made by another company) and provides the speech to the user through an output device (such as a speaker) in the car. In this way, the user can hear the voice response "That's right."

[0121] In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users experience problems. This allows the route guidance algorithm to be updated and more user-friendly route guidance to be provided. Specifically, the collected data is analyzed using machine learning algorithms to identify areas for improvement in the system.

[0122] Specific examples

[0123] For example, if a user asks "Should I turn right at the next fork?" while driving towards their destination, the question is converted into text data "Should I turn right at the next fork?" using a voice recognition engine on the device. The text data is sent to a server via the Internet, and the server analyzes the question using a natural language processing engine to understand what the user wants to confirm. The server references a map information database, confirms that they should turn right at the next fork, and generates the appropriate answer "That's right." The generated answer is again sent to the device as text data, where it is converted into voice using a voice synthesis engine. The user can confirm that the route guidance is accurate by hearing the voice answer "That's right."

[0124] Examples of prompt statements

[0125] "Utterance: 'Should I turn right at the next fork?'"

[0126] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[0127] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0128] Step 1: User Speech Input

[0129] The user sets the destination by speaking into the car navigation system terminal. The input is the user's voice, which can be a specific destination such as "My next destination is Tokyo Tower" or a question such as "Should I turn right at the next fork?" The voice is received through the terminal's microphone input device.

[0130] Step 2: Voice Recognition

[0131] The device converts the user's voice input into text data using a voice recognition engine (e.g., voice recognition software). The input is the user's voice signal, and the output is text data in the form of a string: "Should I turn right at the next fork?" The voice data is converted into a digital signal, and phonemic and linguistic analysis is performed before being output as text.

[0132] Step 3: Send text data to the server

[0133] The device sends the text data converted by the speech recognition engine to the server via the Internet. The input is the text data "Should I turn right at the next fork?" and the output is data in the form of network packets. The device's network module is responsible for this communication and sends the data to the server via the HTTPS protocol.

[0134] Step 4: Analyzing the text data

[0135] The server analyzes the received text data using a natural language processing engine (e.g., a generative AI model). The input is the text data "Should I turn right at the next fork?" and the output is structured data that represents the user's intent. The server performs tokenization and grammatical analysis on the text data to understand what the user wants to confirm.

[0136] Step 5: Query the map database

[0137] The server references a map information database (e.g., a map information API) based on the analysis results. The input is the user's current location and information about the next intersection, and the output is route guidance information including the action to be taken at the next intersection (e.g., turn right). The server sends a query to the map information database to obtain the required information.

[0138] Step 6: Generate directions

[0139] The server generates appropriate route guidance information based on information obtained from the map information database. The input is route guidance information obtained from the map database, and the output is the text-format route guidance information "That's right" to be provided to the user. The server composes this as text data in character string format.

[0140] Step 7: Sending text data from the server to the device

[0141] The generated text data is resent from the server to the terminal. The input is the generated text data "That's right," and the output is data in the form of network packets. The server's network module is responsible for this communication.

[0142] Step 8: Text-to-Speech

[0143] The device converts the text data received from the server into speech using a speech synthesis engine (e.g., text-to-speech software). The input is the text data "That's right," and the output is an audio file (e.g., audio data in WAV or MP3 format). The text data is input into the speech synthesis engine, and output as an audio file.

[0144] Step 9: User voice guidance

[0145] The terminal provides the generated voice data to the user through the car's speaker. The input is the voice file "That's right," and the output is the voice played through the speaker. The user listens to this and confirms the next action.

[0146] Step 10: Feedback and learning

[0147] The server collects question data from all users and periodically analyzes it. The input is the collected question data (e.g., text data history), and the output is feedback data containing improvements to the route guidance algorithm. A machine learning algorithm is used to analyze the data and identify improvements to the system. This allows the server to update the route guidance algorithm and improve the accuracy of the entire system.

[0148] (Application example 1)

[0149] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0150] In food delivery operations, it is difficult for drivers to set destinations and check routes safely and efficiently while driving. Manual resetting and checking while driving is dangerous, so there is a need for voice interaction.

[0151] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0152] In this invention, the server includes means for a user to input voice commands to set a destination and receive route guidance, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and generating route guidance information, means for transmitting the route guidance information as text data to a terminal, means for converting the text data into voice and providing it to the user, means for collecting and analyzing user question data, means for updating a route guidance algorithm based on the analysis results, means for a delivery driver to set a destination and check a route by voice while driving, means for providing the route confirmation results by voice in real time, and means for generating next destination and route information using document generation technology. This enables drivers to set a destination and check a route by voice alone while driving, improving the safety and efficiency of their work.

[0153] "Voice input" is a method by which a user communicates verbal instructions or questions to a system.

[0154] "Text data" is data that has been converted from voice input into text information and is used for analysis and processing.

[0155] A "server" is a computing device that receives transmitted text data, analyzes it, and generates a response.

[0156] "Route guidance information" is information that indicates the route and direction for the user to reach the destination.

[0157] "Device" means a device operated by a user that receives audio input and provides audio output.

[0158] A "voice recognition engine" is software that analyzes voice input and converts it into corresponding text data.

[0159] A "speech synthesis engine" is software that analyzes text data and generates corresponding speech.

[0160] "Question data" refers to information about confirmations or questions that users make to the system.

[0161] "Parsing" is the process of understanding received text data and deciphering its intent.

[0162] A "route guidance algorithm" is a calculation procedure for providing optimal route information to users.

[0163] A "delivery driver" is a driver whose job is to deliver goods to customers.

[0164] "Real-time" refers to near-instant processing or response.

[0165] "Document generation technology" is technology that generates appropriate text responses to user questions.

[0166] "Driving" refers to the state in which the delivery driver is operating the vehicle.

[0167] This invention is a system that allows food delivery drivers to easily set destinations and check routes by voice while driving. To implement the invention, the following specific hardware and software configurations are required.

[0168] Hardware

[0169] Smartphone (device)

[0170] microphone

[0171] speaker

[0172] software

[0173] Speech recognition engine: SpeechRecognition (Python library)

[0174] Speech synthesis engine: gTTS (Google (registered trademark) Text-to-Speech)

[0175] API request tool: requests (Python library)

[0176] Audio output playback utility: mpg321

[0177] Program processing

[0178] 1. User Operation

[0179] The user (delivery driver) speaks instructions into their smartphone, such as "Where's the next delivery destination?" or "Where's the next right turn?"

[0180] 2. Voice Input and Speech Recognition

[0181] The device receives the user's voice input through a microphone. The speech recognition engine (SpeechRecognition) converts this voice input into text data. For example, the voice "Where is the next delivery?" is converted into text data "Where is the next delivery?"

[0182] 3. Sending to the server and analyzing

[0183] The converted text data is sent from the device to the server. The server receives this text data and analyzes it using a natural language processing engine. For example, from the question "Where is the next delivery?", it understands that the user wants to check information about the next delivery address.

[0184] 4. Generating Route Guidance Information

[0185] Based on the analysis results, the server references a map information database and generates information about the next delivery destination and route. The generated route guidance information is sent to the terminal as text data such as "The next delivery destination is ____."

[0186] 5. Speech synthesis and delivery

[0187] The device converts the received text data into speech using a speech synthesis engine (gTTS), and provides the user with a voice message such as, "The next delivery destination is ____." This allows the user to check the next delivery destination without looking at the screen while driving.

[0188] Specific examples

[0189] Specifically, when a delivery driver asks, "Where's the next delivery?", the question is converted by the device into text data, "Where's the next delivery?" The text data is sent to the server, which analyzes the question and confirms the next delivery address. The server generates a response such as "The next delivery address is: XXX" and sends it back to the device as text data. The response is then converted into speech on the device and provided to the user.

[0190] Example prompts for generative AI models

[0191] "When a user asks, 'Where is my next delivery?', please search for the next delivery destination in a map database and generate a sentence that provides appropriate route guidance."

[0192] This will create a system that allows food delivery drivers to safely and efficiently check their destination and set their route while driving.

[0193] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0194] Step 1:

[0195] Using the voice input function of the smartphone, the user asks aloud, "Where is the next delivery destination?" The input is voice data, and the user's question is captured by the smartphone's microphone.

[0196] Step 2:

[0197] The device uses a speech recognition engine (SpeechRecognition) to convert voice data into text data. It analyzes the voice data and generates corresponding text data. For example, a speech saying "Where is the next delivery?" is converted into text data "Where is the next delivery?". This text data is used in the next processing step.

[0198] Step 3:

[0199] The converted text data is sent from the device to the server. The device uses an API request tool (requests) to send this text data to the server. The input is text data, and the output is an HTTP request to the server.

[0200] Step 4:

[0201] The server analyzes the text data received from the device. It uses a natural language processing engine to analyze the received text data and understand the intent of the user's question. Based on the analysis results, it generates a query to search for information on the next delivery destination. This query is used to reference a map information database.

[0202] Step 5:

[0203] The server uses the generated query to refer to a map information database and obtain information about the next delivery destination. The information obtained from the database includes the delivery destination address and route information. The obtained information is saved in text format and passed to the next processing step. For example, a text such as "The next delivery destination is at address: XXX" is generated.

[0204] Step 6:

[0205] The server uses document generation technology to generate an appropriate response based on the acquired delivery destination information. It uses a generative AI model (text generation engine) to create an appropriate response to the question. The response is specific text data such as "The next delivery destination is address: XXX." This text data is then sent back to the terminal.

[0206] Step 7:

[0207] The device converts the received text data into speech using a speech synthesis engine (gTTS). The text data is analyzed and the corresponding voice data is generated. The generated voice data is played back through the smartphone's speaker. For example, the user can hear a voice saying, "The next delivery address is: XXX."

[0208] Step 8:

[0209] The user continues driving based on the next delivery destination information provided by voice. The user can confirm the voice response and follow the appropriate route to the next delivery destination. In this way, a system is realized that allows users to confirm destinations and set routes safely and efficiently while driving.

[0210] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0211] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and also recognizes and provides feedback on the user's emotional state. This system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0212] User Actions

[0213] The user sets their destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0214] Terminal handling

[0215] The device converts the user's voice input into text data using a speech recognition engine. For example, the device converts the speech "Our next destination is Tokyo Tower" into text "Our next destination is Tokyo Tower." The converted text data is designed to be sent to a server. The device then analyzes the user's voice using an emotion engine to determine the user's emotional state. For example, the device may determine that the user is "anxious" based on the tone and speed of their voice. The results of this emotion analysis are also sent to the server.

[0216] Server Processing

[0217] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm the instructions for the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[0218] The server also receives emotional data sent from the device. If the server determines that the user is in an unstable emotional state after analyzing the emotional data, it generates feedback to provide gentle and easy-to-understand route guidance information.

[0219] Feedback and Learning

[0220] The device uses a speech synthesis engine to convert the text data sent from the server into speech and provides it to the user. For example, it might generate a voice saying, "That's right." Additionally, the server analyzes the emotional data and generates feedback, which is then provided to the user via voice. For example, if the user is feeling anxious, it might add a reassuring message such as, "Don't worry, just turn right here."

[0221] The server analyzes the questions and sentiment data collected from all users and periodically updates the algorithm to improve the system. This information helps the navigation algorithm to further evolve and provide optimal feedback to users.

[0222] Specific examples

[0223] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server refers to a map information database, confirms that the next fork should be a right turn, and generates the appropriate response, "That's right." Furthermore, if the device's emotion engine detects "anxiety" from the user's tone of voice, this information is sent to the server, which then generates an additional feedback message, "Don't worry, turn right here." This text data is again sent to the device, where it is converted into speech and provided to the user.

[0224] This process allows users to ask questions by voice while driving and receive appropriate and reliable route guidance. The server also analyzes the question data and emotion data to improve the system's accuracy and user experience. In this way, the car navigation system can continuously evolve and provide higher quality route guidance.

[0225] The processing flow will be explained below.

[0226] Specific processing steps of the program

[0227] 1. User Actions

[0228] Step 1:

[0229] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[0230] 2. Terminal Processing

[0231] Step 2:

[0232] The device receives the user's voice input and captures the audio signal with the microphone.

[0233] Step 3:

[0234] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[0235] Step 4:

[0236] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[0237] Step 5:

[0238] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[0239] Step 6:

[0240] The device again receives the user's voice input, capturing the audio signal with the microphone.

[0241] Step 7:

[0242] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[0243] Step 8:

[0244] The device sends text data and emotion data for voice emotion analysis to the server as an HTTP request.

[0245] 3. Server Processing

[0246] Step 9:

[0247] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[0248] Step 10:

[0249] The server receives the emotion data, which indicates an emotional state such as "anxiety."

[0250] Step 11:

[0251] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[0252] Step 12:

[0253] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[0254] Step 13:

[0255] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[0256] Step 14:

[0257] The server analyzes the emotion data and generates an additional feedback message if the user is in an anxious state: "Don't worry, just turn right here."

[0258] Step 15:

[0259] The server generates text data and sends the feedback message to the terminal as an HTTP response.

[0260] 4. Device reprocessing and user feedback

[0261] Step 16:

[0262] The device receives the text data from the server, including the text data "That's right" and the feedback message "Don't worry, turn right here."

[0263] Step 17:

[0264] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[0265] Step 18:

[0266] The device converts the feedback message into speech, generating a speech message that says, "Don't worry, just turn right here."

[0267] Step 19:

[0268] The device plays generated speech over the speaker, telling the user, "That's right," and "Don't worry, just turn right here."

[0269] 5. Questionnaire data collection and analysis

[0270] Step 20:

[0271] The server stores the user's question data and emotion data. The question "Should I turn right at the next fork?", the answer "Yes, that's right", the emotion "Anxious", and the feedback message "Don't worry, just turn right here" are stored in the database.

[0272] Step 21:

[0273] The server periodically analyzes the stored question data and sentiment data to identify areas where many users have experienced problems.

[0274] Step 22:

[0275] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[0276] These steps will enable users to receive accurate directions and emotion-sensitive feedback in response to voice questions while driving, while the server will continuously analyze data and update algorithms to improve the accuracy of the overall system and user experience.

[0277] Example 2

[0278] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0279] Conventional car navigation systems recognize users' voice input and provide route guidance, but voice recognition alone cannot take into account the user's emotional state, which can lead to anxiety and stress. Even if the system provides accurate answers to users' questions, it lacks emotional feedback, making it difficult to improve the user experience. Furthermore, data analysis based on users' actual usage is required to regularly improve the system.

[0280] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0281] In this invention, the server includes means for converting voice input into text data, means for transmitting the text data and user emotion data to the server, means for analyzing the text data to generate route guidance information, means for generating feedback based on the route guidance information and emotion data, means for transmitting the route guidance information and feedback as text data to the terminal, means for converting the text data into speech and providing it to the user, means for collecting and analyzing user question data and emotion data, and means for updating the route guidance algorithm based on the analysis results. This enables safe and secure route guidance that takes the user's emotional state into consideration, and also allows the system to be continuously improved based on data based on actual usage.

[0282] "User" refers to the driver or user who uses the car navigation system and provides voice input.

[0283] "Voice input" refers to the act of a user providing information by voice to set a destination or ask a question.

[0284] "Text data" refers to character string information converted from voice input by a voice recognition engine.

[0285] "Emotional Data" refers to a user's emotional state as determined by an emotion engine analyzing the user's voice input.

[0286] "Server" refers to a computer system that receives text data and emotion data, analyzes them, and generates feedback.

[0287] "Speech recognition engine" refers to a software or hardware component that converts voice input into text data.

[0288] "Emotion Engine" refers to a software or hardware component that determines a user's emotional state from their voice input.

[0289] A "natural language processing engine" refers to a computational technology that analyzes text data to understand user intent and generate appropriate responses.

[0290] "Feedback" refers to additional reassurance messages or supplemental information provided to users based on emotional data.

[0291] "Map information database" refers to a digital database that stores geographic information and route guidance information.

[0292] "Speech synthesis engine" means a software or hardware component that converts text data into speech.

[0293] "Question Data" refers to information including questions and confirmations made by a User through voice input.

[0294] A "direction guidance algorithm" refers to a set of computational methods and procedures for providing optimal route guidance to a user.

[0295] "Analysis results" refers to conclusions and findings derived from analytical methods based on collected data.

[0296] "Update" refers to the process of incorporating new data and analysis results to improve route guidance algorithms.

[0297] MODE FOR CARRYING OUT THE INVENTION

[0298] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and recognizes the user's emotional state and provides feedback. This system is composed of three entities: the user, the terminal, and the server.

[0299] User Actions

[0300] The user sets the destination by voice into the car navigation system terminal. For example, the user can say, "My next destination is Tokyo Tower." Route guidance will start, and the user can ask questions about unclear route guidance. For example, the user can ask, "Should I turn right at the next fork?"

[0301] Terminal handling

[0302] The device receives the user's voice input and converts the voice into text data using a voice recognition engine (e.g., voice recognition software). The voice input "My next destination is Tokyo Tower" is converted into text data "My next destination is Tokyo Tower." This text data is sent to the server.

[0303] The device also analyzes the user's voice data using an emotion engine (e.g., emotion analysis software) to determine the user's emotional state. For example, if the user is determined to be "anxious" based on the tone and speed of their voice, the emotion analysis results are also sent to the server.

[0304] Server Processing

[0305] The server receives the text data sent from the device and analyzes it using a natural language processing engine (e.g., natural language processing software). The server understands the intent of the user's question, and deciphers, for example, from the question "Should I turn right at the next fork?", that the user wants to confirm whether or not to turn right at the next guidance point.

[0306] The server then refers to a map information database (e.g., geographic information software) to confirm the instructions for the next route guidance point. Based on the confirmation result, it generates text data such as "That's right." This text data is then sent to the terminal.

[0307] Furthermore, the server receives and analyzes the emotional data sent from the device. If the user is in an unstable emotional state, the server generates a feedback message that provides gentle and easy-to-understand guidance information. For example, it generates a reassuring message such as, "Don't worry, just turn right here."

[0308] Device feedback and audio output

[0309] The device converts the text data sent from the server into speech using a speech synthesis engine (e.g., speech synthesis software) and provides it to the user. For example, it may play back a response such as "That's right." In addition, feedback messages based on the user's emotional state may also be provided.

[0310] System training and algorithm updates

[0311] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. This process allows the navigation algorithm to evolve and provide appropriate feedback to users.

[0312] Specific examples

[0313] For example, if a user asks "Should I turn right at the next fork?" while driving, the question is converted into text data by the device and sent to the server. The server analyzes the question, confirms that the user should turn right at the next guidance point, and generates the answer "That's right." If the user's sentiment analysis results indicate "anxiety," the server generates an additional feedback message, "Don't worry, just turn right here." These messages are sent to the device and provided to the user as audio.

[0314] Prompt Sentence Examples

[0315] "Tell me whether to turn right at the next fork. If they seem unsure, add some reassuring feedback."

[0316] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0317] Step 1:

[0318] The user sets their destination by speaking into the device. The user's voice is used as input. For example, they might say, "My next destination is Tokyo Tower." This voice data is input to the device as output.

[0319] Step 2:

[0320] The device converts the user's voice input into text data using a speech recognition engine. Voice data is used as input. Specifically, for example, the Google Speech-to-Text API is used to convert the speech "My next destination is Tokyo Tower" into text "My next destination is Tokyo Tower." This text data is generated as output.

[0321] Step 3:

[0322] The terminal sends the converted text data to the server. The text data is used as input and sent to the server as output.

[0323] Step 4:

[0324] The device analyzes the user's voice data using an emotion analysis engine to determine the user's emotional state. The voice data is used as input. Specifically, for example, emotion analysis software may be used to determine the user's emotional state as "anxiety" based on the tone and rate of the user's voice. Emotion data is generated as output.

[0325] Step 5:

[0326] The terminal transmits emotion data to the server. The emotion data is used as input and transmitted to the server as output.

[0327] Step 6:

[0328] The server receives the text data sent from the device and analyzes it using a natural language processing engine. The text data is used as input. Specifically, the natural language processing software is used to decipher the user's intent from the question, "Should I turn right at the next fork?" The analysis results are generated as output.

[0329] Step 7:

[0330] The server consults a map information database to confirm the instructions for the next route point. The analysis results are used as input. Specifically, geographic information software is used to confirm the correct route for the next route point. The output is generated as "That's right."

[0331] Step 8:

[0332] The server receives the emotion data sent from the device and analyzes the user's emotional state. The emotion data is used as input. As output, a feedback message based on the emotional state is generated. Specifically, if anxiety is detected, the server generates the feedback "Don't worry, turn right here."

[0333] Step 9:

[0334] The server sends the generated route guidance information and feedback messages to the terminal. The route guidance information and feedback are used as inputs. As outputs, these information are sent to the terminal.

[0335] Step 10:

[0336] The device converts the text data sent from the server into speech using a speech synthesis engine. The text data of route guidance information and feedback messages is used as input. Specifically, for example, Amazon Polly is used to convert text such as "That's right" and "Don't worry, turn right here" into speech. The output is speech data.

[0337] Step 11:

[0338] The terminal provides the generated audio data to the user. The audio data is used as input. As output, audio guidance is presented to the user.

[0339] Step 12:

[0340] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. The collected data is used as input. An updated algorithm is generated as output, which improves the route guidance algorithm.

[0341] (Application example 2)

[0342] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0343] Conventional car navigation systems and industrial robot mobility systems lack feedback that takes into account the user's emotional state, which can lead to anxiety and stress. Furthermore, they lack the functionality to respond appropriately to user questions and concerns that arise during route guidance. Furthermore, when a user is in an emotionally unstable state, the system needs to be able to understand this and provide appropriate, reassuring route guidance.

[0344] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0345] In this invention, the server includes means for analyzing the emotional state of the user and adding feedback information to the route guidance information, means for outputting the feedback information by voice, and means for identifying the areas where many users have problems and the emotional states of the users by collecting and analyzing the users' question data and emotional states, and updating the route guidance algorithm, thereby enabling the user to receive appropriate route guidance with peace of mind.

[0346] A "user" is the entity that operates the system, sets destinations, and asks questions.

[0347] A "destination" is the final destination of a trip set by the user.

[0348] "Directions" refers to route information and instructions provided to the user by the system.

[0349] "Voice input means" refers to a device or software that accepts voice instructions from a user.

[0350] "Means for converting into text data" refers to technology or machinery for converting voice data into text format data.

[0351] "Means for sending to the server" refers to the system or protocol for sending data from the terminal to the server.

[0352] The "means for analyzing and generating route guidance information" refers to algorithms or software that analyzes the transmitted text data and generates route information and instructions.

[0353] The "means for transmitting to the terminal as text data" is a system for transmitting the analyzed route guidance information to the terminal as text data.

[0354] "Means of converting text data into audio and providing it to users" refers to technologies and devices that convert text data into audio data and communicate it to users.

[0355] "Means for collecting and analyzing question data" refers to technology for collecting questions asked by users and analyzing them.

[0356] A "means for updating the route guidance algorithm" is a method for improving and updating the system's route provision algorithm based on collected data.

[0357] The "means for analyzing the emotional state and adding feedback information" is a system for analyzing the user's emotions and reflecting corresponding information in the route guidance.

[0358] The "means for outputting feedback information by voice" refers to a technology or device for providing the user with feedback information in the form of voice in response to the analysis results.

[0359] This invention is a system for car navigation systems and factory navigation systems that uses voice input to provide route guidance and provide feedback according to the user's emotional state. It is mainly divided into three entities: the server, the terminal, and the user, each of which plays a specific role.

[0360] server

[0361] The server is responsible for analyzing the user's voice input as text data and generating appropriate route guidance information and feedback. The server has the following functions:

[0362] 1. Speech recognition engine: The server uses a speech recognition engine (e.g., the speech_recognition library) to convert the voice data sent from the terminal into text data.

[0363] 2. Sentiment Analysis Engine: The server uses a sentiment analysis engine (e.g., the BERT model from the transformers library) to determine the emotional state of the user's voice.

[0364] 3. Route guidance generation engine: The server uses a natural language processing engine to analyze the text data and understand the user's intent. It then refers to a map information database, confirms the instructions for the next route guidance point, and generates appropriate route guidance information.

[0365] 4. Feedback generation engine: Based on the results of sentiment analysis, it generates feedback (such as reassuring messages) that adapts to the user's emotional state.

[0366] Terminal

[0367] The terminal acts as an interface between the server and the user and has the following functions:

[0368] 1. Voice acquisition means: Collects the user's voice using a microphone.

[0369] 2. Voice conversion means: Converts voice input into text data and sends it to the server.

[0370] 3. Speech synthesis means: The text data sent from the server is converted into speech using a speech synthesis engine (e.g., the pyttsx3 library) and provided to the user.

[0371] 4. Emotion analysis data transmission means: Transmits the collected emotional state data to the server.

[0372] user

[0373] Users are the users of the system who use voice to navigate and ask questions. User roles include:

[0374] 1. Voice input: Set destinations and ask driving questions by voice.

[0375] 2. Feedback reception: Receives route guidance information and feedback provided by the server and the terminal.

[0376] Specific examples

[0377] For example, if a robot moving around a factory asks, "Where is the next route?", the question is converted into text data by the terminal and sent to the server. The server analyzes the question and understands what the user wants to know. It references a map information database to generate precise instructions such as "Next left turn," and also generates feedback based on emotion analysis, such as "Don't worry, there are 50 meters until the next left turn." This information is sent to the terminal, converted into voice by a speech synthesis engine, and provided to the robot.

[0378] Prompt Sentence Examples

[0379] Use the following example as a prompt to input to your generative AI model:

[0380] What they say: "This area is congested. What's the next safe route?"

[0381] Emotion detection: Anxiety

[0382] Generated feedback: "Next left turn. Don't worry, there are 50 meters until the next left turn."

[0383] Thus, a detailed description of an embodiment of the present invention has been provided, which provides a system that efficiently processes a user's voice input and provides appropriate feedback based on the user's emotional state.

[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0385] Step 1:

[0386] The device uses a microphone as a means of acquiring voice input from the user. When the user gives voice instructions for directions or questions, the device captures this voice. For example, when the user says, "Where is the next route?", the voice is input into the microphone.

[0387] Step 2:

[0388] The device converts the acquired voice data into text data using a speech recognition engine (for example, the speech_recognition library). This text data becomes the string "Where is the next route?" The device creates this text data and prepares to send its status to the server.

[0389] Step 3:

[0390] The device then sends the text data converted by the voice recognition engine to the server. In this case, the text data "Where is the next route?" is sent to the server. At the same time, the device also analyzes the user's emotional state and sends this data to the server. For example, if "anxiety" is detected from the converted voice, that emotional data is also sent.

[0391] Step 4:

[0392] The server analyzes the received text data using a natural language processing engine to understand the user's intent. Specifically, it deciphers the text "Where is the next route?" to understand that the user needs directions. NLP libraries (e.g., transformers) are used for this analysis.

[0393] Step 5:

[0394] The server references the map information database based on the analysis results and generates the next route guidance information. For example, the server obtains information such as "Next left turn" from the map information database and generates it as text data. It also creates additional feedback information based on the received emotion data. For example, it generates feedback information such as "Don't worry, there are 50 meters until the next left turn."

[0395] Step 6:

[0396] The server sends the generated route guidance information and feedback information to the device. Specifically, text data such as "Next left turn" and "Don't worry, there are 50 meters until the next left turn" is sent to the device.

[0397] Step 7:

[0398] The device converts the received text data into speech data using a speech synthesis engine (for example, the pyttsx3 library). The device generates speech data such as "Next left turn. Don't worry, there are 50 meters until the next left turn."

[0399] Step 8:

[0400] The device provides the generated voice data to the user. Specifically, it outputs a voice message to the user through the speaker saying, "Next left turn. Don't worry, there are 50 meters until the next left turn." The user can then continue receiving route guidance with peace of mind after hearing this voice output.

[0401] Step 9:

[0402] The server analyzes the collected question data and emotion data and updates the system's algorithm. For example, if many users feel uneasy at a particular location, the system will provide more detailed route guidance information and stronger feedback. As a result, future users will receive more appropriate route guidance and feedback.

[0403] The above are the specific processing steps of the system that realizes this application example. This flow allows users to ask questions or get directions by voice input, and reach their destination while receiving reassuring feedback that reflects their emotional state.

[0404] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0405] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0406] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0407] [Second embodiment]

[0408] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0409] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0410] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0411] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0412] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0413] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0414] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0415] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0416] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0417] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0418] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0419] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0420] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0421] User Actions

[0422] The user sets the destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0423] Terminal handling

[0424] The device converts the user's voice input into text data using a speech recognition engine. For example, the speech "Should I turn right at the next fork?" is converted into text data such as "Should I turn right at the next fork?" The converted text data is designed to be sent to a server. The device then receives the text data from the server, converts it into speech using a speech synthesis engine, and provides it to the user.

[0425] Server Processing

[0426] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm whether or not to turn right at the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[0427] Feedback and Learning

[0428] The device converts the text data sent from the server into speech and provides the user with a voice response saying, "That's right." In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users have problems. Based on this information, the server updates its route guidance algorithm and provides even easier-to-understand route guidance.

[0429] Specific examples

[0430] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server references a map information database, confirms that the next fork will be a right turn, and generates the appropriate answer, "That's right." The generated answer is then sent back to the device as text data, converted into audio on the device, and provided to the user. The user can confirm that the route guidance is accurate by hearing the audio answer, "That's right."

[0431] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[0432] The processing flow will be explained below.

[0433] Specific processing steps of the program

[0434] 1. User Actions

[0435] Step 1:

[0436] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[0437] 2. Terminal Processing

[0438] Step 2:

[0439] The device receives the user's voice input and captures the audio signal with the microphone.

[0440] Step 3:

[0441] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[0442] Step 4:

[0443] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[0444] Step 5:

[0445] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[0446] Step 6:

[0447] The device again receives the user's voice input, capturing the audio signal with the microphone.

[0448] Step 7:

[0449] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[0450] Step 8:

[0451] The device sends text data to the server as an HTTP request.

[0452] 3. Server Processing

[0453] Step 9:

[0454] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[0455] Step 10:

[0456] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[0457] Step 11:

[0458] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[0459] Step 12:

[0460] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[0461] Step 13:

[0462] The server generates text data and sends it to the terminal as an HTTP response.

[0463] 4. Device reprocessing and user feedback

[0464] Step 14:

[0465] The device receives the text data from the server. The device receives the text data "That's right."

[0466] Step 15:

[0467] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[0468] Step 16:

[0469] The device generates a voice message and plays it over the speaker, telling the user "That's right."

[0470] 5. Questionnaire data collection and analysis

[0471] Step 17:

[0472] The server stores the user's question data. The question "Should I turn right at the next fork?" and the answer "Yes, that's right" are stored in the database.

[0473] Step 18:

[0474] The server periodically analyzes the stored question data to identify areas where many users have experienced problems.

[0475] Step 19:

[0476] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[0477] Through this processing step, users can ask questions by voice and receive appropriate directions, while the server analyzes the question data to improve the accuracy of the system.

[0478] Example 1

[0479] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0480] Conventional car navigation systems often provide inappropriate directions when receiving voice input and are unable to accurately understand the user's intent. Furthermore, users are unable to ask questions by voice if they are unable to understand the directions, which is inconvenient. Furthermore, the lack of a feedback function that allows many users to identify areas where they have problems and improve the navigation algorithm makes it difficult to improve the system's accuracy and user experience.

[0481] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0482] In this invention, the server includes means for converting text data received by the terminal from the server into speech and providing it to the user through an output device in the vehicle, means for analyzing the user's intentions using a natural language processing engine, and means for referencing a map information database to obtain information on the user's current location and the next intersection. This makes it possible to provide appropriate route guidance and answers to destination settings and questions entered through voice input, and to improve the accuracy of the entire system and the user experience through the feedback function.

[0483] 1. "User" means a driver or passenger who uses a car navigation system in a vehicle.

[0484] 2. "Voice input" refers to the act of the user speaking their destination or questions into the car navigation system through a microphone.

[0485] 3. "Text data" means data in the form of a string of characters converted from voice input by a voice recognition engine.

[0486] 4. "Terminal" means an electronic device, including input and output devices, installed in a vehicle in which a car navigation system is installed.

[0487] 5. "Server" means a remote computer system for receiving and analyzing text data.

[0488] 6. "Speech recognition engine" means a software or hardware mechanism that converts voice input into text data.

[0489] 7. A "natural language processing engine" is an artificial intelligence technology that analyzes text data and understands user intent.

[0490] 8. "Map information database" means a database containing road information and geographic information that the server references to generate route guidance information.

[0491] 9. "Route Guidance Information" means information regarding the direction and route generated by the server based on the user's destination and current location.

[0492] 10. "Speech synthesis engine" means a software or hardware mechanism that converts text data into speech.

[0493] 11. "Feedback" is the process of collecting and analyzing user query data to improve the system's navigation algorithms.

[0494] 12. A "guidance algorithm" is a set of procedures or calculation methods for providing users with appropriate directions and routes.

[0495] 13. "Output device" means a device, such as a speaker, that provides voice prompts or other information to the user in the vehicle.

[0496] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0497] First, the user sets their destination by voice into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0498] The device converts the user's voice input into text data using a speech recognition engine (e.g., an ambiguous speech recognition engine or a third-party speech recognition engine). After the voice is converted into text data, the text data is sent to a server via the Internet. The server then analyzes the received text data using a natural language processing engine (e.g., a generative AI model or a third-party natural language processing engine) to understand the intent of the user's question.

[0499] If the analysis of the text data determines that the user wants to confirm whether they should turn right at the next intersection, the server references a map information database (for example, a general map information database API or a map information database API from another company). The server obtains route guidance information based on the user's current location and the information about the next intersection, and generates text data for the appropriate answer, such as "That's right." This text data is then sent back to the device.

[0500] The terminal converts the text data sent from the server into speech using a speech synthesis engine (for example, a rough speech synthesis engine or a speech synthesis engine made by another company) and provides the speech to the user through an output device (such as a speaker) in the car. In this way, the user can hear the voice response "That's right."

[0501] In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users experience problems. This allows the route guidance algorithm to be updated and more user-friendly route guidance to be provided. Specifically, the collected data is analyzed using machine learning algorithms to identify areas for improvement in the system.

[0502] Specific examples

[0503] For example, if a user asks "Should I turn right at the next fork?" while driving towards their destination, the question is converted into text data "Should I turn right at the next fork?" using a voice recognition engine on the device. The text data is sent to a server via the Internet, and the server analyzes the question using a natural language processing engine to understand what the user wants to confirm. The server references a map information database, confirms that they should turn right at the next fork, and generates the appropriate answer "That's right." The generated answer is again sent to the device as text data, where it is converted into voice using a voice synthesis engine. The user can confirm that the route guidance is accurate by hearing the voice answer "That's right."

[0504] Examples of prompt statements

[0505] "Utterance: 'Should I turn right at the next fork?'"

[0506] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[0507] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0508] Step 1: User Speech Input

[0509] The user sets the destination by speaking into the car navigation system terminal. The input is the user's voice, which can be a specific destination such as "My next destination is Tokyo Tower" or a question such as "Should I turn right at the next fork?" The voice is received through the terminal's microphone input device.

[0510] Step 2: Voice Recognition

[0511] The device converts the user's voice input into text data using a voice recognition engine (e.g., voice recognition software). The input is the user's voice signal, and the output is text data in the form of a string: "Should I turn right at the next fork?" The voice data is converted into a digital signal, and phonemic and linguistic analysis is performed before being output as text.

[0512] Step 3: Send text data to the server

[0513] The device sends the text data converted by the speech recognition engine to the server via the Internet. The input is the text data "Should I turn right at the next fork?" and the output is data in the form of network packets. The device's network module is responsible for this communication and sends the data to the server via the HTTPS protocol.

[0514] Step 4: Analyzing the text data

[0515] The server analyzes the received text data using a natural language processing engine (e.g., a generative AI model). The input is the text data "Should I turn right at the next fork?" and the output is structured data that represents the user's intent. The server performs tokenization and grammatical analysis on the text data to understand what the user wants to confirm.

[0516] Step 5: Query the map database

[0517] The server references a map information database (e.g., a map information API) based on the analysis results. The input is the user's current location and information about the next intersection, and the output is route guidance information including the action to be taken at the next intersection (e.g., turn right). The server sends a query to the map information database to obtain the required information.

[0518] Step 6: Generate directions

[0519] The server generates appropriate route guidance information based on information obtained from the map information database. The input is route guidance information obtained from the map database, and the output is the text-format route guidance information "That's right" to be provided to the user. The server composes this as text data in character string format.

[0520] Step 7: Sending text data from the server to the device

[0521] The generated text data is resent from the server to the terminal. The input is the generated text data "That's right," and the output is data in the form of network packets. The server's network module is responsible for this communication.

[0522] Step 8: Text-to-Speech

[0523] The device converts the text data received from the server into speech using a speech synthesis engine (e.g., text-to-speech software). The input is the text data "That's right," and the output is an audio file (e.g., audio data in WAV or MP3 format). The text data is input into the speech synthesis engine, and output as an audio file.

[0524] Step 9: User voice guidance

[0525] The terminal provides the generated voice data to the user through the car's speaker. The input is the voice file "That's right," and the output is the voice played through the speaker. The user listens to this and confirms the next action.

[0526] Step 10: Feedback and learning

[0527] The server collects question data from all users and periodically analyzes it. The input is the collected question data (e.g., text data history), and the output is feedback data containing improvements to the route guidance algorithm. A machine learning algorithm is used to analyze the data and identify improvements to the system. This allows the server to update the route guidance algorithm and improve the accuracy of the entire system.

[0528] (Application example 1)

[0529] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0530] In food delivery operations, it is difficult for drivers to set destinations and check routes safely and efficiently while driving. Manual resetting and checking while driving is dangerous, so there is a need for voice interaction.

[0531] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0532] In this invention, the server includes means for a user to input voice commands to set a destination and receive route guidance, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and generating route guidance information, means for transmitting the route guidance information as text data to a terminal, means for converting the text data into voice and providing it to the user, means for collecting and analyzing user question data, means for updating a route guidance algorithm based on the analysis results, means for a delivery driver to set a destination and check a route by voice while driving, means for providing the route confirmation results by voice in real time, and means for generating next destination and route information using document generation technology. This enables drivers to set a destination and check a route by voice alone while driving, improving the safety and efficiency of their work.

[0533] "Voice input" is a method by which a user communicates verbal instructions or questions to a system.

[0534] "Text data" is data that has been converted from voice input into text information and is used for analysis and processing.

[0535] A "server" is a computing device that receives transmitted text data, analyzes it, and generates a response.

[0536] "Route guidance information" is information that indicates the route and direction for the user to reach the destination.

[0537] "Device" means a device operated by a user that receives audio input and provides audio output.

[0538] A "voice recognition engine" is software that analyzes voice input and converts it into corresponding text data.

[0539] A "speech synthesis engine" is software that analyzes text data and generates corresponding speech.

[0540] "Question data" refers to information about confirmations or questions that users make to the system.

[0541] "Parsing" is the process of understanding received text data and deciphering its intent.

[0542] A "route guidance algorithm" is a calculation procedure for providing optimal route information to users.

[0543] A "delivery driver" is a driver whose job is to deliver goods to customers.

[0544] "Real-time" refers to near-instant processing or response.

[0545] "Document generation technology" is technology that generates appropriate text responses to user questions.

[0546] "Driving" refers to the state in which the delivery driver is operating the vehicle.

[0547] This invention is a system that allows food delivery drivers to easily set destinations and check routes by voice while driving. To implement the invention, the following specific hardware and software configurations are required.

[0548] Hardware

[0549] Smartphone (device)

[0550] microphone

[0551] speaker

[0552] software

[0553] Speech recognition engine: SpeechRecognition (Python library)

[0554] Speech synthesis engine: gTTS (Google Text-to-Speech)

[0555] API request tool: requests (Python library)

[0556] Audio output playback utility: mpg321

[0557] Program processing

[0558] 1. User Operation

[0559] The user (delivery driver) speaks instructions into their smartphone, such as "Where's the next delivery destination?" or "Where's the next right turn?"

[0560] 2. Voice Input and Speech Recognition

[0561] The device receives the user's voice input through a microphone. The speech recognition engine (SpeechRecognition) converts this voice input into text data. For example, the voice "Where is the next delivery?" is converted into text data "Where is the next delivery?"

[0562] 3. Sending to the server and analyzing

[0563] The converted text data is sent from the device to the server. The server receives this text data and analyzes it using a natural language processing engine. For example, from the question "Where is the next delivery?", it understands that the user wants to check information about the next delivery address.

[0564] 4. Generating Route Guidance Information

[0565] Based on the analysis results, the server references a map information database and generates information about the next delivery destination and route. The generated route guidance information is sent to the terminal as text data such as "The next delivery destination is ____."

[0566] 5. Speech synthesis and delivery

[0567] The device converts the received text data into speech using a speech synthesis engine (gTTS), and provides the user with a voice message such as, "The next delivery destination is ____." This allows the user to check the next delivery destination without looking at the screen while driving.

[0568] Specific examples

[0569] Specifically, when a delivery driver asks, "Where's the next delivery?", the question is converted by the device into text data, "Where's the next delivery?" The text data is sent to the server, which analyzes the question and confirms the next delivery address. The server generates a response such as "The next delivery address is: XXX" and sends it back to the device as text data. The response is then converted into speech on the device and provided to the user.

[0570] Example prompts for generative AI models

[0571] "When a user asks, 'Where is my next delivery?', please search for the next delivery destination in a map database and generate a sentence that provides appropriate route guidance."

[0572] This will create a system that allows food delivery drivers to safely and efficiently check their destination and set their route while driving.

[0573] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0574] Step 1:

[0575] Using the voice input function of the smartphone, the user asks aloud, "Where is the next delivery destination?" The input is voice data, and the user's question is captured by the smartphone's microphone.

[0576] Step 2:

[0577] The device uses a speech recognition engine (SpeechRecognition) to convert voice data into text data. It analyzes the voice data and generates corresponding text data. For example, a speech saying "Where is the next delivery?" is converted into text data "Where is the next delivery?". This text data is used in the next processing step.

[0578] Step 3:

[0579] The converted text data is sent from the device to the server. The device uses an API request tool (requests) to send this text data to the server. The input is text data, and the output is an HTTP request to the server.

[0580] Step 4:

[0581] The server analyzes the text data received from the device. It uses a natural language processing engine to analyze the received text data and understand the intent of the user's question. Based on the analysis results, it generates a query to search for information on the next delivery destination. This query is used to reference a map information database.

[0582] Step 5:

[0583] The server uses the generated query to refer to a map information database and obtain information about the next delivery destination. The information obtained from the database includes the delivery destination address and route information. The obtained information is saved in text format and passed to the next processing step. For example, a text such as "The next delivery destination is at address: XXX" is generated.

[0584] Step 6:

[0585] The server uses document generation technology to generate an appropriate response based on the acquired delivery destination information. It uses a generative AI model (text generation engine) to create an appropriate response to the question. The response is specific text data such as "The next delivery destination is address: XXX." This text data is then sent back to the terminal.

[0586] Step 7:

[0587] The device converts the received text data into speech using a speech synthesis engine (gTTS). The text data is analyzed and the corresponding voice data is generated. The generated voice data is played back through the smartphone's speaker. For example, the user can hear a voice saying, "The next delivery address is: XXX."

[0588] Step 8:

[0589] The user continues driving based on the next delivery destination information provided by voice. The user can confirm the voice response and follow the appropriate route to the next delivery destination. In this way, a system is realized that allows users to confirm destinations and set routes safely and efficiently while driving.

[0590] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0591] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and also recognizes and provides feedback on the user's emotional state. This system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0592] User Actions

[0593] The user sets their destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0594] Terminal handling

[0595] The device converts the user's voice input into text data using a speech recognition engine. For example, the device converts the speech "Our next destination is Tokyo Tower" into text "Our next destination is Tokyo Tower." The converted text data is designed to be sent to a server. The device then analyzes the user's voice using an emotion engine to determine the user's emotional state. For example, the device may determine that the user is "anxious" based on the tone and speed of their voice. The results of this emotion analysis are also sent to the server.

[0596] Server Processing

[0597] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm the instructions for the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[0598] The server also receives emotional data sent from the device. If the server determines that the user is in an unstable emotional state after analyzing the emotional data, it generates feedback to provide gentle and easy-to-understand route guidance information.

[0599] Feedback and Learning

[0600] The device uses a speech synthesis engine to convert the text data sent from the server into speech and provides it to the user. For example, it might generate a voice saying, "That's right." Additionally, the server analyzes the emotional data and generates feedback, which is then provided to the user via voice. For example, if the user is feeling anxious, it might add a reassuring message such as, "Don't worry, just turn right here."

[0601] The server analyzes the questions and sentiment data collected from all users and periodically updates the algorithm to improve the system. This information helps the navigation algorithm to further evolve and provide optimal feedback to users.

[0602] Specific examples

[0603] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server refers to a map information database, confirms that the next fork should be a right turn, and generates the appropriate response, "That's right." Furthermore, if the device's emotion engine detects "anxiety" from the user's tone of voice, this information is sent to the server, which then generates an additional feedback message, "Don't worry, turn right here." This text data is again sent to the device, where it is converted into speech and provided to the user.

[0604] This process allows users to ask questions by voice while driving and receive appropriate and reliable route guidance. The server also analyzes the question data and emotion data to improve the system's accuracy and user experience. In this way, the car navigation system can continuously evolve and provide higher quality route guidance.

[0605] The processing flow will be explained below.

[0606] Specific processing steps of the program

[0607] 1. User Actions

[0608] Step 1:

[0609] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[0610] 2. Terminal Processing

[0611] Step 2:

[0612] The device receives the user's voice input and captures the audio signal with the microphone.

[0613] Step 3:

[0614] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[0615] Step 4:

[0616] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[0617] Step 5:

[0618] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[0619] Step 6:

[0620] The device again receives the user's voice input, capturing the audio signal with the microphone.

[0621] Step 7:

[0622] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[0623] Step 8:

[0624] The device sends text data and emotion data for voice emotion analysis to the server as an HTTP request.

[0625] 3. Server Processing

[0626] Step 9:

[0627] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[0628] Step 10:

[0629] The server receives the emotion data, which indicates an emotional state such as "anxiety."

[0630] Step 11:

[0631] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[0632] Step 12:

[0633] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[0634] Step 13:

[0635] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[0636] Step 14:

[0637] The server analyzes the emotion data and generates an additional feedback message if the user is in an anxious state: "Don't worry, just turn right here."

[0638] Step 15:

[0639] The server generates text data and sends the feedback message to the terminal as an HTTP response.

[0640] 4. Device reprocessing and user feedback

[0641] Step 16:

[0642] The device receives the text data from the server, including the text data "That's right" and the feedback message "Don't worry, turn right here."

[0643] Step 17:

[0644] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[0645] Step 18:

[0646] The device converts the feedback message into speech, generating a speech message that says, "Don't worry, just turn right here."

[0647] Step 19:

[0648] The device plays generated speech over the speaker, telling the user, "That's right," and "Don't worry, just turn right here."

[0649] 5. Questionnaire data collection and analysis

[0650] Step 20:

[0651] The server stores the user's question data and emotion data. The question "Should I turn right at the next fork?", the answer "Yes, that's right", the emotion "Anxious", and the feedback message "Don't worry, just turn right here" are stored in the database.

[0652] Step 21:

[0653] The server periodically analyzes the stored question data and sentiment data to identify areas where many users have experienced problems.

[0654] Step 22:

[0655] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[0656] These steps will enable users to receive accurate directions and emotion-sensitive feedback in response to voice questions while driving, while the server will continuously analyze data and update algorithms to improve the accuracy of the overall system and user experience.

[0657] Example 2

[0658] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0659] Conventional car navigation systems recognize users' voice input and provide route guidance, but voice recognition alone cannot take into account the user's emotional state, which can lead to anxiety and stress. Even if the system provides accurate answers to users' questions, it lacks emotional feedback, making it difficult to improve the user experience. Furthermore, data analysis based on users' actual usage is required to regularly improve the system.

[0660] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0661] In this invention, the server includes means for converting voice input into text data, means for transmitting the text data and user emotion data to the server, means for analyzing the text data to generate route guidance information, means for generating feedback based on the route guidance information and emotion data, means for transmitting the route guidance information and feedback as text data to the terminal, means for converting the text data into speech and providing it to the user, means for collecting and analyzing user question data and emotion data, and means for updating the route guidance algorithm based on the analysis results. This enables safe and secure route guidance that takes the user's emotional state into consideration, and also allows the system to be continuously improved based on data based on actual usage.

[0662] "User" refers to the driver or user who uses the car navigation system and provides voice input.

[0663] "Voice input" refers to the act of a user providing information by voice to set a destination or ask a question.

[0664] "Text data" refers to character string information converted from voice input by a voice recognition engine.

[0665] "Emotional Data" refers to a user's emotional state as determined by an emotion engine analyzing the user's voice input.

[0666] "Server" refers to a computer system that receives text data and emotion data, analyzes them, and generates feedback.

[0667] "Speech recognition engine" refers to a software or hardware component that converts voice input into text data.

[0668] "Emotion Engine" refers to a software or hardware component that determines a user's emotional state from their voice input.

[0669] A "natural language processing engine" refers to a computational technology that analyzes text data to understand user intent and generate appropriate responses.

[0670] "Feedback" refers to additional reassurance messages or supplemental information provided to users based on emotional data.

[0671] "Map information database" refers to a digital database that stores geographic information and route guidance information.

[0672] "Speech synthesis engine" means a software or hardware component that converts text data into speech.

[0673] "Question Data" refers to information including questions and confirmations made by a User through voice input.

[0674] A "direction guidance algorithm" refers to a set of computational methods and procedures for providing optimal route guidance to a user.

[0675] "Analysis results" refers to conclusions and findings derived from analytical methods based on collected data.

[0676] "Update" refers to the process of incorporating new data and analysis results to improve route guidance algorithms.

[0677] MODE FOR CARRYING OUT THE INVENTION

[0678] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and recognizes the user's emotional state and provides feedback. This system is composed of three entities: the user, the terminal, and the server.

[0679] User Actions

[0680] The user sets the destination by voice into the car navigation system terminal. For example, the user can say, "My next destination is Tokyo Tower." Route guidance will start, and the user can ask questions about unclear route guidance. For example, the user can ask, "Should I turn right at the next fork?"

[0681] Terminal handling

[0682] The device receives the user's voice input and converts the voice into text data using a voice recognition engine (e.g., voice recognition software). The voice input "My next destination is Tokyo Tower" is converted into text data "My next destination is Tokyo Tower." This text data is sent to the server.

[0683] The device also analyzes the user's voice data using an emotion engine (e.g., emotion analysis software) to determine the user's emotional state. For example, if the user is determined to be "anxious" based on the tone and speed of their voice, the emotion analysis results are also sent to the server.

[0684] Server Processing

[0685] The server receives the text data sent from the device and analyzes it using a natural language processing engine (e.g., natural language processing software). The server understands the intent of the user's question, and deciphers, for example, from the question "Should I turn right at the next fork?", that the user wants to confirm whether or not to turn right at the next guidance point.

[0686] The server then refers to a map information database (e.g., geographic information software) to confirm the instructions for the next route guidance point. Based on the confirmation result, it generates text data such as "That's right." This text data is then sent to the terminal.

[0687] Furthermore, the server receives and analyzes the emotional data sent from the device. If the user is in an unstable emotional state, the server generates a feedback message that provides gentle and easy-to-understand guidance information. For example, it generates a reassuring message such as, "Don't worry, just turn right here."

[0688] Device feedback and audio output

[0689] The device converts the text data sent from the server into speech using a speech synthesis engine (e.g., speech synthesis software) and provides it to the user. For example, it may play back a response such as "That's right." In addition, feedback messages based on the user's emotional state may also be provided.

[0690] System training and algorithm updates

[0691] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. This process allows the navigation algorithm to evolve and provide appropriate feedback to users.

[0692] Specific examples

[0693] For example, if a user asks "Should I turn right at the next fork?" while driving, the question is converted into text data by the device and sent to the server. The server analyzes the question, confirms that the user should turn right at the next guidance point, and generates the answer "That's right." If the user's sentiment analysis results indicate "anxiety," the server generates an additional feedback message, "Don't worry, just turn right here." These messages are sent to the device and provided to the user as audio.

[0694] Prompt Sentence Examples

[0695] "Tell me whether to turn right at the next fork. If they seem unsure, add some reassuring feedback."

[0696] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0697] Step 1:

[0698] The user sets their destination by speaking into the device. The user's voice is used as input. For example, they might say, "My next destination is Tokyo Tower." This voice data is input to the device as output.

[0699] Step 2:

[0700] The device converts the user's voice input into text data using a speech recognition engine. Voice data is used as input. Specifically, for example, the Google Speech-to-Text API is used to convert the speech "My next destination is Tokyo Tower" into text "My next destination is Tokyo Tower." This text data is generated as output.

[0701] Step 3:

[0702] The terminal sends the converted text data to the server. The text data is used as input and sent to the server as output.

[0703] Step 4:

[0704] The device analyzes the user's voice data using an emotion analysis engine to determine the user's emotional state. The voice data is used as input. Specifically, for example, emotion analysis software may be used to determine the user's emotional state as "anxiety" based on the tone and rate of the user's voice. Emotion data is generated as output.

[0705] Step 5:

[0706] The terminal transmits emotion data to the server. The emotion data is used as input and transmitted to the server as output.

[0707] Step 6:

[0708] The server receives the text data sent from the device and analyzes it using a natural language processing engine. The text data is used as input. Specifically, the natural language processing software is used to decipher the user's intent from the question, "Should I turn right at the next fork?" The analysis results are generated as output.

[0709] Step 7:

[0710] The server consults a map information database to confirm the instructions for the next route point. The analysis results are used as input. Specifically, geographic information software is used to confirm the correct route for the next route point. The output is generated as "That's right."

[0711] Step 8:

[0712] The server receives the emotion data sent from the device and analyzes the user's emotional state. The emotion data is used as input. As output, a feedback message based on the emotional state is generated. Specifically, if anxiety is detected, the server generates the feedback "Don't worry, turn right here."

[0713] Step 9:

[0714] The server sends the generated route guidance information and feedback messages to the terminal. The route guidance information and feedback are used as inputs. As outputs, these information are sent to the terminal.

[0715] Step 10:

[0716] The device converts the text data sent from the server into speech using a speech synthesis engine. The text data of route guidance information and feedback messages is used as input. Specifically, for example, Amazon Polly is used to convert text such as "That's right" and "Don't worry, turn right here" into speech. The output is speech data.

[0717] Step 11:

[0718] The terminal provides the generated audio data to the user. The audio data is used as input. As output, audio guidance is presented to the user.

[0719] Step 12:

[0720] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. The collected data is used as input. An updated algorithm is generated as output, which improves the route guidance algorithm.

[0721] (Application example 2)

[0722] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0723] Conventional car navigation systems and industrial robot mobility systems lack feedback that takes into account the user's emotional state, which can lead to anxiety and stress. Furthermore, they lack the functionality to respond appropriately to user questions and concerns that arise during route guidance. Furthermore, when a user is in an emotionally unstable state, the system needs to be able to understand this and provide appropriate, reassuring route guidance.

[0724] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0725] In this invention, the server includes means for analyzing the emotional state of the user and adding feedback information to the route guidance information, means for outputting the feedback information by voice, and means for identifying the areas where many users have problems and the emotional states of the users by collecting and analyzing the users' question data and emotional states, and updating the route guidance algorithm, thereby enabling the user to receive appropriate route guidance with peace of mind.

[0726] A "user" is the entity that operates the system, sets destinations, and asks questions.

[0727] A "destination" is the final destination of a trip set by the user.

[0728] "Directions" refers to route information and instructions provided to the user by the system.

[0729] "Voice input means" refers to a device or software that accepts voice instructions from a user.

[0730] "Means for converting into text data" refers to technology or machinery for converting voice data into text format data.

[0731] "Means for sending to the server" refers to the system or protocol for sending data from the terminal to the server.

[0732] The "means for analyzing and generating route guidance information" refers to algorithms or software that analyzes the transmitted text data and generates route information and instructions.

[0733] The "means for transmitting to the terminal as text data" is a system for transmitting the analyzed route guidance information to the terminal as text data.

[0734] "Means of converting text data into audio and providing it to users" refers to technologies and devices that convert text data into audio data and communicate it to users.

[0735] "Means for collecting and analyzing question data" refers to technology for collecting questions asked by users and analyzing them.

[0736] A "means for updating the route guidance algorithm" is a method for improving and updating the system's route provision algorithm based on collected data.

[0737] The "means for analyzing the emotional state and adding feedback information" is a system for analyzing the user's emotions and reflecting corresponding information in the route guidance.

[0738] The "means for outputting feedback information by voice" refers to a technology or device for providing the user with feedback information in the form of voice in response to the analysis results.

[0739] This invention is a system for car navigation systems and factory navigation systems that uses voice input to provide route guidance and provide feedback according to the user's emotional state. It is mainly divided into three entities: the server, the terminal, and the user, each of which plays a specific role.

[0740] server

[0741] The server is responsible for analyzing the user's voice input as text data and generating appropriate route guidance information and feedback. The server has the following functions:

[0742] 1. Speech recognition engine: The server uses a speech recognition engine (e.g., the speech_recognition library) to convert the voice data sent from the terminal into text data.

[0743] 2. Sentiment Analysis Engine: The server uses a sentiment analysis engine (e.g., the BERT model from the transformers library) to determine the emotional state of the user's voice.

[0744] 3. Route guidance generation engine: The server uses a natural language processing engine to analyze the text data and understand the user's intent. It then refers to a map information database, confirms the instructions for the next route guidance point, and generates appropriate route guidance information.

[0745] 4. Feedback generation engine: Based on the results of sentiment analysis, it generates feedback (such as reassuring messages) that adapts to the user's emotional state.

[0746] Terminal

[0747] The terminal acts as an interface between the server and the user and has the following functions:

[0748] 1. Voice acquisition means: Collects the user's voice using a microphone.

[0749] 2. Voice conversion means: Converts voice input into text data and sends it to the server.

[0750] 3. Speech synthesis means: The text data sent from the server is converted into speech using a speech synthesis engine (e.g., the pyttsx3 library) and provided to the user.

[0751] 4. Emotion analysis data transmission means: Transmits the collected emotional state data to the server.

[0752] user

[0753] Users are the users of the system who use voice to navigate and ask questions. User roles include:

[0754] 1. Voice input: Set destinations and ask driving questions by voice.

[0755] 2. Feedback reception: Receives route guidance information and feedback provided by the server and the terminal.

[0756] Specific examples

[0757] For example, if a robot moving around a factory asks, "Where is the next route?", the question is converted into text data by the terminal and sent to the server. The server analyzes the question and understands what the user wants to know. It references a map information database to generate precise instructions such as "Next left turn," and also generates feedback based on emotion analysis, such as "Don't worry, there are 50 meters until the next left turn." This information is sent to the terminal, converted into voice by a speech synthesis engine, and provided to the robot.

[0758] Prompt Sentence Examples

[0759] Use the following example as a prompt to input to your generative AI model:

[0760] What they say: "This area is congested. What's the next safe route?"

[0761] Emotion detection: Anxiety

[0762] Generated feedback: "Next left turn. Don't worry, there are 50 meters until the next left turn."

[0763] Thus, a detailed description of an embodiment of the present invention has been provided, which provides a system that efficiently processes a user's voice input and provides appropriate feedback based on the user's emotional state.

[0764] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0765] Step 1:

[0766] The device uses a microphone as a means of acquiring voice input from the user. When the user gives voice instructions for directions or questions, the device captures this voice. For example, when the user says, "Where is the next route?", the voice is input into the microphone.

[0767] Step 2:

[0768] The device converts the acquired voice data into text data using a speech recognition engine (for example, the speech_recognition library). This text data becomes the string "Where is the next route?" The device creates this text data and prepares to send its status to the server.

[0769] Step 3:

[0770] The device then sends the text data converted by the voice recognition engine to the server. In this case, the text data "Where is the next route?" is sent to the server. At the same time, the device also analyzes the user's emotional state and sends this data to the server. For example, if "anxiety" is detected from the converted voice, that emotional data is also sent.

[0771] Step 4:

[0772] The server analyzes the received text data using a natural language processing engine to understand the user's intent. Specifically, it deciphers the text "Where is the next route?" to understand that the user needs directions. NLP libraries (e.g., transformers) are used for this analysis.

[0773] Step 5:

[0774] The server references the map information database based on the analysis results and generates the next route guidance information. For example, the server obtains information such as "Next left turn" from the map information database and generates it as text data. It also creates additional feedback information based on the received emotion data. For example, it generates feedback information such as "Don't worry, there are 50 meters until the next left turn."

[0775] Step 6:

[0776] The server sends the generated route guidance information and feedback information to the device. Specifically, text data such as "Next left turn" and "Don't worry, there are 50 meters until the next left turn" is sent to the device.

[0777] Step 7:

[0778] The device converts the received text data into speech data using a speech synthesis engine (for example, the pyttsx3 library). The device generates speech data such as "Next left turn. Don't worry, there are 50 meters until the next left turn."

[0779] Step 8:

[0780] The device provides the generated voice data to the user. Specifically, it outputs a voice message to the user through the speaker saying, "Next left turn. Don't worry, there are 50 meters until the next left turn." The user can then continue receiving route guidance with peace of mind after hearing this voice output.

[0781] Step 9:

[0782] The server analyzes the collected question data and emotion data and updates the system's algorithm. For example, if many users feel uneasy at a particular location, the system will provide more detailed route guidance information and stronger feedback. As a result, future users will receive more appropriate route guidance and feedback.

[0783] The above are the specific processing steps of the system that realizes this application example. This flow allows users to ask questions or get directions by voice input, and reach their destination while receiving reassuring feedback that reflects their emotional state.

[0784] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0785] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0786] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0787] [Third embodiment]

[0788] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0789] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0790] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0791] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0792] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0793] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0794] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0795] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0796] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0797] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0798] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0799] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0800] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0801] User Actions

[0802] The user sets the destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0803] Terminal handling

[0804] The device converts the user's voice input into text data using a speech recognition engine. For example, the speech "Should I turn right at the next fork?" is converted into text data such as "Should I turn right at the next fork?" The converted text data is designed to be sent to a server. The device then receives the text data from the server, converts it into speech using a speech synthesis engine, and provides it to the user.

[0805] Server Processing

[0806] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm whether or not to turn right at the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[0807] Feedback and Learning

[0808] The device converts the text data sent from the server into speech and provides the user with a voice response saying, "That's right." In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users have problems. Based on this information, the server updates its route guidance algorithm and provides even easier-to-understand route guidance.

[0809] Specific examples

[0810] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server references a map information database, confirms that the next fork will be a right turn, and generates the appropriate answer, "That's right." The generated answer is then sent back to the device as text data, converted into audio on the device, and provided to the user. The user can confirm that the route guidance is accurate by hearing the audio answer, "That's right."

[0811] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[0812] The processing flow will be explained below.

[0813] Specific processing steps of the program

[0814] 1. User Actions

[0815] Step 1:

[0816] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[0817] 2. Terminal Processing

[0818] Step 2:

[0819] The device receives the user's voice input and captures the audio signal with the microphone.

[0820] Step 3:

[0821] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[0822] Step 4:

[0823] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[0824] Step 5:

[0825] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[0826] Step 6:

[0827] The device again receives the user's voice input, capturing the audio signal with the microphone.

[0828] Step 7:

[0829] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[0830] Step 8:

[0831] The device sends text data to the server as an HTTP request.

[0832] 3. Server Processing

[0833] Step 9:

[0834] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[0835] Step 10:

[0836] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[0837] Step 11:

[0838] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[0839] Step 12:

[0840] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[0841] Step 13:

[0842] The server generates text data and sends it to the terminal as an HTTP response.

[0843] 4. Device reprocessing and user feedback

[0844] Step 14:

[0845] The device receives the text data from the server. The device receives the text data "That's right."

[0846] Step 15:

[0847] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[0848] Step 16:

[0849] The device generates a voice message and plays it over the speaker, telling the user "That's right."

[0850] 5. Questionnaire data collection and analysis

[0851] Step 17:

[0852] The server stores the user's question data. The question "Should I turn right at the next fork?" and the answer "Yes, that's right" are stored in the database.

[0853] Step 18:

[0854] The server periodically analyzes the stored question data to identify areas where many users have experienced problems.

[0855] Step 19:

[0856] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[0857] Through this processing step, users can ask questions by voice and receive appropriate directions, while the server analyzes the question data to improve the accuracy of the system.

[0858] Example 1

[0859] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0860] Conventional car navigation systems often provide inappropriate directions when receiving voice input and are unable to accurately understand the user's intent. Furthermore, users are unable to ask questions by voice if they are unable to understand the directions, which is inconvenient. Furthermore, the lack of a feedback function that allows many users to identify areas where they have problems and improve the navigation algorithm makes it difficult to improve the system's accuracy and user experience.

[0861] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0862] In this invention, the server includes means for converting text data received by the terminal from the server into speech and providing it to the user through an output device in the vehicle, means for analyzing the user's intentions using a natural language processing engine, and means for referencing a map information database to obtain information on the user's current location and the next intersection. This makes it possible to provide appropriate route guidance and answers to destination settings and questions entered through voice input, and to improve the accuracy of the entire system and the user experience through the feedback function.

[0863] 1. "User" means a driver or passenger who uses a car navigation system in a vehicle.

[0864] 2. "Voice input" refers to the act of the user speaking their destination or questions into the car navigation system through a microphone.

[0865] 3. "Text data" means data in the form of a string of characters converted from voice input by a voice recognition engine.

[0866] 4. "Terminal" means an electronic device, including input and output devices, installed in a vehicle in which a car navigation system is installed.

[0867] 5. "Server" means a remote computer system for receiving and analyzing text data.

[0868] 6. "Speech recognition engine" means a software or hardware mechanism that converts voice input into text data.

[0869] 7. A "natural language processing engine" is an artificial intelligence technology that analyzes text data and understands user intent.

[0870] 8. "Map information database" means a database containing road information and geographic information that the server references to generate route guidance information.

[0871] 9. "Route Guidance Information" means information regarding the direction and route generated by the server based on the user's destination and current location.

[0872] 10. "Speech synthesis engine" means a software or hardware mechanism that converts text data into speech.

[0873] 11. "Feedback" is the process of collecting and analyzing user query data to improve the system's navigation algorithms.

[0874] 12. A "guidance algorithm" is a set of procedures or calculation methods for providing users with appropriate directions and routes.

[0875] 13. "Output device" means a device, such as a speaker, that provides voice prompts or other information to the user in the vehicle.

[0876] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0877] First, the user sets their destination by voice into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0878] The device converts the user's voice input into text data using a speech recognition engine (e.g., an ambiguous speech recognition engine or a third-party speech recognition engine). After the voice is converted into text data, the text data is sent to a server via the Internet. The server then analyzes the received text data using a natural language processing engine (e.g., a generative AI model or a third-party natural language processing engine) to understand the intent of the user's question.

[0879] If the analysis of the text data determines that the user wants to confirm whether they should turn right at the next intersection, the server references a map information database (for example, a general map information database API or a map information database API from another company). The server obtains route guidance information based on the user's current location and the information about the next intersection, and generates text data for the appropriate answer, such as "That's right." This text data is then sent back to the device.

[0880] The terminal converts the text data sent from the server into speech using a speech synthesis engine (for example, a rough speech synthesis engine or a speech synthesis engine made by another company) and provides the speech to the user through an output device (such as a speaker) in the car. In this way, the user can hear the voice response "That's right."

[0881] In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users experience problems. This allows the route guidance algorithm to be updated and more user-friendly route guidance to be provided. Specifically, the collected data is analyzed using machine learning algorithms to identify areas for improvement in the system.

[0882] Specific examples

[0883] For example, if a user asks "Should I turn right at the next fork?" while driving towards their destination, the question is converted into text data "Should I turn right at the next fork?" using a voice recognition engine on the device. The text data is sent to a server via the Internet, and the server analyzes the question using a natural language processing engine to understand what the user wants to confirm. The server references a map information database, confirms that they should turn right at the next fork, and generates the appropriate answer "That's right." The generated answer is again sent to the device as text data, where it is converted into voice using a voice synthesis engine. The user can confirm that the route guidance is accurate by hearing the voice answer "That's right."

[0884] Examples of prompt statements

[0885] "Utterance: 'Should I turn right at the next fork?'"

[0886] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[0887] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0888] Step 1: User Speech Input

[0889] The user sets the destination by speaking into the car navigation system terminal. The input is the user's voice, which can be a specific destination such as "My next destination is Tokyo Tower" or a question such as "Should I turn right at the next fork?" The voice is received through the terminal's microphone input device.

[0890] Step 2: Voice Recognition

[0891] The device converts the user's voice input into text data using a voice recognition engine (e.g., voice recognition software). The input is the user's voice signal, and the output is text data in the form of a string: "Should I turn right at the next fork?" The voice data is converted into a digital signal, and phonemic and linguistic analysis is performed before being output as text.

[0892] Step 3: Send text data to the server

[0893] The device sends the text data converted by the speech recognition engine to the server via the Internet. The input is the text data "Should I turn right at the next fork?" and the output is data in the form of network packets. The device's network module is responsible for this communication and sends the data to the server via the HTTPS protocol.

[0894] Step 4: Analyzing the text data

[0895] The server analyzes the received text data using a natural language processing engine (e.g., a generative AI model). The input is the text data "Should I turn right at the next fork?" and the output is structured data that represents the user's intent. The server performs tokenization and grammatical analysis on the text data to understand what the user wants to confirm.

[0896] Step 5: Query the map database

[0897] The server references a map information database (e.g., a map information API) based on the analysis results. The input is the user's current location and information about the next intersection, and the output is route guidance information including the action to be taken at the next intersection (e.g., turn right). The server sends a query to the map information database to obtain the required information.

[0898] Step 6: Generate directions

[0899] The server generates appropriate route guidance information based on information obtained from the map information database. The input is route guidance information obtained from the map database, and the output is the text-format route guidance information "That's right" to be provided to the user. The server composes this as text data in character string format.

[0900] Step 7: Sending text data from the server to the device

[0901] The generated text data is resent from the server to the terminal. The input is the generated text data "That's right," and the output is data in the form of network packets. The server's network module is responsible for this communication.

[0902] Step 8: Text-to-Speech

[0903] The device converts the text data received from the server into speech using a speech synthesis engine (e.g., text-to-speech software). The input is the text data "That's right," and the output is an audio file (e.g., audio data in WAV or MP3 format). The text data is input into the speech synthesis engine, and output as an audio file.

[0904] Step 9: User voice guidance

[0905] The terminal provides the generated voice data to the user through the car's speaker. The input is the voice file "That's right," and the output is the voice played through the speaker. The user listens to this and confirms the next action.

[0906] Step 10: Feedback and learning

[0907] The server collects question data from all users and periodically analyzes it. The input is the collected question data (e.g., text data history), and the output is feedback data containing improvements to the route guidance algorithm. A machine learning algorithm is used to analyze the data and identify improvements to the system. This allows the server to update the route guidance algorithm and improve the accuracy of the entire system.

[0908] (Application example 1)

[0909] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0910] In food delivery operations, it is difficult for drivers to set destinations and check routes safely and efficiently while driving. Manual resetting and checking while driving is dangerous, so there is a need for voice interaction.

[0911] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0912] In this invention, the server includes means for a user to input voice commands to set a destination and receive route guidance, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and generating route guidance information, means for transmitting the route guidance information as text data to a terminal, means for converting the text data into voice and providing it to the user, means for collecting and analyzing user question data, means for updating a route guidance algorithm based on the analysis results, means for a delivery driver to set a destination and check a route by voice while driving, means for providing the route confirmation results by voice in real time, and means for generating next destination and route information using document generation technology. This enables drivers to set a destination and check a route by voice alone while driving, improving the safety and efficiency of their work.

[0913] "Voice input" is a method by which a user communicates verbal instructions or questions to a system.

[0914] "Text data" is data that has been converted from voice input into text information and is used for analysis and processing.

[0915] A "server" is a computing device that receives transmitted text data, analyzes it, and generates a response.

[0916] "Route guidance information" is information that indicates the route and direction for the user to reach the destination.

[0917] "Device" means a device operated by a user that receives audio input and provides audio output.

[0918] A "voice recognition engine" is software that analyzes voice input and converts it into corresponding text data.

[0919] A "speech synthesis engine" is software that analyzes text data and generates corresponding speech.

[0920] "Question data" refers to information about confirmations or questions that users make to the system.

[0921] "Parsing" is the process of understanding received text data and deciphering its intent.

[0922] A "route guidance algorithm" is a calculation procedure for providing optimal route information to users.

[0923] A "delivery driver" is a driver whose job is to deliver goods to customers.

[0924] "Real-time" refers to near-instant processing or response.

[0925] "Document generation technology" is technology that generates appropriate text responses to user questions.

[0926] "Driving" refers to the state in which the delivery driver is operating the vehicle.

[0927] This invention is a system that allows food delivery drivers to easily set destinations and check routes by voice while driving. To implement the invention, the following specific hardware and software configurations are required.

[0928] Hardware

[0929] Smartphone (device)

[0930] microphone

[0931] speaker

[0932] software

[0933] Speech recognition engine: SpeechRecognition (Python library)

[0934] Speech synthesis engine: gTTS (Google Text-to-Speech)

[0935] API request tool: requests (Python library)

[0936] Audio output playback utility: mpg321

[0937] Program processing

[0938] 1. User Operation

[0939] The user (delivery driver) speaks instructions into their smartphone, such as "Where's the next delivery destination?" or "Where's the next right turn?"

[0940] 2. Voice Input and Speech Recognition

[0941] The device receives the user's voice input through a microphone. The speech recognition engine (SpeechRecognition) converts this voice input into text data. For example, the voice "Where is the next delivery?" is converted into text data "Where is the next delivery?"

[0942] 3. Sending to the server and analyzing

[0943] The converted text data is sent from the device to the server. The server receives this text data and analyzes it using a natural language processing engine. For example, from the question "Where is the next delivery?", it understands that the user wants to check information about the next delivery address.

[0944] 4. Generating Route Guidance Information

[0945] Based on the analysis results, the server references a map information database and generates information about the next delivery destination and route. The generated route guidance information is sent to the terminal as text data such as "The next delivery destination is ____."

[0946] 5. Speech synthesis and delivery

[0947] The device converts the received text data into speech using a speech synthesis engine (gTTS), and provides the user with a voice message such as, "The next delivery destination is ____." This allows the user to check the next delivery destination without looking at the screen while driving.

[0948] Specific examples

[0949] Specifically, when a delivery driver asks, "Where's the next delivery?", the question is converted by the device into text data, "Where's the next delivery?" The text data is sent to the server, which analyzes the question and confirms the next delivery address. The server generates a response such as "The next delivery address is: XXX" and sends it back to the device as text data. The response is then converted into speech on the device and provided to the user.

[0950] Example prompts for generative AI models

[0951] "When a user asks, 'Where is my next delivery?', please search for the next delivery destination in a map database and generate a sentence that provides appropriate route guidance."

[0952] This will create a system that allows food delivery drivers to safely and efficiently check their destination and set their route while driving.

[0953] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0954] Step 1:

[0955] Using the voice input function of the smartphone, the user asks aloud, "Where is the next delivery destination?" The input is voice data, and the user's question is captured by the smartphone's microphone.

[0956] Step 2:

[0957] The device uses a speech recognition engine (SpeechRecognition) to convert voice data into text data. It analyzes the voice data and generates corresponding text data. For example, a speech saying "Where is the next delivery?" is converted into text data "Where is the next delivery?". This text data is used in the next processing step.

[0958] Step 3:

[0959] The converted text data is sent from the device to the server. The device uses an API request tool (requests) to send this text data to the server. The input is text data, and the output is an HTTP request to the server.

[0960] Step 4:

[0961] The server analyzes the text data received from the device. It uses a natural language processing engine to analyze the received text data and understand the intent of the user's question. Based on the analysis results, it generates a query to search for information on the next delivery destination. This query is used to reference a map information database.

[0962] Step 5:

[0963] The server uses the generated query to refer to a map information database and obtain information about the next delivery destination. The information obtained from the database includes the delivery destination address and route information. The obtained information is saved in text format and passed to the next processing step. For example, a text such as "The next delivery destination is at address: XXX" is generated.

[0964] Step 6:

[0965] The server uses document generation technology to generate an appropriate response based on the acquired delivery destination information. It uses a generative AI model (text generation engine) to create an appropriate response to the question. The response is specific text data such as "The next delivery destination is address: XXX." This text data is then sent back to the terminal.

[0966] Step 7:

[0967] The device converts the received text data into speech using a speech synthesis engine (gTTS). The text data is analyzed and the corresponding voice data is generated. The generated voice data is played back through the smartphone's speaker. For example, the user can hear a voice saying, "The next delivery address is: XXX."

[0968] Step 8:

[0969] The user continues driving based on the next delivery destination information provided by voice. The user can confirm the voice response and follow the appropriate route to the next delivery destination. In this way, a system is realized that allows users to confirm destinations and set routes safely and efficiently while driving.

[0970] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0971] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and also recognizes and provides feedback on the user's emotional state. This system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[0972] User Actions

[0973] The user sets their destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[0974] Terminal handling

[0975] The device converts the user's voice input into text data using a speech recognition engine. For example, the device converts the speech "Our next destination is Tokyo Tower" into text "Our next destination is Tokyo Tower." The converted text data is designed to be sent to a server. The device then analyzes the user's voice using an emotion engine to determine the user's emotional state. For example, the device may determine that the user is "anxious" based on the tone and speed of their voice. The results of this emotion analysis are also sent to the server.

[0976] Server Processing

[0977] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm the instructions for the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[0978] The server also receives emotional data sent from the device. If the server determines that the user is in an unstable emotional state after analyzing the emotional data, it generates feedback to provide gentle and easy-to-understand route guidance information.

[0979] Feedback and Learning

[0980] The device uses a speech synthesis engine to convert the text data sent from the server into speech and provides it to the user. For example, it might generate a voice saying, "That's right." Additionally, the server analyzes the emotional data and generates feedback, which is then provided to the user via voice. For example, if the user is feeling anxious, it might add a reassuring message such as, "Don't worry, just turn right here."

[0981] The server analyzes the questions and sentiment data collected from all users and periodically updates the algorithm to improve the system. This information helps the navigation algorithm to further evolve and provide optimal feedback to users.

[0982] Specific examples

[0983] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server refers to a map information database, confirms that the next fork should be a right turn, and generates the appropriate response, "That's right." Furthermore, if the device's emotion engine detects "anxiety" from the user's tone of voice, this information is sent to the server, which then generates an additional feedback message, "Don't worry, turn right here." This text data is again sent to the device, where it is converted into speech and provided to the user.

[0984] This process allows users to ask questions by voice while driving and receive appropriate and reliable route guidance. The server also analyzes the question data and emotion data to improve the system's accuracy and user experience. In this way, the car navigation system can continuously evolve and provide higher quality route guidance.

[0985] The processing flow will be explained below.

[0986] Specific processing steps of the program

[0987] 1. User Actions

[0988] Step 1:

[0989] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[0990] 2. Terminal Processing

[0991] Step 2:

[0992] The device receives the user's voice input and captures the audio signal with the microphone.

[0993] Step 3:

[0994] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[0995] Step 4:

[0996] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[0997] Step 5:

[0998] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[0999] Step 6:

[1000] The device again receives the user's voice input, capturing the audio signal with the microphone.

[1001] Step 7:

[1002] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[1003] Step 8:

[1004] The device sends text data and emotion data for voice emotion analysis to the server as an HTTP request.

[1005] 3. Server Processing

[1006] Step 9:

[1007] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[1008] Step 10:

[1009] The server receives the emotion data, which indicates an emotional state such as "anxiety."

[1010] Step 11:

[1011] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[1012] Step 12:

[1013] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[1014] Step 13:

[1015] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[1016] Step 14:

[1017] The server analyzes the emotion data and generates an additional feedback message if the user is in an anxious state: "Don't worry, just turn right here."

[1018] Step 15:

[1019] The server generates text data and sends the feedback message to the terminal as an HTTP response.

[1020] 4. Device reprocessing and user feedback

[1021] Step 16:

[1022] The device receives the text data from the server, including the text data "That's right" and the feedback message "Don't worry, turn right here."

[1023] Step 17:

[1024] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[1025] Step 18:

[1026] The device converts the feedback message into speech, generating a speech message that says, "Don't worry, just turn right here."

[1027] Step 19:

[1028] The device plays generated speech over the speaker, telling the user, "That's right," and "Don't worry, just turn right here."

[1029] 5. Questionnaire data collection and analysis

[1030] Step 20:

[1031] The server stores the user's question data and emotion data. The question "Should I turn right at the next fork?", the answer "Yes, that's right", the emotion "Anxious", and the feedback message "Don't worry, just turn right here" are stored in the database.

[1032] Step 21:

[1033] The server periodically analyzes the stored question data and sentiment data to identify areas where many users have experienced problems.

[1034] Step 22:

[1035] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[1036] These steps will enable users to receive accurate directions and emotion-sensitive feedback in response to voice questions while driving, while the server will continuously analyze data and update algorithms to improve the accuracy of the overall system and user experience.

[1037] Example 2

[1038] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1039] Conventional car navigation systems recognize users' voice input and provide route guidance, but voice recognition alone cannot take into account the user's emotional state, which can lead to anxiety and stress. Even if the system provides accurate answers to users' questions, it lacks emotional feedback, making it difficult to improve the user experience. Furthermore, data analysis based on users' actual usage is required to regularly improve the system.

[1040] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1041] In this invention, the server includes means for converting voice input into text data, means for transmitting the text data and user emotion data to the server, means for analyzing the text data to generate route guidance information, means for generating feedback based on the route guidance information and emotion data, means for transmitting the route guidance information and feedback as text data to the terminal, means for converting the text data into speech and providing it to the user, means for collecting and analyzing user question data and emotion data, and means for updating the route guidance algorithm based on the analysis results. This enables safe and secure route guidance that takes the user's emotional state into consideration, and also allows the system to be continuously improved based on data based on actual usage.

[1042] "User" refers to the driver or user who uses the car navigation system and provides voice input.

[1043] "Voice input" refers to the act of a user providing information by voice to set a destination or ask a question.

[1044] "Text data" refers to character string information converted from voice input by a voice recognition engine.

[1045] "Emotional Data" refers to a user's emotional state as determined by an emotion engine analyzing the user's voice input.

[1046] "Server" refers to a computer system that receives text data and emotion data, analyzes them, and generates feedback.

[1047] "Speech recognition engine" refers to a software or hardware component that converts voice input into text data.

[1048] "Emotion Engine" refers to a software or hardware component that determines a user's emotional state from their voice input.

[1049] A "natural language processing engine" refers to a computational technology that analyzes text data to understand user intent and generate appropriate responses.

[1050] "Feedback" refers to additional reassurance messages or supplemental information provided to users based on emotional data.

[1051] "Map information database" refers to a digital database that stores geographic information and route guidance information.

[1052] "Speech synthesis engine" means a software or hardware component that converts text data into speech.

[1053] "Question Data" refers to information including questions and confirmations made by a User through voice input.

[1054] A "direction guidance algorithm" refers to a set of computational methods and procedures for providing optimal route guidance to a user.

[1055] "Analysis results" refers to conclusions and findings derived from analytical methods based on collected data.

[1056] "Update" refers to the process of incorporating new data and analysis results to improve route guidance algorithms.

[1057] MODE FOR CARRYING OUT THE INVENTION

[1058] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and recognizes the user's emotional state and provides feedback. This system is composed of three entities: the user, the terminal, and the server.

[1059] User Actions

[1060] The user sets the destination by voice into the car navigation system terminal. For example, the user can say, "My next destination is Tokyo Tower." Route guidance will start, and the user can ask questions about unclear route guidance. For example, the user can ask, "Should I turn right at the next fork?"

[1061] Terminal handling

[1062] The device receives the user's voice input and converts the voice into text data using a voice recognition engine (e.g., voice recognition software). The voice input "My next destination is Tokyo Tower" is converted into text data "My next destination is Tokyo Tower." This text data is sent to the server.

[1063] The device also analyzes the user's voice data using an emotion engine (e.g., emotion analysis software) to determine the user's emotional state. For example, if the user is determined to be "anxious" based on the tone and speed of their voice, the emotion analysis results are also sent to the server.

[1064] Server Processing

[1065] The server receives the text data sent from the device and analyzes it using a natural language processing engine (e.g., natural language processing software). The server understands the intent of the user's question, and deciphers, for example, from the question "Should I turn right at the next fork?", that the user wants to confirm whether or not to turn right at the next guidance point.

[1066] The server then refers to a map information database (e.g., geographic information software) to confirm the instructions for the next route guidance point. Based on the confirmation result, it generates text data such as "That's right." This text data is then sent to the terminal.

[1067] Furthermore, the server receives and analyzes the emotional data sent from the device. If the user is in an unstable emotional state, the server generates a feedback message that provides gentle and easy-to-understand guidance information. For example, it generates a reassuring message such as, "Don't worry, just turn right here."

[1068] Device feedback and audio output

[1069] The device converts the text data sent from the server into speech using a speech synthesis engine (e.g., speech synthesis software) and provides it to the user. For example, it may play back a response such as "That's right." In addition, feedback messages based on the user's emotional state may also be provided.

[1070] System training and algorithm updates

[1071] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. This process allows the navigation algorithm to evolve and provide appropriate feedback to users.

[1072] Specific examples

[1073] For example, if a user asks "Should I turn right at the next fork?" while driving, the question is converted into text data by the device and sent to the server. The server analyzes the question, confirms that the user should turn right at the next guidance point, and generates the answer "That's right." If the user's sentiment analysis results indicate "anxiety," the server generates an additional feedback message, "Don't worry, just turn right here." These messages are sent to the device and provided to the user as audio.

[1074] Prompt Sentence Examples

[1075] "Tell me whether to turn right at the next fork. If they seem unsure, add some reassuring feedback."

[1076] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1077] Step 1:

[1078] The user sets their destination by speaking into the device. The user's voice is used as input. For example, they might say, "My next destination is Tokyo Tower." This voice data is input to the device as output.

[1079] Step 2:

[1080] The device converts the user's voice input into text data using a speech recognition engine. Voice data is used as input. Specifically, for example, the Google Speech-to-Text API is used to convert the speech "My next destination is Tokyo Tower" into text "My next destination is Tokyo Tower." This text data is generated as output.

[1081] Step 3:

[1082] The terminal sends the converted text data to the server. The text data is used as input and sent to the server as output.

[1083] Step 4:

[1084] The device analyzes the user's voice data using an emotion analysis engine to determine the user's emotional state. The voice data is used as input. Specifically, for example, emotion analysis software may be used to determine the user's emotional state as "anxiety" based on the tone and rate of the user's voice. Emotion data is generated as output.

[1085] Step 5:

[1086] The terminal transmits emotion data to the server. The emotion data is used as input and transmitted to the server as output.

[1087] Step 6:

[1088] The server receives the text data sent from the device and analyzes it using a natural language processing engine. The text data is used as input. Specifically, the natural language processing software is used to decipher the user's intent from the question, "Should I turn right at the next fork?" The analysis results are generated as output.

[1089] Step 7:

[1090] The server consults a map information database to confirm the instructions for the next route point. The analysis results are used as input. Specifically, geographic information software is used to confirm the correct route for the next route point. The output is generated as "That's right."

[1091] Step 8:

[1092] The server receives the emotion data sent from the device and analyzes the user's emotional state. The emotion data is used as input. As output, a feedback message based on the emotional state is generated. Specifically, if anxiety is detected, the server generates the feedback "Don't worry, turn right here."

[1093] Step 9:

[1094] The server sends the generated route guidance information and feedback messages to the terminal. The route guidance information and feedback are used as inputs. As outputs, these information are sent to the terminal.

[1095] Step 10:

[1096] The device converts the text data sent from the server into speech using a speech synthesis engine. The text data of route guidance information and feedback messages is used as input. Specifically, for example, Amazon Polly is used to convert text such as "That's right" and "Don't worry, turn right here" into speech. The output is speech data.

[1097] Step 11:

[1098] The terminal provides the generated audio data to the user. The audio data is used as input. As output, audio guidance is presented to the user.

[1099] Step 12:

[1100] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. The collected data is used as input. An updated algorithm is generated as output, which improves the route guidance algorithm.

[1101] (Application example 2)

[1102] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1103] Conventional car navigation systems and industrial robot mobility systems lack feedback that takes into account the user's emotional state, which can lead to anxiety and stress. Furthermore, they lack the functionality to respond appropriately to user questions and concerns that arise during route guidance. Furthermore, when a user is in an emotionally unstable state, the system needs to be able to understand this and provide appropriate, reassuring route guidance.

[1104] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1105] In this invention, the server includes means for analyzing the emotional state of the user and adding feedback information to the route guidance information, means for outputting the feedback information by voice, and means for identifying the areas where many users have problems and the emotional states of the users by collecting and analyzing the users' question data and emotional states, and updating the route guidance algorithm, thereby enabling the user to receive appropriate route guidance with peace of mind.

[1106] A "user" is the entity that operates the system, sets destinations, and asks questions.

[1107] A "destination" is the final destination of a trip set by the user.

[1108] "Directions" refers to route information and instructions provided to the user by the system.

[1109] "Voice input means" refers to a device or software that accepts voice instructions from a user.

[1110] "Means for converting into text data" refers to technology or machinery for converting voice data into text format data.

[1111] "Means for sending to the server" refers to the system or protocol for sending data from the terminal to the server.

[1112] The "means for analyzing and generating route guidance information" refers to algorithms or software that analyzes the transmitted text data and generates route information and instructions.

[1113] The "means for transmitting to the terminal as text data" is a system for transmitting the analyzed route guidance information to the terminal as text data.

[1114] "Means of converting text data into audio and providing it to users" refers to technologies and devices that convert text data into audio data and communicate it to users.

[1115] "Means for collecting and analyzing question data" refers to technology for collecting questions asked by users and analyzing them.

[1116] A "means for updating the route guidance algorithm" is a method for improving and updating the system's route provision algorithm based on collected data.

[1117] The "means for analyzing the emotional state and adding feedback information" is a system for analyzing the user's emotions and reflecting corresponding information in the route guidance.

[1118] The "means for outputting feedback information by voice" refers to a technology or device for providing the user with feedback information in the form of voice in response to the analysis results.

[1119] This invention is a system for car navigation systems and factory navigation systems that uses voice input to provide route guidance and provide feedback according to the user's emotional state. It is mainly divided into three entities: the server, the terminal, and the user, each of which plays a specific role.

[1120] server

[1121] The server is responsible for analyzing the user's voice input as text data and generating appropriate route guidance information and feedback. The server has the following functions:

[1122] 1. Speech recognition engine: The server uses a speech recognition engine (e.g., the speech_recognition library) to convert the voice data sent from the terminal into text data.

[1123] 2. Sentiment Analysis Engine: The server uses a sentiment analysis engine (e.g., the BERT model from the transformers library) to determine the emotional state of the user's voice.

[1124] 3. Route guidance generation engine: The server uses a natural language processing engine to analyze the text data and understand the user's intent. It then refers to a map information database, confirms the instructions for the next route guidance point, and generates appropriate route guidance information.

[1125] 4. Feedback generation engine: Based on the results of sentiment analysis, it generates feedback (such as reassuring messages) that adapts to the user's emotional state.

[1126] Terminal

[1127] The terminal acts as an interface between the server and the user and has the following functions:

[1128] 1. Voice acquisition means: Collects the user's voice using a microphone.

[1129] 2. Voice conversion means: Converts voice input into text data and sends it to the server.

[1130] 3. Speech synthesis means: The text data sent from the server is converted into speech using a speech synthesis engine (e.g., the pyttsx3 library) and provided to the user.

[1131] 4. Emotion analysis data transmission means: Transmits the collected emotional state data to the server.

[1132] user

[1133] Users are the users of the system who use voice to navigate and ask questions. User roles include:

[1134] 1. Voice input: Set destinations and ask driving questions by voice.

[1135] 2. Feedback reception: Receives route guidance information and feedback provided by the server and the terminal.

[1136] Specific examples

[1137] For example, if a robot moving around a factory asks, "Where is the next route?", the question is converted into text data by the terminal and sent to the server. The server analyzes the question and understands what the user wants to know. It references a map information database to generate precise instructions such as "Next left turn," and also generates feedback based on emotion analysis, such as "Don't worry, there are 50 meters until the next left turn." This information is sent to the terminal, converted into voice by a speech synthesis engine, and provided to the robot.

[1138] Prompt Sentence Examples

[1139] Use the following example as a prompt to input to your generative AI model:

[1140] What they say: "This area is congested. What's the next safe route?"

[1141] Emotion detection: Anxiety

[1142] Generated feedback: "Next left turn. Don't worry, there are 50 meters until the next left turn."

[1143] Thus, a detailed description of an embodiment of the present invention has been provided, which provides a system that efficiently processes a user's voice input and provides appropriate feedback based on the user's emotional state.

[1144] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1145] Step 1:

[1146] The device uses a microphone as a means of acquiring voice input from the user. When the user gives voice instructions for directions or questions, the device captures this voice. For example, when the user says, "Where is the next route?", the voice is input into the microphone.

[1147] Step 2:

[1148] The device converts the acquired voice data into text data using a speech recognition engine (for example, the speech_recognition library). This text data becomes the string "Where is the next route?" The device creates this text data and prepares to send its status to the server.

[1149] Step 3:

[1150] The device then sends the text data converted by the voice recognition engine to the server. In this case, the text data "Where is the next route?" is sent to the server. At the same time, the device also analyzes the user's emotional state and sends this data to the server. For example, if "anxiety" is detected from the converted voice, that emotional data is also sent.

[1151] Step 4:

[1152] The server analyzes the received text data using a natural language processing engine to understand the user's intent. Specifically, it deciphers the text "Where is the next route?" to understand that the user needs directions. NLP libraries (e.g., transformers) are used for this analysis.

[1153] Step 5:

[1154] The server references the map information database based on the analysis results and generates the next route guidance information. For example, the server obtains information such as "Next left turn" from the map information database and generates it as text data. It also creates additional feedback information based on the received emotion data. For example, it generates feedback information such as "Don't worry, there are 50 meters until the next left turn."

[1155] Step 6:

[1156] The server sends the generated route guidance information and feedback information to the device. Specifically, text data such as "Next left turn" and "Don't worry, there are 50 meters until the next left turn" is sent to the device.

[1157] Step 7:

[1158] The device converts the received text data into speech data using a speech synthesis engine (for example, the pyttsx3 library). The device generates speech data such as "Next left turn. Don't worry, there are 50 meters until the next left turn."

[1159] Step 8:

[1160] The device provides the generated voice data to the user. Specifically, it outputs a voice message to the user through the speaker saying, "Next left turn. Don't worry, there are 50 meters until the next left turn." The user can then continue receiving route guidance with peace of mind after hearing this voice output.

[1161] Step 9:

[1162] The server analyzes the collected question data and emotion data and updates the system's algorithm. For example, if many users feel uneasy at a particular location, the system will provide more detailed route guidance information and stronger feedback. As a result, future users will receive more appropriate route guidance and feedback.

[1163] The above are the specific processing steps of the system that realizes this application example. This flow allows users to ask questions or get directions by voice input, and reach their destination while receiving reassuring feedback that reflects their emotional state.

[1164] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1165] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1166] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1167] [Fourth embodiment]

[1168] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1169] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1170] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1171] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1172] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1173] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1174] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1175] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1176] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1177] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1178] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1179] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1180] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1181] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[1182] User Actions

[1183] The user sets the destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[1184] Terminal handling

[1185] The device converts the user's voice input into text data using a speech recognition engine. For example, the speech "Should I turn right at the next fork?" is converted into text data such as "Should I turn right at the next fork?" The converted text data is designed to be sent to a server. The device then receives the text data from the server, converts it into speech using a speech synthesis engine, and provides it to the user.

[1186] Server Processing

[1187] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm whether or not to turn right at the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[1188] Feedback and Learning

[1189] The device converts the text data sent from the server into speech and provides the user with a voice response saying, "That's right." In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users have problems. Based on this information, the server updates its route guidance algorithm and provides even easier-to-understand route guidance.

[1190] Specific examples

[1191] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server references a map information database, confirms that the next fork will be a right turn, and generates the appropriate answer, "That's right." The generated answer is then sent back to the device as text data, converted into audio on the device, and provided to the user. The user can confirm that the route guidance is accurate by hearing the audio answer, "That's right."

[1192] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[1193] The processing flow will be explained below.

[1194] Specific processing steps of the program

[1195] 1. User Actions

[1196] Step 1:

[1197] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[1198] 2. Terminal Processing

[1199] Step 2:

[1200] The device receives the user's voice input and captures the audio signal with the microphone.

[1201] Step 3:

[1202] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[1203] Step 4:

[1204] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[1205] Step 5:

[1206] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[1207] Step 6:

[1208] The device again receives the user's voice input, capturing the audio signal with the microphone.

[1209] Step 7:

[1210] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[1211] Step 8:

[1212] The device sends text data to the server as an HTTP request.

[1213] 3. Server Processing

[1214] Step 9:

[1215] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[1216] Step 10:

[1217] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[1218] Step 11:

[1219] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[1220] Step 12:

[1221] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[1222] Step 13:

[1223] The server generates text data and sends it to the terminal as an HTTP response.

[1224] 4. Device reprocessing and user feedback

[1225] Step 14:

[1226] The device receives the text data from the server. The device receives the text data "That's right."

[1227] Step 15:

[1228] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[1229] Step 16:

[1230] The device generates a voice message and plays it over the speaker, telling the user "That's right."

[1231] 5. Questionnaire data collection and analysis

[1232] Step 17:

[1233] The server stores the user's question data. The question "Should I turn right at the next fork?" and the answer "Yes, that's right" are stored in the database.

[1234] Step 18:

[1235] The server periodically analyzes the stored question data to identify areas where many users have experienced problems.

[1236] Step 19:

[1237] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[1238] Through this processing step, users can ask questions by voice and receive appropriate directions, while the server analyzes the question data to improve the accuracy of the system.

[1239] Example 1

[1240] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1241] Conventional car navigation systems often provide inappropriate directions when receiving voice input and are unable to accurately understand the user's intent. Furthermore, users are unable to ask questions by voice if they are unable to understand the directions, which is inconvenient. Furthermore, the lack of a feedback function that allows many users to identify areas where they have problems and improve the navigation algorithm makes it difficult to improve the system's accuracy and user experience.

[1242] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1243] In this invention, the server includes means for converting text data received by the terminal from the server into speech and providing it to the user through an output device in the vehicle, means for analyzing the user's intentions using a natural language processing engine, and means for referencing a map information database to obtain information on the user's current location and the next intersection. This makes it possible to provide appropriate route guidance and answers to destination settings and questions entered through voice input, and to improve the accuracy of the entire system and the user experience through the feedback function.

[1244] 1. "User" means a driver or passenger who uses a car navigation system in a vehicle.

[1245] 2. "Voice input" refers to the act of the user speaking their destination or questions into the car navigation system through a microphone.

[1246] 3. "Text data" means data in the form of a string of characters converted from voice input by a voice recognition engine.

[1247] 4. "Terminal" means an electronic device, including input and output devices, installed in a vehicle in which a car navigation system is installed.

[1248] 5. "Server" means a remote computer system for receiving and analyzing text data.

[1249] 6. "Speech recognition engine" means a software or hardware mechanism that converts voice input into text data.

[1250] 7. A "natural language processing engine" is an artificial intelligence technology that analyzes text data and understands user intent.

[1251] 8. "Map information database" means a database containing road information and geographic information that the server references to generate route guidance information.

[1252] 9. "Route Guidance Information" means information regarding the direction and route generated by the server based on the user's destination and current location.

[1253] 10. "Speech synthesis engine" means a software or hardware mechanism that converts text data into speech.

[1254] 11. "Feedback" is the process of collecting and analyzing user query data to improve the system's navigation algorithms.

[1255] 12. A "guidance algorithm" is a set of procedures or calculation methods for providing users with appropriate directions and routes.

[1256] 13. "Output device" means a device, such as a speaker, that provides voice prompts or other information to the user in the vehicle.

[1257] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance. The system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[1258] First, the user sets their destination by voice into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins and the user continues driving. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[1259] The device converts the user's voice input into text data using a speech recognition engine (e.g., an ambiguous speech recognition engine or a third-party speech recognition engine). After the voice is converted into text data, the text data is sent to a server via the Internet. The server then analyzes the received text data using a natural language processing engine (e.g., a generative AI model or a third-party natural language processing engine) to understand the intent of the user's question.

[1260] If the analysis of the text data determines that the user wants to confirm whether they should turn right at the next intersection, the server references a map information database (for example, a general map information database API or a map information database API from another company). The server obtains route guidance information based on the user's current location and the information about the next intersection, and generates text data for the appropriate answer, such as "That's right." This text data is then sent back to the device.

[1261] The terminal converts the text data sent from the server into speech using a speech synthesis engine (for example, a rough speech synthesis engine or a speech synthesis engine made by another company) and provides the speech to the user through an output device (such as a speaker) in the car. In this way, the user can hear the voice response "That's right."

[1262] In addition, the server collects question data from all users and periodically analyzes it to identify areas where many users experience problems. This allows the route guidance algorithm to be updated and more user-friendly route guidance to be provided. Specifically, the collected data is analyzed using machine learning algorithms to identify areas for improvement in the system.

[1263] Specific examples

[1264] For example, if a user asks "Should I turn right at the next fork?" while driving towards their destination, the question is converted into text data "Should I turn right at the next fork?" using a voice recognition engine on the device. The text data is sent to a server via the Internet, and the server analyzes the question using a natural language processing engine to understand what the user wants to confirm. The server references a map information database, confirms that they should turn right at the next fork, and generates the appropriate answer "That's right." The generated answer is again sent to the device as text data, where it is converted into voice using a voice synthesis engine. The user can confirm that the route guidance is accurate by hearing the voice answer "That's right."

[1265] Examples of prompt statements

[1266] "Utterance: 'Should I turn right at the next fork?'"

[1267] Through this process of voice input, analysis, and feedback, users can receive detailed spoken directions while driving, greatly improving driving safety and convenience. Furthermore, the server analyzes the collected question data and updates the navigation algorithm, improving the accuracy of the entire system and the user experience. By repeating this process, the car navigation system continuously evolves, providing users with higher-quality directions.

[1268] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1269] Step 1: User Speech Input

[1270] The user sets the destination by speaking into the car navigation system terminal. The input is the user's voice, which can be a specific destination such as "My next destination is Tokyo Tower" or a question such as "Should I turn right at the next fork?" The voice is received through the terminal's microphone input device.

[1271] Step 2: Voice Recognition

[1272] The device converts the user's voice input into text data using a voice recognition engine (e.g., voice recognition software). The input is the user's voice signal, and the output is text data in the form of a string: "Should I turn right at the next fork?" The voice data is converted into a digital signal, and phonemic and linguistic analysis is performed before being output as text.

[1273] Step 3: Send text data to the server

[1274] The device sends the text data converted by the speech recognition engine to the server via the Internet. The input is the text data "Should I turn right at the next fork?" and the output is data in the form of network packets. The device's network module is responsible for this communication and sends the data to the server via the HTTPS protocol.

[1275] Step 4: Analyzing the text data

[1276] The server analyzes the received text data using a natural language processing engine (e.g., a generative AI model). The input is the text data "Should I turn right at the next fork?" and the output is structured data that represents the user's intent. The server performs tokenization and grammatical analysis on the text data to understand what the user wants to confirm.

[1277] Step 5: Query the map database

[1278] The server references a map information database (e.g., a map information API) based on the analysis results. The input is the user's current location and information about the next intersection, and the output is route guidance information including the action to be taken at the next intersection (e.g., turn right). The server sends a query to the map information database to obtain the required information.

[1279] Step 6: Generate directions

[1280] The server generates appropriate route guidance information based on information obtained from the map information database. The input is route guidance information obtained from the map database, and the output is the text-format route guidance information "That's right" to be provided to the user. The server composes this as text data in character string format.

[1281] Step 7: Sending text data from the server to the device

[1282] The generated text data is resent from the server to the terminal. The input is the generated text data "That's right," and the output is data in the form of network packets. The server's network module is responsible for this communication.

[1283] Step 8: Text-to-Speech

[1284] The device converts the text data received from the server into speech using a speech synthesis engine (e.g., text-to-speech software). The input is the text data "That's right," and the output is an audio file (e.g., audio data in WAV or MP3 format). The text data is input into the speech synthesis engine, and output as an audio file.

[1285] Step 9: User voice guidance

[1286] The terminal provides the generated voice data to the user through the car's speaker. The input is the voice file "That's right," and the output is the voice played through the speaker. The user listens to this and confirms the next action.

[1287] Step 10: Feedback and learning

[1288] The server collects question data from all users and periodically analyzes it. The input is the collected question data (e.g., text data history), and the output is feedback data containing improvements to the route guidance algorithm. A machine learning algorithm is used to analyze the data and identify improvements to the system. This allows the server to update the route guidance algorithm and improve the accuracy of the entire system.

[1289] (Application example 1)

[1290] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1291] In food delivery operations, it is difficult for drivers to set destinations and check routes safely and efficiently while driving. Manual resetting and checking while driving is dangerous, so there is a need for voice interaction.

[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1293] In this invention, the server includes means for a user to input voice commands to set a destination and receive route guidance, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and generating route guidance information, means for transmitting the route guidance information as text data to a terminal, means for converting the text data into voice and providing it to the user, means for collecting and analyzing user question data, means for updating a route guidance algorithm based on the analysis results, means for a delivery driver to set a destination and check a route by voice while driving, means for providing the route confirmation results by voice in real time, and means for generating next destination and route information using document generation technology. This enables drivers to set a destination and check a route by voice alone while driving, improving the safety and efficiency of their work.

[1294] "Voice input" is a method by which a user communicates verbal instructions or questions to a system.

[1295] "Text data" is data that has been converted from voice input into text information and is used for analysis and processing.

[1296] A "server" is a computing device that receives transmitted text data, analyzes it, and generates a response.

[1297] "Route guidance information" is information that indicates the route and direction for the user to reach the destination.

[1298] "Device" means a device operated by a user that receives audio input and provides audio output.

[1299] A "voice recognition engine" is software that analyzes voice input and converts it into corresponding text data.

[1300] A "speech synthesis engine" is software that analyzes text data and generates corresponding speech.

[1301] "Question data" refers to information about confirmations and questions that users send to the system.

[1302] "Analysis" is a process of understanding the received text data and decoding its intention.

[1303] "Route guidance algorithm" is a calculation procedure for providing users with optimal route information.

[1304] "Delivery driver" refers to a driver responsible for delivering goods to customers.

[1305] "Real-time" means that processing and responses are carried out almost immediately.

[1306] "Document generation technology" is a technology for generating appropriate text responses to users' questions.

[1307] "During driving" refers to the state when a delivery driver is operating a vehicle.

[1308] This invention is a system that enables a driver for food delivery to easily set a destination and confirm a route by voice during driving. To implement the invention, the following specific hardware and software configurations are required.

[1309] Hardware

[1310] Smartphone (terminal)

[1311] Microphone

[1312] Speaker

[1313] Software

[1314] Speech recognition engine: SpeechRecognition (Python library)

[1315] Speech synthesis engine: gTTS (Google Text-to-Speech)

[1316] API request tool: requests (Python library)

[1317] Audio output playback utility: mpg321

[1318] Program processing

[1319] 1. User Operation

[1320] The user (delivery driver) speaks instructions into their smartphone, such as "Where's the next delivery destination?" or "Where's the next right turn?"

[1321] 2. Voice Input and Speech Recognition

[1322] The device receives the user's voice input through a microphone. The speech recognition engine (SpeechRecognition) converts this voice input into text data. For example, the voice "Where is the next delivery?" is converted into text data "Where is the next delivery?"

[1323] 3. Sending to the server and analyzing

[1324] The converted text data is sent from the device to the server. The server receives this text data and analyzes it using a natural language processing engine. For example, from the question "Where is the next delivery?", it understands that the user wants to check information about the next delivery address.

[1325] 4. Generating Route Guidance Information

[1326] Based on the analysis results, the server references a map information database and generates information about the next delivery destination and route. The generated route guidance information is sent to the terminal as text data such as "The next delivery destination is ____."

[1327] 5. Speech synthesis and delivery

[1328] The device converts the received text data into speech using a speech synthesis engine (gTTS), and provides the user with a voice message such as, "The next delivery destination is ____." This allows the user to check the next delivery destination without looking at the screen while driving.

[1329] Specific examples

[1330] Specifically, when a delivery driver asks, "Where's the next delivery?", the question is converted by the device into text data, "Where's the next delivery?" The text data is sent to the server, which analyzes the question and confirms the next delivery address. The server generates a response such as "The next delivery address is: XXX" and sends it back to the device as text data. The response is then converted into speech on the device and provided to the user.

[1331] Example prompts for generative AI models

[1332] "When a user asks, 'Where is my next delivery?', please search for the next delivery destination in a map database and generate a sentence that provides appropriate route guidance."

[1333] This will create a system that allows food delivery drivers to safely and efficiently check their destination and set their route while driving.

[1334] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1335] Step 1:

[1336] Using the voice input function of the smartphone, the user asks aloud, "Where is the next delivery destination?" The input is voice data, and the user's question is captured by the smartphone's microphone.

[1337] Step 2:

[1338] The device uses a speech recognition engine (SpeechRecognition) to convert voice data into text data. It analyzes the voice data and generates corresponding text data. For example, a speech saying "Where is the next delivery?" is converted into text data "Where is the next delivery?". This text data is used in the next processing step.

[1339] Step 3:

[1340] The converted text data is sent from the device to the server. The device uses an API request tool (requests) to send this text data to the server. The input is text data, and the output is an HTTP request to the server.

[1341] Step 4:

[1342] The server analyzes the text data received from the device. It uses a natural language processing engine to analyze the received text data and understand the intent of the user's question. Based on the analysis results, it generates a query to search for information on the next delivery destination. This query is used to reference a map information database.

[1343] Step 5:

[1344] The server uses the generated query to refer to a map information database and obtain information about the next delivery destination. The information obtained from the database includes the delivery destination address and route information. The obtained information is saved in text format and passed to the next processing step. For example, a text such as "The next delivery destination is at address: XXX" is generated.

[1345] Step 6:

[1346] The server uses document generation technology to generate an appropriate response based on the acquired delivery destination information. It uses a generative AI model (text generation engine) to create an appropriate response to the question. The response is specific text data such as "The next delivery destination is address: XXX." This text data is then sent back to the terminal.

[1347] Step 7:

[1348] The device converts the received text data into speech using a speech synthesis engine (gTTS). The text data is analyzed and the corresponding voice data is generated. The generated voice data is played back through the smartphone's speaker. For example, the user can hear a voice saying, "The next delivery address is: XXX."

[1349] Step 8:

[1350] The user continues driving based on the next delivery destination information provided by voice. The user can confirm the voice response and follow the appropriate route to the next delivery destination. In this way, a system is realized that allows users to confirm destinations and set routes safely and efficiently while driving.

[1351] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1352] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and also recognizes and provides feedback on the user's emotional state. This system is composed of three entities: the user, the terminal, and the server, each of which plays a specific role.

[1353] User Actions

[1354] The user sets their destination by speaking into the car navigation system terminal. For example, they might say, "My next destination is Tokyo Tower." After the destination is set, route guidance begins. If the route guidance is unclear or if they need clarification, they can ask a question by voice. For example, they might say, "Should I turn right at the next fork?"

[1355] Terminal handling

[1356] The device converts the user's voice input into text data using a speech recognition engine. For example, the device converts the speech "Our next destination is Tokyo Tower" into text "Our next destination is Tokyo Tower." The converted text data is designed to be sent to a server. The device then analyzes the user's voice using an emotion engine to determine the user's emotional state. For example, the device may determine that the user is "anxious" based on the tone and speed of their voice. The results of this emotion analysis are also sent to the server.

[1357] Server Processing

[1358] The server receives the text data sent from the device. The received text data is analyzed by a natural language processing engine to understand the intent of the user's question. For example, from the question "Should I turn right at the next fork?", it is deciphered that the user wants to confirm whether or not to turn right at the next route guidance point. The server then references a map information database to confirm the instructions for the next route guidance point. Based on the confirmation results, it generates appropriate route guidance information and generates text data such as "That's right." This text data is then sent back to the device.

[1359] The server also receives emotional data sent from the device. If the server determines that the user is in an unstable emotional state after analyzing the emotional data, it generates feedback to provide gentle and easy-to-understand route guidance information.

[1360] Feedback and Learning

[1361] The device uses a speech synthesis engine to convert the text data sent from the server into speech and provides it to the user. For example, it might generate a voice saying, "That's right." Additionally, the server analyzes the emotional data and generates feedback, which is then provided to the user via voice. For example, if the user is feeling anxious, it might add a reassuring message such as, "Don't worry, just turn right here."

[1362] The server analyzes the questions and sentiment data collected from all users and periodically updates the algorithm to improve the system. This information helps the navigation algorithm to further evolve and provide optimal feedback to users.

[1363] Specific examples

[1364] For example, if a user asks "Should I turn right at the next fork?" while driving towards a destination, the question is converted by the device into text data, "Should I turn right at the next fork?" The text data is sent to a server, which analyzes the question and understands what the user wants to confirm. The server refers to a map information database, confirms that the next fork should be a right turn, and generates the appropriate response, "That's right." Furthermore, if the device's emotion engine detects "anxiety" from the user's tone of voice, this information is sent to the server, which then generates an additional feedback message, "Don't worry, turn right here." This text data is again sent to the device, where it is converted into speech and provided to the user.

[1365] This process allows users to ask questions by voice while driving and receive appropriate and reliable route guidance. The server also analyzes the question data and emotion data to improve the system's accuracy and user experience. In this way, the car navigation system can continuously evolve and provide higher quality route guidance.

[1366] The processing flow will be explained below.

[1367] Specific processing steps of the program

[1368] 1. User Actions

[1369] Step 1:

[1370] The user sets the destination by voice input, saying, "My next destination is Tokyo Tower."

[1371] 2. Terminal Processing

[1372] Step 2:

[1373] The device receives the user's voice input and captures the audio signal with the microphone.

[1374] Step 3:

[1375] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Our next destination is Tokyo Tower" is converted into text "Our next destination is Tokyo Tower."

[1376] Step 4:

[1377] The device sets the destination and starts navigation, calculates the route from the map information database, and starts providing directions.

[1378] Step 5:

[1379] If a user has a question while driving, they can ask it by voice, such as, "Should I turn right at the next fork?"

[1380] Step 6:

[1381] The device again receives the user's voice input, capturing the audio signal with the microphone.

[1382] Step 7:

[1383] The device uses a speech recognition engine to convert the voice signal into text data. For example, the voice saying "Should I turn right at the next fork?" is converted into text "Should I turn right at the next fork?"

[1384] Step 8:

[1385] The device sends text data and emotion data for voice emotion analysis to the server as an HTTP request.

[1386] 3. Server Processing

[1387] Step 9:

[1388] The server receives the text data from the device, such as "Should I turn right at the next fork?"

[1389] Step 10:

[1390] The server receives the emotion data, which indicates an emotional state such as "anxiety."

[1391] Step 11:

[1392] The server uses a natural language processing engine to analyze the text data and decipher the user's intent. For example, it analyzes the question, "Should I turn right at the next fork?"

[1393] Step 12:

[1394] The server refers to the map information database to check the instructions for the next route guidance point. If there is a right turn at the next fork, that information is acquired.

[1395] Step 13:

[1396] The server generates an appropriate answer based on the information it has obtained. It creates text data such as "That's right."

[1397] Step 14:

[1398] The server analyzes the emotion data and generates an additional feedback message if the user is in an anxious state: "Don't worry, just turn right here."

[1399] Step 15:

[1400] The server generates text data and sends the feedback message to the terminal as an HTTP response.

[1401] 4. Device reprocessing and user feedback

[1402] Step 16:

[1403] The device receives the text data from the server, including the text data "That's right" and the feedback message "Don't worry, turn right here."

[1404] Step 17:

[1405] The device uses a speech synthesis engine to convert the text data into speech, generating the voice "That's right."

[1406] Step 18:

[1407] The device converts the feedback message into speech, generating a speech message that says, "Don't worry, just turn right here."

[1408] Step 19:

[1409] The device plays generated speech over the speaker, telling the user, "That's right," and "Don't worry, just turn right here."

[1410] 5. Questionnaire data collection and analysis

[1411] Step 20:

[1412] The server stores the user's question data and emotion data. The question "Should I turn right at the next fork?", the answer "Yes, that's right", the emotion "Anxious", and the feedback message "Don't worry, just turn right here" are stored in the database.

[1413] Step 21:

[1414] The server periodically analyzes the stored question data and sentiment data to identify areas where many users have experienced problems.

[1415] Step 22:

[1416] The server updates the route guidance algorithm based on the analysis results, providing even easier route guidance in the future.

[1417] These steps will enable users to receive accurate directions and emotion-sensitive feedback in response to voice questions while driving, while the server will continuously analyze data and update algorithms to improve the accuracy of the overall system and user experience.

[1418] Example 2

[1419] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1420] Conventional car navigation systems recognize users' voice input and provide route guidance, but voice recognition alone cannot take into account the user's emotional state, which can lead to anxiety and stress. Even if the system provides accurate answers to users' questions, it lacks emotional feedback, making it difficult to improve the user experience. Furthermore, data analysis based on users' actual usage is required to regularly improve the system.

[1421] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1422] In this invention, the server includes means for converting voice input into text data, means for transmitting the text data and user emotion data to the server, means for analyzing the text data to generate route guidance information, means for generating feedback based on the route guidance information and emotion data, means for transmitting the route guidance information and feedback as text data to the terminal, means for converting the text data into speech and providing it to the user, means for collecting and analyzing user question data and emotion data, and means for updating the route guidance algorithm based on the analysis results. This enables safe and secure route guidance that takes the user's emotional state into consideration, and also allows the system to be continuously improved based on data based on actual usage.

[1423] "User" refers to the driver or user who uses the car navigation system and provides voice input.

[1424] "Voice input" refers to the act of a user providing information by voice to set a destination or ask a question.

[1425] "Text data" refers to character string information converted from voice input by a voice recognition engine.

[1426] "Emotional Data" refers to a user's emotional state as determined by an emotion engine analyzing the user's voice input.

[1427] "Server" refers to a computer system that receives text data and emotion data, analyzes them, and generates feedback.

[1428] "Speech recognition engine" refers to a software or hardware component that converts voice input into text data.

[1429] "Emotion Engine" refers to a software or hardware component that determines a user's emotional state from their voice input.

[1430] A "natural language processing engine" refers to a computational technology that analyzes text data to understand user intent and generate appropriate responses.

[1431] "Feedback" refers to additional reassurance messages or supplemental information provided to users based on emotional data.

[1432] "Map information database" refers to a digital database that stores geographic information and route guidance information.

[1433] "Speech synthesis engine" means a software or hardware component that converts text data into speech.

[1434] "Question Data" refers to information including questions and confirmations made by a User through voice input.

[1435] A "direction guidance algorithm" refers to a set of computational methods and procedures for providing optimal route guidance to a user.

[1436] "Analysis results" refers to conclusions and findings derived from analytical methods based on collected data.

[1437] "Update" refers to the process of incorporating new data and analysis results to improve route guidance algorithms.

[1438] MODE FOR CARRYING OUT THE INVENTION

[1439] This invention is a car navigation system that allows users to set their destination and ask questions by voice when receiving route guidance, and recognizes the user's emotional state and provides feedback. This system is composed of three entities: the user, the terminal, and the server.

[1440] User Actions

[1441] The user sets the destination by voice into the car navigation system terminal. For example, the user can say, "My next destination is Tokyo Tower." Route guidance will start, and the user can ask questions about unclear route guidance. For example, the user can ask, "Should I turn right at the next fork?"

[1442] Terminal handling

[1443] The device receives the user's voice input and converts the voice into text data using a voice recognition engine (e.g., voice recognition software). The voice input "My next destination is Tokyo Tower" is converted into text data "My next destination is Tokyo Tower." This text data is sent to the server.

[1444] The device also analyzes the user's voice data using an emotion engine (e.g., emotion analysis software) to determine the user's emotional state. For example, if the user is determined to be "anxious" based on the tone and speed of their voice, the emotion analysis results are also sent to the server.

[1445] Server Processing

[1446] The server receives the text data sent from the device and analyzes it using a natural language processing engine (e.g., natural language processing software). The server understands the intent of the user's question, and deciphers, for example, from the question "Should I turn right at the next fork?", that the user wants to confirm whether or not to turn right at the next guidance point.

[1447] The server then refers to a map information database (e.g., geographic information software) to confirm the instructions for the next route guidance point. Based on the confirmation result, it generates text data such as "That's right." This text data is then sent to the terminal.

[1448] Furthermore, the server receives and analyzes the emotional data sent from the device. If the user is in an unstable emotional state, the server generates a feedback message that provides gentle and easy-to-understand guidance information. For example, it generates a reassuring message such as, "Don't worry, just turn right here."

[1449] Device feedback and audio output

[1450] The device converts the text data sent from the server into speech using a speech synthesis engine (e.g., speech synthesis software) and provides it to the user. For example, it may play back a response such as "That's right." In addition, feedback messages based on the user's emotional state may also be provided.

[1451] System training and algorithm updates

[1452] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. This process allows the navigation algorithm to evolve and provide appropriate feedback to users.

[1453] Specific examples

[1454] For example, if a user asks "Should I turn right at the next fork?" while driving, the question is converted into text data by the device and sent to the server. The server analyzes the question, confirms that the user should turn right at the next guidance point, and generates the answer "That's right." If the user's sentiment analysis results indicate "anxiety," the server generates an additional feedback message, "Don't worry, just turn right here." These messages are sent to the device and provided to the user as audio.

[1455] Prompt Sentence Examples

[1456] "Tell me whether to turn right at the next fork. If they seem unsure, add some reassuring feedback."

[1457] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1458] Step 1:

[1459] The user sets their destination by speaking into the device. The user's voice is used as input. For example, they might say, "My next destination is Tokyo Tower." This voice data is input to the device as output.

[1460] Step 2:

[1461] The device converts the user's voice input into text data using a speech recognition engine. Voice data is used as input. Specifically, for example, the Google Speech-to-Text API is used to convert the speech "My next destination is Tokyo Tower" into text "My next destination is Tokyo Tower." This text data is generated as output.

[1462] Step 3:

[1463] The terminal sends the converted text data to the server. The text data is used as input and sent to the server as output.

[1464] Step 4:

[1465] The device analyzes the user's voice data using an emotion analysis engine to determine the user's emotional state. The voice data is used as input. Specifically, for example, emotion analysis software may be used to determine the user's emotional state as "anxiety" based on the tone and rate of the user's voice. Emotion data is generated as output.

[1466] Step 5:

[1467] The terminal transmits emotion data to the server. The emotion data is used as input and transmitted to the server as output.

[1468] Step 6:

[1469] The server receives the text data sent from the device and analyzes it using a natural language processing engine. The text data is used as input. Specifically, the natural language processing software is used to decipher the user's intent from the question, "Should I turn right at the next fork?" The analysis results are generated as output.

[1470] Step 7:

[1471] The server consults a map information database to confirm the instructions for the next route point. The analysis results are used as input. Specifically, geographic information software is used to confirm the correct route for the next route point. The output is generated as "That's right."

[1472] Step 8:

[1473] The server receives the emotion data sent from the device and analyzes the user's emotional state. The emotion data is used as input. As output, a feedback message based on the emotional state is generated. Specifically, if anxiety is detected, the server generates the feedback "Don't worry, turn right here."

[1474] Step 9:

[1475] The server sends the generated route guidance information and feedback messages to the terminal. The route guidance information and feedback are used as inputs. As outputs, these information are sent to the terminal.

[1476] Step 10:

[1477] The device converts the text data sent from the server into speech using a speech synthesis engine. The text data of route guidance information and feedback messages is used as input. Specifically, for example, Amazon Polly is used to convert text such as "That's right" and "Don't worry, turn right here" into speech. The output is speech data.

[1478] Step 11:

[1479] The terminal provides the generated audio data to the user. The audio data is used as input. As output, audio guidance is presented to the user.

[1480] Step 12:

[1481] The server analyzes the question data and sentiment data collected from all users and periodically updates the algorithm to improve the system. The collected data is used as input. An updated algorithm is generated as output, which improves the route guidance algorithm.

[1482] (Application example 2)

[1483] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1484] Conventional car navigation systems and industrial robot mobility systems lack feedback that takes into account the user's emotional state, which can lead to anxiety and stress. Furthermore, they lack the functionality to respond appropriately to user questions and concerns that arise during route guidance. Furthermore, when a user is in an emotionally unstable state, the system needs to be able to understand this and provide appropriate, reassuring route guidance.

[1485] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1486] In this invention, the server includes means for analyzing the emotional state of the user and adding feedback information to the route guidance information, means for outputting the feedback information by voice, and means for identifying the areas where many users have problems and the emotional states of the users by collecting and analyzing the users' question data and emotional states, and updating the route guidance algorithm, thereby enabling the user to receive appropriate route guidance with peace of mind.

[1487] A "user" is the entity that operates the system, sets destinations, and asks questions.

[1488] A "destination" is the final destination of a trip set by the user.

[1489] "Directions" refers to route information and instructions provided to the user by the system.

[1490] "Voice input means" refers to a device or software that accepts voice instructions from a user.

[1491] "Means for converting into text data" refers to technology or machinery for converting voice data into text format data.

[1492] "Means for sending to the server" refers to the system or protocol for sending data from the terminal to the server.

[1493] The "means for analyzing and generating route guidance information" refers to algorithms or software that analyzes the transmitted text data and generates route information and instructions.

[1494] The "means for transmitting to the terminal as text data" is a system for transmitting the analyzed route guidance information to the terminal as text data.

[1495] "Means of converting text data into audio and providing it to users" refers to technologies and devices that convert text data into audio data and communicate it to users.

[1496] "Means for collecting and analyzing question data" refers to technology for collecting questions asked by users and analyzing them.

[1497] A "means for updating the route guidance algorithm" is a method for improving and updating the system's route provision algorithm based on collected data.

[1498] The "means for analyzing the emotional state and adding feedback information" is a system for analyzing the user's emotions and reflecting corresponding information in the route guidance.

[1499] The "means for outputting feedback information by voice" refers to a technology or device for providing the user with feedback information in the form of voice in response to the analysis results.

[1500] This invention is a system for car navigation systems and factory navigation systems that uses voice input to provide route guidance and provide feedback according to the user's emotional state. It is mainly divided into three entities: the server, the terminal, and the user, each of which plays a specific role.

[1501] server

[1502] The server is responsible for analyzing the user's voice input as text data and generating appropriate route guidance information and feedback. The server has the following functions:

[1503] 1. Speech recognition engine: The server uses a speech recognition engine (e.g., the speech_recognition library) to convert the voice data sent from the terminal into text data.

[1504] 2. Sentiment Analysis Engine: The server uses a sentiment analysis engine (e.g., the BERT model from the transformers library) to determine the emotional state of the user's voice.

[1505] 3. Route guidance generation engine: The server uses a natural language processing engine to analyze the text data and understand the user's intent. It then refers to a map information database, confirms the instructions for the next route guidance point, and generates appropriate route guidance information.

[1506] 4. Feedback generation engine: Based on the results of sentiment analysis, it generates feedback (such as reassuring messages) that adapts to the user's emotional state.

[1507] Terminal

[1508] The terminal acts as an interface between the server and the user and has the following functions:

[1509] 1. Voice acquisition means: Collects the user's voice using a microphone.

[1510] 2. Voice conversion means: Converts voice input into text data and sends it to the server.

[1511] 3. Speech synthesis means: The text data sent from the server is converted into speech using a speech synthesis engine (e.g., the pyttsx3 library) and provided to the user.

[1512] 4. Emotion analysis data transmission means: Transmits the collected emotional state data to the server.

[1513] user

[1514] Users are the users of the system who use voice to navigate and ask questions. User roles include:

[1515] 1. Voice input: Set destinations and ask driving questions by voice.

[1516] 2. Feedback reception: Receives route guidance information and feedback provided by the server and the terminal.

[1517] Specific examples

[1518] For example, if a robot moving around a factory asks, "Where is the next route?", the question is converted into text data by the terminal and sent to the server. The server analyzes the question and understands what the user wants to know. It references a map information database to generate precise instructions such as "Next left turn," and also generates feedback based on emotion analysis, such as "Don't worry, there are 50 meters until the next left turn." This information is sent to the terminal, converted into voice by a speech synthesis engine, and provided to the robot.

[1519] Prompt Sentence Examples

[1520] Use the following example as a prompt to input to your generative AI model:

[1521] What they say: "This area is congested. What's the next safe route?"

[1522] Emotion detection: Anxiety

[1523] Generated feedback: "Next left turn. Don't worry, there are 50 meters until the next left turn."

[1524] Thus, a detailed description of an embodiment of the present invention has been provided, which provides a system that efficiently processes a user's voice input and provides appropriate feedback based on the user's emotional state.

[1525] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1526] Step 1:

[1527] The device uses a microphone as a means of acquiring voice input from the user. When the user gives voice instructions for directions or questions, the device captures this voice. For example, when the user says, "Where is the next route?", the voice is input into the microphone.

[1528] Step 2:

[1529] The device converts the acquired voice data into text data using a speech recognition engine (for example, the speech_recognition library). This text data becomes the string "Where is the next route?" The device creates this text data and prepares to send its status to the server.

[1530] Step 3:

[1531] The device then sends the text data converted by the voice recognition engine to the server. In this case, the text data "Where is the next route?" is sent to the server. At the same time, the device also analyzes the user's emotional state and sends this data to the server. For example, if "anxiety" is detected from the converted voice, that emotional data is also sent.

[1532] Step 4:

[1533] The server analyzes the received text data using a natural language processing engine to understand the user's intent. Specifically, it deciphers the text "Where is the next route?" to understand that the user needs directions. NLP libraries (e.g., transformers) are used for this analysis.

[1534] Step 5:

[1535] The server references the map information database based on the analysis results and generates the next route guidance information. For example, the server obtains information such as "Next left turn" from the map information database and generates it as text data. It also creates additional feedback information based on the received emotion data. For example, it generates feedback information such as "Don't worry, there are 50 meters until the next left turn."

[1536] Step 6:

[1537] The server sends the generated route guidance information and feedback information to the device. Specifically, text data such as "Next left turn" and "Don't worry, there are 50 meters until the next left turn" is sent to the device.

[1538] Step 7:

[1539] The device converts the received text data into speech data using a speech synthesis engine (for example, the pyttsx3 library). The device generates speech data such as "Next left turn. Don't worry, there are 50 meters until the next left turn."

[1540] Step 8:

[1541] The device provides the generated voice data to the user. Specifically, it outputs a voice message to the user through the speaker saying, "Next left turn. Don't worry, there are 50 meters until the next left turn." The user can then continue receiving route guidance with peace of mind after hearing this voice output.

[1542] Step 9:

[1543] The server analyzes the collected question data and emotion data and updates the system's algorithm. For example, if many users feel uneasy at a particular location, the system will provide more detailed route guidance information and stronger feedback. As a result, future users will receive more appropriate route guidance and feedback.

[1544] The above are the specific processing steps of the system that realizes this application example. This flow allows users to ask questions or get directions by voice input, and reach their destination while receiving reassuring feedback that reflects their emotional state.

[1545] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1546] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1547] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1548] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1549] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1550] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1551] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1552] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1553] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1554] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1555] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1556] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1557] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1558] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1559] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1560] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1561] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1562] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1563] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1564] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1565] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1566] The following is further disclosed regarding the above embodiment.

[1567] (Claim 1)

[1568] a means for a user to enter voice input to set a destination and receive directions;

[1569] means for converting the voice input into text data;

[1570] means for transmitting the text data to a server;

[1571] means for analyzing the text data and generating route guidance information;

[1572] means for transmitting the route guidance information to a terminal as text data;

[1573] means for converting the text data into speech and providing the speech to a user;

[1574] A means for collecting and analyzing user question data;

[1575] means for updating a route guidance algorithm based on the analysis results;

[1576] A system including:

[1577] (Claim 2)

[1578] 10. The system of claim 1, wherein the system receives the user's voice input and converts it into text data using a speech recognition engine.

[1579] (Claim 3)

[1580] The system according to claim 1, wherein the system collects and analyzes the user question data to identify areas where many users have experienced problems and updates the route guidance algorithm.

[1581] "Example 1"

[1582] (Claim 1)

[1583] a means for a user to enter voice input to set a destination and receive directions;

[1584] means for converting the voice input into text data;

[1585] means for transmitting the text data to a server;

[1586] means for analyzing the text data and generating route guidance information;

[1587] means for transmitting the route guidance information to a terminal as text data;

[1588] means for converting the text data into speech and providing the speech to a user;

[1589] A means for collecting and analyzing user question data;

[1590] means for updating a route guidance algorithm based on the analysis results;

[1591] a means for converting the text data received by the terminal from the server into speech and providing the speech to the user through an output device in the vehicle;

[1592] A means of analyzing user intent using a natural language processing engine;

[1593] A means for referencing a map information database to obtain information on the user's current location and the next intersection;

[1594] A system including:

[1595] (Claim 2)

[1596] 10. The system of claim 1, wherein the system receives a user's voice input and converts it into text data using a speech recognition engine.

[1597] (Claim 3)

[1598] The system according to claim 1, wherein the system collects and analyzes user question data to identify areas where many users have experienced problems and updates the route guidance algorithm.

[1599] "Application Example 1"

[1600] (Claim 1)

[1601] a means for a user to enter voice input to set a destination and receive directions;

[1602] means for converting the voice input into text data;

[1603] means for transmitting the text data to a server;

[1604] means for analyzing the text data and generating route guidance information;

[1605] means for transmitting the route guidance information to a terminal as text data;

[1606] means for converting the text data into speech and providing the speech to a user;

[1607] A means for collecting and analyzing user question data;

[1608] means for updating a route guidance algorithm based on the analysis results;

[1609] A means for delivery drivers to set destinations and check routes by voice while driving,

[1610] means for providing the route confirmation result by voice in real time;

[1611] A means for generating next destination and route information using document generation technology;

[1612] A system including:

[1613] (Claim 2)

[1614] 10. The system of claim 1, wherein the system receives the user's voice input and converts it into text data using a speech recognition engine.

[1615] (Claim 3)

[1616] The system according to claim 1, wherein the system collects and analyzes the user question data to identify areas where many users have experienced problems and updates the route guidance algorithm.

[1617] "Example 2: Combining Emotion Engines"

[1618] (Claim 1)

[1619] a means for a user to enter voice input to set a destination and receive directions;

[1620] means for converting the voice input into text data;

[1621] means for transmitting the text data and the user's emotion data to a server;

[1622] means for analyzing the text data and generating route guidance information;

[1623] means for generating feedback based on the navigation information and the emotion data;

[1624] means for transmitting the route guidance information and feedback to a terminal as text data;

[1625] means for converting the text data into speech and providing the speech to a user;

[1626] A means for collecting and analyzing user question data and emotion data;

[1627] means for updating a route guidance algorithm based on the analysis results;

[1628] A system including:

[1629] (Claim 2)

[1630] 10. The system of claim 1, wherein the system receives the user's voice input and converts it into text data using a speech recognition engine.

[1631] (Claim 3)

[1632] The system according to claim 1, wherein the system collects and analyzes the user question data and emotion data to identify areas where many users have experienced problems and updates the route guidance algorithm.

[1633] "Application example 2 when combining emotion engines"

[1634] (Claim 1)

[1635] a means for a user to enter voice input to set a destination and receive directions;

[1636] means for converting the voice input into text data;

[1637] means for transmitting the text data to a server;

[1638] means for analyzing the text data and generating route guidance information;

[1639] means for transmitting the route guidance information to a terminal as text data;

[1640] means for converting the text data into speech and providing the speech to a user;

[1641] A means for collecting and analyzing user question data;

[1642] means for updating a route guidance algorithm based on the analysis results;

[1643] means for analyzing the emotional state of the user and adding feedback information to the navigation information;

[1644] means for outputting the feedback information by voice;

[1645] A system including:

[1646] (Claim 2)

[1647] 10. The system of claim 1, wherein the system receives the user's voice input and converts it into text data using a speech recognition engine.

[1648] (Claim 3)

[1649] The system of claim 1, wherein the system collects and analyzes the user's question data and emotional state to identify areas where many users have experienced problems and their emotional state, and updates the route guidance algorithm. [Explanation of symbols]

[1650] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to enter voice input to set a destination and receive directions; means for converting the voice input into text data; means for transmitting the text data to a server; means for analyzing the text data and generating route guidance information; means for transmitting the route guidance information to a terminal as text data; means for converting the text data into speech and providing the speech to a user; A means for collecting and analyzing user question data; means for updating a route guidance algorithm based on the analysis results; A system including:

2. The system of claim 1 , wherein the system receives the user's voice input and converts it into text data using a speech recognition engine.

3. The system according to claim 1, wherein the system collects and analyzes the user question data to identify areas where many users have experienced problems and updates the route guidance algorithm.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A