system

The multilingual real-time navigation system addresses language barriers and real-time traffic challenges by converting voice input to text, translating, and using augmented reality for intuitive guidance, ensuring smooth and stress-free travel.

JP2026068325APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Tourists and visitors in foreign countries face challenges in understanding local languages and obtaining accurate real-time traffic information, leading to stress and difficulty in navigation.

Method used

A multilingual real-time navigation system that uses speech recognition to convert voice input into text, translates it into multiple languages, incorporates real-time traffic information, and provides visually enhanced guidance using augmented reality technology.

Benefits of technology

Enables tourists to reach their destinations efficiently and stress-free by providing intuitive navigation information tailored to their language and emotional state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068325000001_ABST
    Figure 2026068325000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Equipped with a speech recognition engine, and a means for converting the user's voice input into text information, A means for identifying a destination based on the aforementioned textual information and creating guidance information by translating it into multiple languages, A means for presenting the aforementioned guidance information to the user via a visual display device, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When a visitor searches for a destination in a foreign country, there are problems such as not being able to understand Japanese and not being able to obtain accurate real-time traffic information, which causes stress and difficulties during movement. In particular, there is a need for a navigation system suitable for tourists who speak various languages, and a method for solving this is required.

Means for Solving the Problems

[0005] This invention solves the above problems by using a speech recognition engine to convert the user's voice input into text information, analyzing it to identify the destination, and providing guidance information translated into multiple languages. Furthermore, it supports efficient and stress-free travel by using a visual display device to present route guidance to the user in a visually easy-to-understand manner, and by incorporating the optimal travel route using real-time traffic information. By utilizing augmented reality technology, it is possible to visually emphasize the guidance information and provide it in a way that is easy for the user to intuitively understand.

[0006] A "speech recognition engine" is a program or device that has the function of analyzing speech input and converting it into a string of characters.

[0007] A "user" is an individual who attempts to obtain navigation information using digital signage.

[0008] "Textual information" refers to text data generated through voice input and is used for analyzing navigation information.

[0009] A "destination" refers to a specific place that the user wishes to visit or travel to.

[0010] Translation is the process of converting written information from one language into another language and conveying its meaning.

[0011] "Guidance information" refers to data that includes navigation and destination-related information, and is intended to assist users in their travels.

[0012] A "visual display device" is a display device attached to digital signage, and is a means of presenting visual information.

[0013] "Real-time traffic information" refers to the latest data on current traffic conditions and is used to help choose a travel route.

[0014] Augmented reality technology is a technology that enhances the real world by overlaying computer graphics onto the real environment. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiment for Implementing the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a multilingual real-time navigation system designed to enable tourists and visitors to smoothly reach their destinations in foreign lands. Specifically, a terminal installed as part of digital signage accepts voice input from users.

[0037] Users can use their device's voice input function to communicate their destination in their native language. For example, they can give instructions such as, "Tell me how to get to the art museum."

[0038] The terminal uses a speech recognition engine to convert the user's voice into text. The converted text is then sent directly to the server.

[0039] The server identifies the destination from the received text information and performs the corresponding multilingual translation process. This translation utilizes generative AI to generate multilingual guidance information. Furthermore, real-time traffic information is incorporated to calculate the optimal route.

[0040] Navigation information generated by the server is transmitted to the terminal in a visually easy-to-understand format and provided to the user through a visual display on the terminal.

[0041] For example, if a user says "I want to go to Central Park" in front of a digital signage screen, the speech recognition engine converts the speech into text and sends it to the server. The server recognizes "Central Park" as the destination and generates multilingual directions. These directions include the best transportation options and recommended routes based on current traffic conditions.

[0042] Users can follow the displayed information and travel to their destination without stress. If augmented reality technology is used in this process, the guidance information is visually overlaid, allowing users to understand it intuitively.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user speaks their destination into the voice input device on the digital signage. At this time, the user can naturally specify the destination in their native language.

[0046] Step 2:

[0047] The device receives voice input and uses its built-in speech recognition engine to convert the voice data into text. The converted text contains information about the destination specified by the user.

[0048] Step 3:

[0049] The terminal constructs the converted character information as a data packet and sends it to the server. This packet contains the user's voice input recorded as a string of characters.

[0050] Step 4:

[0051] The server analyzes the received text information to identify the destination. It then compares it with a tourism database to retrieve detailed information related to the destination.

[0052] Step 5:

[0053] The server uses a generative AI model to translate the analyzed destination information into multiple languages. Based on the translation results, it constructs navigation information and makes it available in a displayable format.

[0054] Step 6:

[0055] The server acquires real-time traffic information and calculates the optimal travel route. This calculation takes into account traffic congestion and the operating status of public transportation.

[0056] Step 7:

[0057] The server sends the generated navigation information to the terminal. This data includes visual map information and translated text.

[0058] Step 8:

[0059] The terminal displays the received information on a visual display device. The displayed information may also include visual effects utilizing augmented reality technology.

[0060] Step 9:

[0061] The user begins their journey based on the displayed navigation information. Additional voice instructions may be given to help them understand the route to their destination.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] A challenge exists in that users who speak different languages ​​often find it difficult to reach their destinations accurately and smoothly in unfamiliar places. Therefore, there is a need for a system that provides real-time multilingual support and optimal routes. Furthermore, there is a need for a method of presenting guidance information in an intuitively understandable format.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes a voice conversion means that converts the user's voice input into text information, a means that analyzes the text information to identify the destination and uses a generative AI model to translate it into multiple languages ​​and create guidance information, and a means that acquires real-time information and calculates the optimal travel route based on the guidance information. This makes it possible for users from different language regions to reach their destination intuitively and quickly.

[0067] "Voice conversion means" refers to a technology or device for converting a user's voice input into text information.

[0068] A "generative AI model" is a generative artificial intelligence model intended for natural language processing and translation, and it performs information conversion between different languages.

[0069] "Real-time information" refers to information that is acquired and updated immediately in accordance with the current situation.

[0070] "Guidance information" refers to information that provides users with directions for travel, including routes to their destination and related information.

[0071] "Visual display means" refers to a device or technology for visually presenting guidance information to users.

[0072] Augmented reality technology is a technology that overlays digital information onto visual information from the real world.

[0073] A "prompt message" is a text message used to convey instructions or requests to a generative AI model.

[0074] This invention is a multilingual real-time navigation system for users who speak different languages ​​to reach their destination. This system provides guidance information to the user by combining voice conversion means, a generative AI model, multilingual translation, real-time information acquisition, and visual display means.

[0075] The user voice-inputs their destination in their native language into a digital display terminal. For example, they might say, "Please tell me how to get to the museum." This voice input is converted into text by the speech recognition engine built into the terminal. The speech recognition engine used here includes various speech recognition technologies as general software.

[0076] The terminal sends data to the server based on the converted text information. The server uses a generative AI model to analyze the received text information. The generative AI model incorporates technologies that enable diverse language conversion. The server uses this model to create multilingual guidance information. This information includes the optimal travel route to the destination and current traffic conditions.

[0077] The server utilizes a traffic information acquisition system to analyze traffic conditions using real-time data and calculate the optimal travel route. In this case, the route is adjusted to optimize it according to the traffic conditions.

[0078] The generated navigation information is transmitted to the terminal in a visually easy-to-understand format and displayed to the user by the terminal. The displayed content includes digital maps and directional icons so that users can understand it immediately. This allows users to reach their destination smoothly.

[0079] As a concrete example, consider a user prompt: "I want to go to Central Park." Based on this input, the server uses a generative AI model to create directions and provide the user with the most appropriate navigation.

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The user voice-inputs their destination in their native language into a digital display terminal. The input voice is received by a "voice recognition engine," which then serves as initial data for processing in the next step.

[0083] Step 2:

[0084] The device activates its speech recognition engine and converts the received audio into text information in real time. This conversion process analyzes the audio signal and converts the spoken content into text format. The resulting text information is specific text data, such as "Please tell me how to get to the art museum." The main data calculation here is the conversion of the audio signal into a string of characters.

[0085] Step 3:

[0086] The terminal sends the converted character information to the server. The character information, as input, is transferred to the server using a secure communication protocol. The output becomes data for destination identification and translation in the next step.

[0087] Step 4:

[0088] The server analyzes the received text information. In this step, it uses text analysis technology to identify the destination based on the input text information. Subsequently, a generative AI model is used to translate the information into multiple languages. The output is the translated guidance information data. The data processing here involves text analysis and language conversion.

[0089] Step 5:

[0090] The server calls a traffic information acquisition system to obtain real-time traffic information. Using destination information as input, it calculates the optimal route considering traffic conditions. As a result of this step, navigation guidance incorporating traffic information is generated. The output is optimal travel route data.

[0091] Step 6:

[0092] The server sends the generated multilingual navigation information to the terminal. This transmission process provides information including the selected optimal route and directions to the destination. The data output is provided in a format suitable for visual display.

[0093] Step 7:

[0094] The terminal displays navigation information to the user using visual display means. Specifically, maps, transportation information, walking routes, etc., are displayed on the terminal's screen. The input is navigation data received from the server, and the output is visual information that the user can refer to in real time.

[0095] (Application Example 1)

[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] Modern travelers and visitors face challenges in navigating foreign lands, including the need for multilingual support and real-time traffic information. This challenge stems from the need for fast and accurate navigation to ensure travelers with diverse language and cultural backgrounds reach their destinations smoothly. Furthermore, there is a lack of effective methods for integrating real-world and digital information and presenting it visually and clearly.

[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0099] In this invention, the server includes means for converting the user's voice input into text information, means for identifying the destination and generating guidance information by translating it into multiple languages, and means for visually overlaying and presenting the guidance information using AR technology. As a result, even in a foreign country, users can receive appropriate guidance in an intuitively understandable form and reach their destination quickly and efficiently.

[0100] "Speech recognition means" refers to a technology that converts a user's voice into text information, and is a means that analyzes voice input as a string of characters.

[0101] "Means for identifying a destination" refers to methods for identifying a destination from information entered by the user and deriving information about related locations.

[0102] "Methods for translating into multiple languages" refer to technologies that convert information expressed in one language into different languages, making it understandable to users who speak those languages.

[0103] "Means for generating guidance information" refers to technologies for creating detailed route guidance and traffic information related to a destination, and is a means of constructing information to be provided to users.

[0104] "AR technology" refers to augmented reality technology, which is a technology that overlays digital information onto the real world environment.

[0105] "Visual overlay presentation methods" refer to technologies that integrate and display digital information within the user's field of vision in a way that is intuitively understandable.

[0106] The system for carrying out this invention consists of a voice recognition means, a destination identification means, a multi-language translation means, a guidance information generation means, and a visual presentation means that utilizes AR technology.

[0107] The user enters their destination by voice via their smart device. The device is equipped with a speech recognition engine and uses the Google® Speech-to-Text API to convert the voice data into text. The converted text is then sent to the server.

[0108] The server analyzes the received text information to identify the destination. The Google Translate API is used to generate multilingual guidance information for the destination using AI generation. This makes it possible to provide guidance in different languages ​​selected by the user. Furthermore, the optimal travel route is calculated while considering real-time traffic information.

[0109] Unity and Vuforia are used to leverage augmented reality technology as a visual presentation tool. The generated navigation information is presented to the user through smart glasses or other visual devices. This allows users to intuitively receive navigation information in a form overlaid on the real world.

[0110] As a concrete example, if a user uses smart glasses and says, "I want to go to the museum," the voice is converted to text and sent to a server. The server generates the optimal route and multilingual navigation to the destination "museum" and displays it on the smart glasses using augmented reality. This allows the user to intuitively understand the route.

[0111] Examples of prompts for a generative AI model:

[0112] "Destination: Museum. Please generate multilingual navigation directions for the best route from my current location."

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] The user uses the microphone function of their smart device to input destination information by voice. The input voice data is received by the device's voice recognition engine.

[0116] Step 2:

[0117] The device uses the Google Speech-to-Text API to convert audio data into text. This data processing involves real-time analysis of the audio waveform to generate text, which in turn generates a string of characters. This string is then output and sent to the server.

[0118] Step 3:

[0119] The server analyzes the received string information to identify the destination. Natural language processing is used for the string data, extracting the destination and recognizing it as output. The recognized destination information is then passed on to the next step.

[0120] Step 4:

[0121] The server acquires real-time traffic information based on destination and current location information and calculates the optimal route. The output is generated as optimal route information. In this process, the best route is derived through data calculations performed by referencing an external traffic information database.

[0122] Step 5:

[0123] The server uses the Google Translate API to access a generation AI model and create multilingual navigation guidance information. It sends prompt text to the generation AI model, which then generates multilingual guidance information about the destination as output.

[0124] Step 6:

[0125] The device uses Unity and Vuforia to display the obtained multilingual guidance information as a visual presentation on smart glasses. In practice, AR technology is used to overlay digital information onto the real-world scenery, providing the user with intuitive navigation.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention is a multilingual real-time navigation system designed to make travel more comfortable for visitors in foreign countries, and it includes a function to recognize user emotions. The system's main components are a terminal embedded in digital signage and a server connected to it.

[0128] The user speaks their destination into the voice input on the digital signage. This voice input is converted into text by a speech recognition engine. The voice often contains hints indicating the user's emotions, and this information is also analyzed by an emotion engine.

[0129] The device converts speech to text while simultaneously using an emotion engine to analyze the user's emotions. This analysis determines the user's emotional state, such as whether they are relaxed or stressed. This allows for the incorporation of emotion-responsive feedback into navigation information.

[0130] The server analyzes text and emotional information received from the terminal to identify the destination. It then performs multilingual translation and generates navigation information. This generated navigation information incorporates real-time traffic data, suggesting the optimal travel route based on the user's emotional state.

[0131] For example, if a user enters "I want to get to the next tourist spot quickly" in a somewhat tense voice, the emotion engine will detect the stress level, and the server will generate guidance information that prioritizes the fastest route. On the other hand, if a user says in a relaxed voice, "Tell me a route with nice scenery," it is possible to incorporate scenic routes into the guidance information to provide a relaxing sightseeing experience.

[0132] The terminal displays information received from the server to the user on a visual display device. The displayed guidance information is visually enhanced using augmented reality technology, making it easier for the user to intuitively understand the navigation information.

[0133] Thus, the present invention aims to reduce the stress of travel and enrich the travel experience of visitors by providing a navigation system with emotion recognition capabilities.

[0134] The following describes the processing flow.

[0135] Step 1:

[0136] The user speaks their desired destination in their native language into the voice input device on the digital signage. In this process, the user's voice may naturally convey emotions.

[0137] Step 2:

[0138] The device receives voice input and converts the voice data into text using a speech recognition engine. This converted text contains information about the destination.

[0139] Step 3:

[0140] The device uses an emotion engine to analyze the emotional characteristics during voice input. The analysis identifies the user's emotional state, such as whether they are relaxed or tense.

[0141] Step 4:

[0142] The terminal packages the processed text information and emotional state data into a packet and sends it to the server. This packet contains both destination information and emotional data.

[0143] Step 5:

[0144] The server identifies the destination based on the received text information and compares it with a tourism database. Furthermore, it uses a generative AI model to translate destination-related information into multiple languages.

[0145] Step 6:

[0146] The server customizes navigation information based on the user's emotional state. For example, if the user is stressed, it prioritizes the fastest route; if they are relaxed, it suggests a route that includes tourist attractions.

[0147] Step 7:

[0148] The server calculates the optimal travel route, taking real-time traffic information into account. This calculation takes into account current travel conditions such as traffic congestion and public transport.

[0149] Step 8:

[0150] The server sends the generated, customized guidance information to the terminal. The data includes detailed maps and translated guidance text.

[0151] Step 9:

[0152] The terminal displays the received guidance information to the user on a visual display device. Using augmented reality technology, the guidance information is visually overlaid, allowing the user to visually confirm the guidance on the spot.

[0153] Step 10:

[0154] Based on the provided guidance information, users can comfortably begin their journey. The provision of emotionally considerate services allows them to enjoy a stress-free travel experience.

[0155] (Example 2)

[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0157] When traveling in a foreign country, there is a need to provide personalized navigation information that responds to the user's emotional state while addressing language barriers and real-time, ever-changing traffic conditions. Furthermore, there is a lack of easily understandable visual methods, making it a challenge to provide a less stressful travel experience.

[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0159] In this invention, the server includes means for recording the user's voice input and converting the voice into text information using a voice recognition mechanism; means including a mechanism for analyzing emotional information from the voice; and means for identifying a destination based on the text information and emotional information, translating it into multiple languages, and creating guidance information that corresponds to the user's emotional state. This makes it possible to provide users with intuitive, easy-to-understand, and emotionally sensitive real-time navigation information.

[0160] A "speech recognition mechanism" is a technology that converts a user's voice input into text data, and is a means of changing speech into a format that can be processed by a computer.

[0161] "Emotional information" is data that represents the user's psychological state by analyzing the emotional nuances extracted from the user's voice.

[0162] "Multilingual translation" refers to the process of converting information expressed in one language into multiple other languages, and is a technology that enables communication between people who speak different languages.

[0163] "Visual display devices" are devices used to visually present digital information to users, and include output devices such as screens and displays.

[0164] "Real-time travel information" refers to the latest travel data based on current traffic and road conditions, and is information used to respond to changing travel environments.

[0165] Augmented reality technology is a technology that overlays digital information onto the real world environment, enabling users to experience digital information in a realistic context.

[0166] "User emotional state" refers to the user's psychological and emotional condition and is an important indicator for personalizing navigation information.

[0167] This invention is a multilingual real-time navigation system that smoothly supports travel in foreign lands and is equipped with emotion recognition capabilities. The system mainly consists of a terminal embedded in digital signage and a server connected to it.

[0168] The user speaks their destination to the voice input device on the digital signage. The voice input is converted into text in real time by a speech recognition mechanism utilizing Google Cloud Speech-to-Text API and other technologies. In addition, data indicating the user's emotional state is analyzed from the voice using Azure Cognitive Services and other technologies. This allows for the determination of the user's psychological state, such as tension or relaxation.

[0169] The device sends this analyzed text data and sentiment information to the server. SSL / TLS protocol is used for communication to ensure data security. Upon receiving this data, the server uses the Alpaca API to extract keywords for destination information and the DeepL API to perform multilingual translation. Furthermore, by obtaining real-time travel status information via the Google Maps API, it generates the optimal travel route. This combination of data creates guidance information tailored to the user's emotional state. For example, if the user is in a hurry, the fastest route can be suggested; if they prefer a relaxing route, a scenic route can be proposed.

[0170] The terminal receives guidance information from the server and displays it visually on digital signage using augmented reality technology. This allows users to intuitively receive and easily understand the information.

[0171] This system aims to reduce stress and provide a richer travel experience by overcoming language barriers when traveling in foreign countries and offering optimal navigation tailored to the user's emotions. Examples of prompts generated using a generative AI model are as follows:

[0172] "When a user says they want to quickly move on to the next tourist spot, and the emotion engine detects a stressed state, please generate the optimal route to provide."

[0173] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0174] Step 1:

[0175] The user voice-inputs their destination into the microphone on the digital signage. The input data is the user's voice, which is then transmitted to the terminal. The terminal activates a speech recognition mechanism and, utilizing the Google Cloud Speech-to-Text API and other tools, converts the voice into text data in real time. As a result, the text information extracted from the voice input is output.

[0176] Step 2:

[0177] The device inputs the text information generated by speech recognition into the emotion engine. The emotion engine uses Azure Cognitive Services and other tools to analyze the emotional nuances contained in the speech. This process outputs emotional information such as whether the user is relaxed or tense.

[0178] Step 3:

[0179] The terminal sends text information and sentiment information together to the server. The input data consists of stringified destination information and sentiment data. The SSL / TLS protocol is used to securely transmit the data to the server. The output is the destination information and sentiment information received on the server side.

[0180] Step 4:

[0181] The server analyzes the received destination information and extracts relevant keywords using the Alpaca API. Based on this input data, it outputs processed key data. Next, it translates the destination information into multiple languages ​​using the DeepL API and outputs the translated data.

[0182] Step 5:

[0183] The server uses the Google Maps API to obtain real-time travel information. This input includes the user's destination information. Based on the acquired travel data, the server generates the optimal travel route, taking into account the user's emotional state. The output is personalized guidance information.

[0184] Step 6:

[0185] The terminal displays navigation information received from the server on a digital signage screen. The input is navigation information from the server, and the output is a navigation display enhanced using augmented reality technology. This allows the user to intuitively understand directions to their destination and travel comfortably.

[0186] (Application Example 2)

[0187] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0188] When traveling in a foreign country, there is a need to alleviate the language barrier and emotional stress that users face and provide a comfortable travel experience. Conventional navigation systems present routes in a single way without considering the user's emotional state, making it difficult to meet the diverse needs of users.

[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0190] In this invention, the server includes a speech recognition module and means for converting the user's voice input into text information, means for identifying a destination based on the text information and determining the user's emotional state using emotion analysis technology, and means for creating guidance information by translating it into multiple languages ​​and incorporating feedback according to the user's emotional state. As a result, even in a foreign country, the user will be guided to the optimal travel route according to their emotional state, enabling a stress-free travel experience without feeling the language barrier.

[0191] A "speech recognition module" is a device that receives voice input from a user, analyzes it, and converts it into text information.

[0192] "Feedback tailored to the user's emotional state" refers to a function that analyzes the user's emotions and provides appropriate information and guidance based on that state.

[0193] "Emotional analysis technology" is a technology that analyzes the tone of a user's voice and the content of their speech to identify their emotions.

[0194] "Translating into multiple languages" is the process of converting one language into several other languages, providing information in a way that is understandable to users whose native languages ​​are different.

[0195] To realize this invention, it is necessary to build a system that effectively combines a speech recognition module, emotion analysis technology, a multilingual translation engine, and augmented reality technology. The server receives the user's voice input, analyzes it to identify the destination and determine the emotion. The voice input is converted into text information using a speech recognition engine (e.g., Google Speech-to-Text). This text information is then used to determine the user's emotional state through an emotion analysis engine. The emotion analysis takes into account the tone and content of the voice.

[0196] The device uses a multilingual translation engine to provide guidance information that can be understood in the user's native language. The generated guidance information includes the optimal travel route, taking real-time traffic information into account. This allows users to travel to their destination comfortably and efficiently without experiencing language barriers.

[0197] Augmented reality technology is employed for visual presentation. The terminal visually highlights and presents guidance information, allowing users to intuitively understand the information. This technology makes even complex information easy to grasp.

[0198] As a concrete example, if a native Japanese speaker uses voice input to say, "I want to get to the next tourist spot quickly" while at a tourist destination abroad, the server will analyze the slightly tense tone of voice and suggest the fastest route. Example prompt: "Translate 'How do I get to the Eiffel Tower?' into French, including sentiment analysis."

[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0200] Step 1:

[0201] The user provides voice input, and the device receives the audio. The input is the user's voice data, which the device sends to a speech recognition engine (e.g., Google Speech-to-Text) to be converted into text. This converted text becomes the output of the next process.

[0202] Step 2:

[0203] The server receives text information and uses an emotion analysis engine to determine the user's emotional state. The input is text information, and emotion analysis technology is used to analyze the user's tone of voice and word choices to determine whether the user is stressed or relaxed. The output is the analyzed emotional state.

[0204] Step 3:

[0205] The server analyzes text information to identify the destination and extract location information. The input is text information, and natural language processing is used to identify the place name or facility name of the destination. The output is the location information of the destination.

[0206] Step 4:

[0207] The server uses a multilingual translation engine to translate the analysis results into a language the user can understand. The input consists of the guidance information and the user's native language, and the translation is performed using a language model. The output is the translated guidance information.

[0208] Step 5:

[0209] The server acquires real-time traffic information and calculates the optimal travel route based on the identified destination and the user's emotional state. The input consists of real-time traffic data and emotional analysis results. A route planning algorithm is used to identify the best route for the user. The output is the optimal travel route.

[0210] Step 6:

[0211] The device uses augmented reality technology to visually present guidance information to the user. Input consists of translated guidance information and the optimal travel route, presented in a visually enhanced format through the display device. Output is an intuitive visual display of guidance information for the user.

[0212] This entire process allows users to obtain a comfortable mode of transportation that suits their needs, without being hindered by language barriers.

[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0216] [Second Embodiment]

[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0229] This invention is a multilingual real-time navigation system designed to enable tourists and visitors to smoothly reach their destinations in foreign lands. Specifically, a terminal installed as part of digital signage accepts voice input from users.

[0230] Users can use their device's voice input function to communicate their destination in their native language. For example, they can give instructions such as, "Tell me how to get to the art museum."

[0231] The terminal uses a speech recognition engine to convert the user's voice into text. The converted text is then sent directly to the server.

[0232] The server identifies the destination from the received text information and performs the corresponding multilingual translation process. This translation utilizes generative AI to generate multilingual guidance information. Furthermore, real-time traffic information is incorporated to calculate the optimal route.

[0233] Navigation information generated by the server is transmitted to the terminal in a visually easy-to-understand format and provided to the user through a visual display on the terminal.

[0234] For example, if a user says "I want to go to Central Park" in front of a digital signage screen, the speech recognition engine converts the speech into text and sends it to the server. The server recognizes "Central Park" as the destination and generates multilingual directions. These directions include the best transportation options and recommended routes based on current traffic conditions.

[0235] Users can follow the displayed information and travel to their destination without stress. If augmented reality technology is used in this process, the guidance information is visually overlaid, allowing users to understand it intuitively.

[0236] The following describes the processing flow.

[0237] Step 1:

[0238] The user speaks their destination into the voice input device on the digital signage. At this time, the user can naturally specify the destination in their native language.

[0239] Step 2:

[0240] The device receives voice input and uses its built-in speech recognition engine to convert the voice data into text. The converted text contains information about the destination specified by the user.

[0241] Step 3:

[0242] The terminal constructs the converted character information as a data packet and sends it to the server. This packet contains the user's voice input recorded as a string of characters.

[0243] Step 4:

[0244] The server analyzes the received text information to identify the destination. It then compares it with a tourism database to retrieve detailed information related to the destination.

[0245] Step 5:

[0246] The server uses a generative AI model to translate the analyzed destination information into multiple languages. Based on the translation results, it constructs navigation information and makes it available in a displayable format.

[0247] Step 6:

[0248] The server acquires real-time traffic information and calculates the optimal travel route. This calculation takes into account traffic congestion and the operating status of public transportation.

[0249] Step 7:

[0250] The server sends the generated navigation information to the terminal. This data includes visual map information and translated text.

[0251] Step 8:

[0252] The terminal displays the received information on a visual display device. The displayed information may also include visual effects utilizing augmented reality technology.

[0253] Step 9:

[0254] The user begins their journey based on the displayed navigation information. Additional voice instructions may be given to help them understand the route to their destination.

[0255] (Example 1)

[0256] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0257] A challenge exists in that users who speak different languages ​​often find it difficult to reach their destinations accurately and smoothly in unfamiliar places. Therefore, there is a need for a system that provides real-time multilingual support and optimal routes. Furthermore, there is a need for a method of presenting guidance information in an intuitively understandable format.

[0258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0259] In this invention, the server includes a voice conversion means that converts the user's voice input into text information, a means that analyzes the text information to identify the destination and uses a generative AI model to translate it into multiple languages ​​and create guidance information, and a means that acquires real-time information and calculates the optimal travel route based on the guidance information. This makes it possible for users from different language regions to reach their destination intuitively and quickly.

[0260] "Voice conversion means" refers to a technology or device for converting a user's voice input into text information.

[0261] A "generative AI model" is a generative artificial intelligence model intended for natural language processing and translation, and it performs information conversion between different languages.

[0262] "Real-time information" refers to information that is acquired and updated immediately in accordance with the current situation.

[0263] "Guidance information" refers to information that provides users with directions for travel, including routes to their destination and related information.

[0264] "Visual display means" refers to a device or technology for visually presenting guidance information to users.

[0265] Augmented reality technology is a technology that overlays digital information onto visual information from the real world.

[0266] A "prompt message" is a text message used to convey instructions or requests to a generative AI model.

[0267] This invention is a multilingual real-time navigation system for users who speak different languages ​​to reach their destination. This system provides guidance information to the user by combining voice conversion means, a generative AI model, multilingual translation, real-time information acquisition, and visual display means.

[0268] The user voice-inputs their destination in their native language into a digital display terminal. For example, they might say, "Please tell me how to get to the museum." This voice input is converted into text by the speech recognition engine built into the terminal. The speech recognition engine used here includes various speech recognition technologies as general software.

[0269] The terminal sends data to the server based on the converted text information. The server uses a generative AI model to analyze the received text information. The generative AI model incorporates technologies that enable diverse language conversion. The server uses this model to create multilingual guidance information. This information includes the optimal travel route to the destination and current traffic conditions.

[0270] The server utilizes a traffic information acquisition system to analyze traffic conditions using real-time data and calculate the optimal travel route. In this case, the route is adjusted to optimize it according to the traffic conditions.

[0271] The generated navigation information is transmitted to the terminal in a visually easy-to-understand format and displayed to the user by the terminal. The displayed content includes digital maps and directional icons so that users can understand it immediately. This allows users to reach their destination smoothly.

[0272] As a concrete example, consider a user prompt: "I want to go to Central Park." Based on this input, the server uses a generative AI model to create directions and provide the user with the most appropriate navigation.

[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0274] Step 1:

[0275] The user voice-inputs their destination in their native language into a digital display terminal. The input voice is received by a "voice recognition engine," which then serves as initial data for processing in the next step.

[0276] Step 2:

[0277] The device activates its speech recognition engine and converts the received audio into text information in real time. This conversion process analyzes the audio signal and converts the spoken content into text format. The resulting text information is specific text data, such as "Please tell me how to get to the art museum." The main data calculation here is the conversion of the audio signal into a string of characters.

[0278] Step 3:

[0279] The terminal sends the converted character information to the server. The character information, as input, is transferred to the server using a secure communication protocol. The output becomes data for destination identification and translation in the next step.

[0280] Step 4:

[0281] The server analyzes the received character information. Based on the character information input in this step, text analysis technology is used to identify the destination. Subsequently, the process of translating into multiple languages proceeds using the generative AI model. The output obtained is the data of the translated guidance information. The data processing here is text analysis and language conversion.

[0282] Step 5:

[0283] The server calls the traffic information acquisition system to obtain real-time traffic information. Using the destination information as input, it calculates the optimal route considering the traffic situation. As a result of this step, a navigation guidance incorporating traffic information is generated. The output is the optimal movement route data.

[0284] Step 6:

[0285] The server transmits the generated multilingual navigation information to the terminal. In this transmission process, information including the selected optimal route and guidance to the destination is provided. The data output is provided in a format suitable for visual display.

[0286] Step 7:

[0287] The terminal uses visual display means to display navigation information to the user. Specifically, a map, transportation information, walking routes, etc. are displayed on the terminal's display. The input is the navigation data received from the server, and the output is visual information that the user can refer to in real time.

[0288] (Application Example 1)

[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0290] Modern travelers and visitors face challenges in navigating foreign lands, including the need for multilingual support and real-time traffic information. This challenge stems from the need for fast and accurate navigation to ensure travelers with diverse language and cultural backgrounds reach their destinations smoothly. Furthermore, there is a lack of effective methods for integrating real-world and digital information and presenting it visually and clearly.

[0291] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0292] In this invention, the server includes means for converting the user's voice input into text information, means for identifying the destination and generating guidance information by translating it into multiple languages, and means for visually overlaying and presenting the guidance information using AR technology. As a result, even in a foreign country, users can receive appropriate guidance in an intuitively understandable form and reach their destination quickly and efficiently.

[0293] "Speech recognition means" refers to a technology that converts a user's voice into text information, and is a means that analyzes voice input as a string of characters.

[0294] "Means for identifying a destination" refers to methods for identifying a destination from information entered by the user and deriving information about related locations.

[0295] "Methods for translating into multiple languages" refer to technologies that convert information expressed in one language into different languages, making it understandable to users who speak those languages.

[0296] "Means for generating guidance information" refers to technologies for creating detailed route guidance and traffic information related to a destination, and is a means of constructing information to be provided to users.

[0297] "AR technology" refers to augmented reality technology, which is a technology that overlays digital information onto the real world environment.

[0298] "Visual overlay presentation methods" refer to technologies that integrate and display digital information within the user's field of vision in a way that is intuitively understandable.

[0299] The system for carrying out this invention consists of a voice recognition means, a destination identification means, a multi-language translation means, a guidance information generation means, and a visual presentation means that utilizes AR technology.

[0300] The user enters their destination by voice via their smart device. The device is equipped with a speech recognition engine and uses the Google Speech-to-Text API to convert the voice data into text. The converted text is then sent to the server.

[0301] The server analyzes the received text information to identify the destination. The Google Translate API is used to generate multilingual guidance information for the destination using AI generation. This makes it possible to provide guidance in different languages ​​selected by the user. Furthermore, the optimal travel route is calculated while considering real-time traffic information.

[0302] Unity and Vuforia are used to leverage augmented reality technology as a visual presentation tool. The generated navigation information is presented to the user through smart glasses or other visual devices. This allows users to intuitively receive navigation information in a form overlaid on the real world.

[0303] As a concrete example, if a user uses smart glasses and says, "I want to go to the museum," the voice is converted to text and sent to a server. The server generates the optimal route and multilingual navigation to the destination "museum" and displays it on the smart glasses using augmented reality. This allows the user to intuitively understand the route.

[0304] Examples of prompts for a generative AI model:

[0305] "Destination: Museum. Please generate multilingual navigation guidance for the best route from the current location."

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] The user uses the microphone function of the smart device to input destination information by voice. The input voice data is received by the voice recognition engine of the terminal.

[0309] Step 2:

[0310] The terminal converts the voice data into character information using the Google Speech-to-Text API. The data processing performed at this stage is text conversion by real-time analysis of the voice waveform, and as a result, string information is generated. This string information is output and sent to the server.

[0311] Step 3:

[0312] The server analyzes the received string information to identify the destination. Here, natural language processing for the string is used as data calculation, and the destination is extracted and recognized as output. The recognized destination information is passed on to the next step.

[0313] Step 4:

[0314] The server obtains real-time traffic information based on the destination information and the current location information, and calculates the optimal route. The output is generated as optimal route information. In this process, an external traffic information database is referenced, and the best route is derived by the data calculation performed.

[0315] Step 5:

[0316] The server uses the Google Translate API to access a generation AI model and create multilingual navigation guidance information. It sends prompt text to the generation AI model, which then generates multilingual guidance information about the destination as output.

[0317] Step 6:

[0318] The device uses Unity and Vuforia to display the obtained multilingual guidance information as a visual presentation on smart glasses. In practice, AR technology is used to overlay digital information onto the real-world scenery, providing the user with intuitive navigation.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention is a multilingual real-time navigation system designed to make travel more comfortable for visitors in foreign countries, and it includes a function to recognize user emotions. The system's main components are a terminal embedded in digital signage and a server connected to it.

[0321] The user speaks their destination into the voice input on the digital signage. This voice input is converted into text by a speech recognition engine. The voice often contains hints indicating the user's emotions, and this information is also analyzed by an emotion engine.

[0322] The device converts speech to text while simultaneously using an emotion engine to analyze the user's emotions. This analysis determines the user's emotional state, such as whether they are relaxed or stressed. This allows for the incorporation of emotion-responsive feedback into navigation information.

[0323] The server analyzes text and emotional information received from the terminal to identify the destination. It then performs multilingual translation and generates navigation information. This generated navigation information incorporates real-time traffic data, suggesting the optimal travel route based on the user's emotional state.

[0324] For example, if a user enters "I want to get to the next tourist spot quickly" in a somewhat tense voice, the emotion engine will detect the stress level, and the server will generate guidance information that prioritizes the fastest route. On the other hand, if a user says in a relaxed voice, "Tell me a route with nice scenery," it is possible to incorporate scenic routes into the guidance information to provide a relaxing sightseeing experience.

[0325] The terminal displays information received from the server to the user on a visual display device. The displayed guidance information is visually enhanced using augmented reality technology, making it easier for the user to intuitively understand the navigation information.

[0326] Thus, the present invention aims to reduce the stress of travel and enrich the travel experience of visitors by providing a navigation system with emotion recognition capabilities.

[0327] The following describes the processing flow.

[0328] Step 1:

[0329] The user speaks their desired destination in their native language into the voice input device on the digital signage. In this process, the user's voice may naturally convey emotions.

[0330] Step 2:

[0331] The device receives voice input and converts the voice data into text using a speech recognition engine. This converted text contains information about the destination.

[0332] Step 3:

[0333] The device uses an emotion engine to analyze the emotional characteristics during voice input. The analysis identifies the user's emotional state, such as whether they are relaxed or tense.

[0334] Step 4:

[0335] The terminal packages the processed text information and emotional state data into a packet and sends it to the server. This packet contains both destination information and emotional data.

[0336] Step 5:

[0337] The server identifies the destination based on the received text information and compares it with a tourism database. Furthermore, it uses a generative AI model to translate destination-related information into multiple languages.

[0338] Step 6:

[0339] The server customizes navigation information based on the user's emotional state. For example, if the user is stressed, it prioritizes the fastest route; if they are relaxed, it suggests a route that includes tourist attractions.

[0340] Step 7:

[0341] The server calculates the optimal travel route, taking real-time traffic information into account. This calculation takes into account current travel conditions such as traffic congestion and public transport.

[0342] Step 8:

[0343] The server sends the generated, customized guidance information to the terminal. The data includes detailed maps and translated guidance text.

[0344] Step 9:

[0345] The terminal displays the received guidance information to the user on a visual display device. Using augmented reality technology, the guidance information is visually overlaid, allowing the user to visually confirm the guidance on the spot.

[0346] Step 10:

[0347] Based on the provided guidance information, users can comfortably begin their journey. The provision of emotionally considerate services allows them to enjoy a stress-free travel experience.

[0348] (Example 2)

[0349] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0350] When traveling in a foreign country, there is a need to provide personalized navigation information that responds to the user's emotional state while addressing language barriers and real-time, ever-changing traffic conditions. Furthermore, there is a lack of easily understandable visual methods, making it a challenge to provide a less stressful travel experience.

[0351] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0352] In this invention, the server includes means for recording the user's voice input and converting the voice into text information using a voice recognition mechanism; means including a mechanism for analyzing emotional information from the voice; and means for identifying a destination based on the text information and emotional information, translating it into multiple languages, and creating guidance information that corresponds to the user's emotional state. This makes it possible to provide users with intuitive, easy-to-understand, and emotionally sensitive real-time navigation information.

[0353] A "speech recognition mechanism" is a technology that converts a user's voice input into text data, and is a means of changing speech into a format that can be processed by a computer.

[0354] "Emotional information" is data that represents the user's psychological state by analyzing the emotional nuances extracted from the user's voice.

[0355] "Multilingual translation" refers to the process of converting information expressed in one language into multiple other languages, and is a technology that enables communication between people who speak different languages.

[0356] "Visual display devices" are devices used to visually present digital information to users, and include output devices such as screens and displays.

[0357] "Real-time travel information" refers to the latest travel data based on current traffic and road conditions, and is information used to respond to changing travel environments.

[0358] Augmented reality technology is a technology that overlays digital information onto the real world environment, enabling users to experience digital information in a realistic context.

[0359] "User emotional state" refers to the user's psychological and emotional condition and is an important indicator for personalizing navigation information.

[0360] This invention is a multilingual real-time navigation system that smoothly supports travel in foreign lands and is equipped with emotion recognition capabilities. The system mainly consists of a terminal embedded in digital signage and a server connected to it.

[0361] The user speaks their destination to the voice input device on the digital signage. The voice input is converted into text in real time by a speech recognition mechanism utilizing the Google Cloud Speech-to-Text API and other tools. In addition, data indicating the user's emotional state is analyzed from the voice using Azure Cognitive Services and other tools. This allows the system to determine the user's psychological state, such as tension or relaxation.

[0362] The device sends this analyzed text data and sentiment information to the server. SSL / TLS protocol is used for communication to ensure data security. Upon receiving this data, the server uses the Alpaca API to extract keywords for destination information and the DeepL API to perform multilingual translation. Furthermore, by obtaining real-time travel status information via the Google Maps API, it generates the optimal travel route. This combination of data creates guidance information tailored to the user's emotional state. For example, if the user is in a hurry, the fastest route can be suggested; if they prefer a relaxing route, a scenic route can be proposed.

[0363] The terminal receives guidance information from the server and displays it visually on digital signage using augmented reality technology. This allows users to intuitively receive and easily understand the information.

[0364] This system aims to reduce stress and provide a richer travel experience by overcoming language barriers when traveling in foreign countries and offering optimal navigation tailored to the user's emotions. Examples of prompts generated using a generative AI model are as follows:

[0365] "When a user says they want to quickly move on to the next tourist spot, and the emotion engine detects a stressed state, please generate the optimal route to provide."

[0366] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0367] Step 1:

[0368] The user voice-inputs their destination into the microphone on the digital signage. The input data is the user's voice, which is then transmitted to the terminal. The terminal activates a speech recognition mechanism and, utilizing the Google Cloud Speech-to-Text API and other tools, converts the voice into text data in real time. As a result, the text information extracted from the voice input is output.

[0369] Step 2:

[0370] The device inputs the text information generated by speech recognition into the emotion engine. The emotion engine uses Azure Cognitive Services and other tools to analyze the emotional nuances contained in the speech. This process outputs emotional information such as whether the user is relaxed or tense.

[0371] Step 3:

[0372] The terminal sends text information and sentiment information together to the server. The input data consists of stringified destination information and sentiment data. The SSL / TLS protocol is used to securely transmit the data to the server. The output is the destination information and sentiment information received on the server side.

[0373] Step 4:

[0374] The server analyzes the received destination information and extracts relevant keywords using the Alpaca API. Based on this input data, it outputs processed key data. Next, it translates the destination information into multiple languages ​​using the DeepL API and outputs the translated data.

[0375] Step 5:

[0376] The server uses the Google Maps API to obtain real-time travel information. This input includes the user's destination information. Based on the acquired travel data, the server generates the optimal travel route, taking into account the user's emotional state. The output is personalized guidance information.

[0377] Step 6:

[0378] The terminal displays navigation information received from the server on a digital signage screen. The input is navigation information from the server, and the output is a navigation display enhanced using augmented reality technology. This allows the user to intuitively understand directions to their destination and travel comfortably.

[0379] (Application Example 2)

[0380] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0381] When traveling in a foreign country, there is a need to alleviate the language barrier and emotional stress that users face and provide a comfortable travel experience. Conventional navigation systems present routes in a single way without considering the user's emotional state, making it difficult to meet the diverse needs of users.

[0382] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0383] In this invention, the server includes a speech recognition module and means for converting the user's voice input into text information, means for identifying a destination based on the text information and determining the user's emotional state using emotion analysis technology, and means for creating guidance information by translating it into multiple languages ​​and incorporating feedback according to the user's emotional state. As a result, even in a foreign country, the user will be guided to the optimal travel route according to their emotional state, enabling a stress-free travel experience without feeling the language barrier.

[0384] A "speech recognition module" is a device that receives voice input from a user, analyzes it, and converts it into text information.

[0385] "Feedback tailored to the user's emotional state" refers to a function that analyzes the user's emotions and provides appropriate information and guidance based on that state.

[0386] "Emotional analysis technology" is a technology that analyzes the tone of a user's voice and the content of their speech to identify their emotions.

[0387] "Translating into multiple languages" is the process of converting one language into several other languages, providing information in a way that is understandable to users whose native languages ​​are different.

[0388] To realize this invention, it is necessary to build a system that effectively combines a speech recognition module, emotion analysis technology, a multilingual translation engine, and augmented reality technology. The server receives the user's voice input, analyzes it to identify the destination and determine the emotion. The voice input is converted into text information using a speech recognition engine (e.g., Google Speech-to-Text). This text information is then used to determine the user's emotional state through an emotion analysis engine. The emotion analysis takes into account the tone and content of the voice.

[0389] The device uses a multilingual translation engine to provide guidance information that can be understood in the user's native language. The generated guidance information includes the optimal travel route, taking real-time traffic information into account. This allows users to travel to their destination comfortably and efficiently without experiencing language barriers.

[0390] Augmented reality technology is employed for visual presentation. The terminal visually highlights and presents guidance information, allowing users to intuitively understand the information. This technology makes even complex information easy to grasp.

[0391] As a concrete example, if a native Japanese speaker uses voice input to say, "I want to get to the next tourist spot quickly" while at a tourist destination abroad, the server will analyze the slightly tense tone of voice and suggest the fastest route. Example prompt: "Translate 'How do I get to the Eiffel Tower?' into French, including sentiment analysis."

[0392] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0393] Step 1:

[0394] The user provides voice input, and the device receives the audio. The input is the user's voice data, which the device sends to a speech recognition engine (e.g., Google Speech-to-Text) to be converted into text. This converted text becomes the output of the next process.

[0395] Step 2:

[0396] The server receives text information and uses an emotion analysis engine to determine the user's emotional state. The input is text information, and emotion analysis technology is used to analyze the user's tone of voice and word choices to determine whether the user is stressed or relaxed. The output is the analyzed emotional state.

[0397] Step 3:

[0398] The server analyzes text information to identify the destination and extract location information. The input is text information, and natural language processing is used to identify the place name or facility name of the destination. The output is the location information of the destination.

[0399] Step 4:

[0400] The server uses a multilingual translation engine to translate the analysis results into a language the user can understand. The input consists of the guidance information and the user's native language, and the translation is performed using a language model. The output is the translated guidance information.

[0401] Step 5:

[0402] The server acquires real-time traffic information and calculates the optimal travel route based on the identified destination and the user's emotional state. The input consists of real-time traffic data and emotional analysis results. A route planning algorithm is used to identify the best route for the user. The output is the optimal travel route.

[0403] Step 6:

[0404] The device uses augmented reality technology to visually present guidance information to the user. Input consists of translated guidance information and the optimal travel route, presented in a visually enhanced format through the display device. Output is an intuitive visual display of guidance information for the user.

[0405] This entire process allows users to obtain a comfortable mode of transportation that suits their needs, without being hindered by language barriers.

[0406] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0407] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0408] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0409] [Third Embodiment]

[0410] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0411] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0412] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0413] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0414] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0415] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0416] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0417] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0418] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0419] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0420] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0421] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0422] This invention is a multilingual real-time navigation system designed to enable tourists and visitors to smoothly reach their destinations in foreign lands. Specifically, a terminal installed as part of digital signage accepts voice input from users.

[0423] Users can use their device's voice input function to communicate their destination in their native language. For example, they can give instructions such as, "Tell me how to get to the art museum."

[0424] The terminal uses a speech recognition engine to convert the user's voice into text. The converted text is then sent directly to the server.

[0425] The server identifies the destination from the received text information and performs the corresponding multilingual translation process. This translation utilizes generative AI to generate multilingual guidance information. Furthermore, real-time traffic information is incorporated to calculate the optimal route.

[0426] Navigation information generated by the server is transmitted to the terminal in a visually easy-to-understand format and provided to the user through a visual display on the terminal.

[0427] For example, if a user says "I want to go to Central Park" in front of a digital signage screen, the speech recognition engine converts the speech into text and sends it to the server. The server recognizes "Central Park" as the destination and generates multilingual directions. These directions include the best transportation options and recommended routes based on current traffic conditions.

[0428] Users can follow the displayed information and travel to their destination without stress. If augmented reality technology is used in this process, the guidance information is visually overlaid, allowing users to understand it intuitively.

[0429] The following describes the processing flow.

[0430] Step 1:

[0431] The user speaks their destination into the voice input device on the digital signage. At this time, the user can naturally specify the destination in their native language.

[0432] Step 2:

[0433] The device receives voice input and uses its built-in speech recognition engine to convert the voice data into text. The converted text contains information about the destination specified by the user.

[0434] Step 3:

[0435] The terminal constructs the converted character information as a data packet and sends it to the server. This packet contains the user's voice input recorded as a string of characters.

[0436] Step 4:

[0437] The server analyzes the received text information to identify the destination. It then compares it with a tourism database to retrieve detailed information related to the destination.

[0438] Step 5:

[0439] The server uses a generative AI model to translate the analyzed destination information into multiple languages. Based on the translation results, it constructs navigation information and makes it available in a displayable format.

[0440] Step 6:

[0441] The server acquires real-time traffic information and calculates the optimal travel route. This calculation takes into account traffic congestion and the operating status of public transportation.

[0442] Step 7:

[0443] The server sends the generated navigation information to the terminal. This data includes visual map information and translated text.

[0444] Step 8:

[0445] The terminal displays the received information on a visual display device. The displayed information may also include visual effects utilizing augmented reality technology.

[0446] Step 9:

[0447] The user begins their journey based on the displayed navigation information. Additional voice instructions may be given to help them understand the route to their destination.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0450] A challenge exists in that users who speak different languages ​​often find it difficult to reach their destinations accurately and smoothly in unfamiliar places. Therefore, there is a need for a system that provides real-time multilingual support and optimal routes. Furthermore, there is a need for a method of presenting guidance information in an intuitively understandable format.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes a voice conversion means that converts the user's voice input into text information, a means that analyzes the text information to identify the destination and uses a generative AI model to translate it into multiple languages ​​and create guidance information, and a means that acquires real-time information and calculates the optimal travel route based on the guidance information. This makes it possible for users from different language regions to reach their destination intuitively and quickly.

[0453] "Voice conversion means" refers to a technology or device for converting a user's voice input into text information.

[0454] A "generative AI model" is a generative artificial intelligence model intended for natural language processing and translation, and it performs information conversion between different languages.

[0455] "Real-time information" refers to information that is acquired and updated immediately in accordance with the current situation.

[0456] "Guidance information" refers to information that provides users with directions for travel, including routes to their destination and related information.

[0457] "Visual display means" refers to a device or technology for visually presenting guidance information to users.

[0458] Augmented reality technology is a technology that overlays digital information onto visual information from the real world.

[0459] A "prompt message" is a text message used to convey instructions or requests to a generative AI model.

[0460] This invention is a multilingual real-time navigation system for users who speak different languages ​​to reach their destination. This system provides guidance information to the user by combining voice conversion means, a generative AI model, multilingual translation, real-time information acquisition, and visual display means.

[0461] The user voice-inputs their destination in their native language into a digital display terminal. For example, they might say, "Please tell me how to get to the museum." This voice input is converted into text by the speech recognition engine built into the terminal. The speech recognition engine used here includes various speech recognition technologies as general software.

[0462] The terminal sends data to the server based on the converted text information. The server uses a generative AI model to analyze the received text information. The generative AI model incorporates technologies that enable diverse language conversion. The server uses this model to create multilingual guidance information. This information includes the optimal travel route to the destination and current traffic conditions.

[0463] The server utilizes a traffic information acquisition system to analyze traffic conditions using real-time data and calculate the optimal travel route. In this case, the route is adjusted to optimize it according to the traffic conditions.

[0464] The generated navigation information is transmitted to the terminal in a visually easy-to-understand format and displayed to the user by the terminal. The displayed content includes digital maps and directional icons so that users can understand it immediately. This allows users to reach their destination smoothly.

[0465] As a concrete example, consider a user prompt: "I want to go to Central Park." Based on this input, the server uses a generative AI model to create directions and provide the user with the most appropriate navigation.

[0466] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0467] Step 1:

[0468] The user voice-inputs their destination in their native language into a digital display terminal. The input voice is received by a "voice recognition engine," which then serves as initial data for processing in the next step.

[0469] Step 2:

[0470] The device activates its speech recognition engine and converts the received audio into text information in real time. This conversion process analyzes the audio signal and converts the spoken content into text format. The resulting text information is specific text data, such as "Please tell me how to get to the art museum." The main data calculation here is the conversion of the audio signal into a string of characters.

[0471] Step 3:

[0472] The terminal sends the converted character information to the server. The character information, as input, is transferred to the server using a secure communication protocol. The output becomes data for destination identification and translation in the next step.

[0473] Step 4:

[0474] The server analyzes the received text information. In this step, it uses text analysis technology to identify the destination based on the input text information. Subsequently, a generative AI model is used to translate the information into multiple languages. The output is the translated guidance information data. The data processing here involves text analysis and language conversion.

[0475] Step 5:

[0476] The server calls a traffic information acquisition system to obtain real-time traffic information. Using destination information as input, it calculates the optimal route considering traffic conditions. As a result of this step, navigation guidance incorporating traffic information is generated. The output is optimal travel route data.

[0477] Step 6:

[0478] The server sends the generated multilingual navigation information to the terminal. This transmission process provides information including the selected optimal route and directions to the destination. The data output is provided in a format suitable for visual display.

[0479] Step 7:

[0480] The terminal displays navigation information to the user using visual display means. Specifically, maps, transportation information, walking routes, etc., are displayed on the terminal's screen. The input is navigation data received from the server, and the output is visual information that the user can refer to in real time.

[0481] (Application Example 1)

[0482] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0483] Modern travelers and visitors face challenges in navigating foreign lands, including the need for multilingual support and real-time traffic information. This challenge stems from the need for fast and accurate navigation to ensure travelers with diverse language and cultural backgrounds reach their destinations smoothly. Furthermore, there is a lack of effective methods for integrating real-world and digital information and presenting it visually and clearly.

[0484] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0485] In this invention, the server includes means for converting the user's voice input into text information, means for identifying the destination and generating guidance information by translating it into multiple languages, and means for visually overlaying and presenting the guidance information using AR technology. As a result, even in a foreign country, users can receive appropriate guidance in an intuitively understandable form and reach their destination quickly and efficiently.

[0486] "Speech recognition means" refers to a technology that converts a user's voice into text information, and is a means that analyzes voice input as a string of characters.

[0487] "Means for identifying a destination" refers to methods for identifying a destination from information entered by the user and deriving information about related locations.

[0488] "Methods for translating into multiple languages" refer to technologies that convert information expressed in one language into different languages, making it understandable to users who speak those languages.

[0489] "Means for generating guidance information" refers to technologies for creating detailed route guidance and traffic information related to a destination, and is a means of constructing information to be provided to users.

[0490] "AR technology" refers to augmented reality technology, which is a technology that overlays digital information onto the real world environment.

[0491] "Visual overlay presentation methods" refer to technologies that integrate and display digital information within the user's field of vision in a way that is intuitively understandable.

[0492] The system for carrying out this invention consists of a voice recognition means, a destination identification means, a multi-language translation means, a guidance information generation means, and a visual presentation means that utilizes AR technology.

[0493] The user enters their destination by voice via their smart device. The device is equipped with a speech recognition engine and uses the Google Speech-to-Text API to convert the voice data into text. The converted text is then sent to the server.

[0494] The server analyzes the received text information to identify the destination. The Google Translate API is used to generate multilingual guidance information for the destination using AI generation. This makes it possible to provide guidance in different languages ​​selected by the user. Furthermore, the optimal travel route is calculated while considering real-time traffic information.

[0495] Unity and Vuforia are used to leverage augmented reality technology as a visual presentation tool. The generated navigation information is presented to the user through smart glasses or other visual devices. This allows users to intuitively receive navigation information in a form overlaid on the real world.

[0496] As a concrete example, if a user uses smart glasses and says, "I want to go to the museum," the voice is converted to text and sent to a server. The server generates the optimal route and multilingual navigation to the destination "museum" and displays it on the smart glasses using augmented reality. This allows the user to intuitively understand the route.

[0497] Examples of prompts for a generative AI model:

[0498] "Destination: Museum. Please generate multilingual navigation directions for the best route from my current location."

[0499] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0500] Step 1:

[0501] The user uses the microphone function of their smart device to input destination information by voice. The input voice data is received by the device's voice recognition engine.

[0502] Step 2:

[0503] The device uses the Google Speech-to-Text API to convert audio data into text. This data processing involves real-time analysis of the audio waveform to generate text, which in turn generates a string of characters. This string is then output and sent to the server.

[0504] Step 3:

[0505] The server analyzes the received string information to identify the destination. Natural language processing is used for the string data, extracting the destination and recognizing it as output. The recognized destination information is then passed on to the next step.

[0506] Step 4:

[0507] The server acquires real-time traffic information based on destination and current location information and calculates the optimal route. The output is generated as optimal route information. In this process, the best route is derived through data calculations performed by referencing an external traffic information database.

[0508] Step 5:

[0509] The server uses the Google Translate API to access a generation AI model and create multilingual navigation guidance information. It sends prompt text to the generation AI model, which then generates multilingual guidance information about the destination as output.

[0510] Step 6:

[0511] The device uses Unity and Vuforia to display the obtained multilingual guidance information as a visual presentation on smart glasses. In practice, AR technology is used to overlay digital information onto the real-world scenery, providing the user with intuitive navigation.

[0512] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0513] This invention is a multilingual real-time navigation system designed to make travel more comfortable for visitors in foreign countries, and it includes a function to recognize user emotions. The system's main components are a terminal embedded in digital signage and a server connected to it.

[0514] The user speaks their destination into the voice input on the digital signage. This voice input is converted into text by a speech recognition engine. The voice often contains hints indicating the user's emotions, and this information is also analyzed by an emotion engine.

[0515] The device converts speech to text while simultaneously using an emotion engine to analyze the user's emotions. This analysis determines the user's emotional state, such as whether they are relaxed or stressed. This allows for the incorporation of emotion-responsive feedback into navigation information.

[0516] The server analyzes text and emotional information received from the terminal to identify the destination. It then performs multilingual translation and generates navigation information. This generated navigation information incorporates real-time traffic data, suggesting the optimal travel route based on the user's emotional state.

[0517] For example, if a user enters "I want to get to the next tourist spot quickly" in a somewhat tense voice, the emotion engine will detect the stress level, and the server will generate guidance information that prioritizes the fastest route. On the other hand, if a user says in a relaxed voice, "Tell me a route with nice scenery," it is possible to incorporate scenic routes into the guidance information to provide a relaxing sightseeing experience.

[0518] The terminal displays information received from the server to the user on a visual display device. The displayed guidance information is visually enhanced using augmented reality technology, making it easier for the user to intuitively understand the navigation information.

[0519] Thus, the present invention aims to reduce the stress of travel and enrich the travel experience of visitors by providing a navigation system with emotion recognition capabilities.

[0520] The following describes the processing flow.

[0521] Step 1:

[0522] The user speaks their desired destination in their native language into the voice input device on the digital signage. In this process, the user's voice may naturally convey emotions.

[0523] Step 2:

[0524] The device receives voice input and converts the voice data into text using a speech recognition engine. This converted text contains information about the destination.

[0525] Step 3:

[0526] The device uses an emotion engine to analyze the emotional characteristics during voice input. The analysis identifies the user's emotional state, such as whether they are relaxed or tense.

[0527] Step 4:

[0528] The terminal packages the processed text information and emotional state data into a packet and sends it to the server. This packet contains both destination information and emotional data.

[0529] Step 5:

[0530] The server identifies the destination based on the received text information and compares it with a tourism database. Furthermore, it uses a generative AI model to translate destination-related information into multiple languages.

[0531] Step 6:

[0532] The server customizes navigation information based on the user's emotional state. For example, if the user is stressed, it prioritizes the fastest route; if they are relaxed, it suggests a route that includes tourist attractions.

[0533] Step 7:

[0534] The server calculates the optimal travel route, taking real-time traffic information into account. This calculation takes into account current travel conditions such as traffic congestion and public transport.

[0535] Step 8:

[0536] The server sends the generated, customized guidance information to the terminal. The data includes detailed maps and translated guidance text.

[0537] Step 9:

[0538] The terminal displays the received guidance information to the user on a visual display device. Using augmented reality technology, the guidance information is visually overlaid, allowing the user to visually confirm the guidance on the spot.

[0539] Step 10:

[0540] Based on the provided guidance information, users can comfortably begin their journey. The provision of emotionally considerate services allows them to enjoy a stress-free travel experience.

[0541] (Example 2)

[0542] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0543] When traveling in a foreign country, there is a need to provide personalized navigation information that responds to the user's emotional state while addressing language barriers and real-time, ever-changing traffic conditions. Furthermore, there is a lack of easily understandable visual methods, making it a challenge to provide a less stressful travel experience.

[0544] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0545] In this invention, the server includes means for recording the user's voice input and converting the voice into text information using a voice recognition mechanism; means including a mechanism for analyzing emotional information from the voice; and means for identifying a destination based on the text information and emotional information, translating it into multiple languages, and creating guidance information that corresponds to the user's emotional state. This makes it possible to provide users with intuitive, easy-to-understand, and emotionally sensitive real-time navigation information.

[0546] A "speech recognition mechanism" is a technology that converts a user's voice input into text data, and is a means of changing speech into a format that can be processed by a computer.

[0547] "Emotional information" is data that represents the user's psychological state by analyzing the emotional nuances extracted from the user's voice.

[0548] "Multilingual translation" refers to the process of converting information expressed in one language into multiple other languages, and is a technology that enables communication between people who speak different languages.

[0549] "Visual display devices" are devices used to visually present digital information to users, and include output devices such as screens and displays.

[0550] "Real-time travel information" refers to the latest travel data based on current traffic and road conditions, and is information used to respond to changing travel environments.

[0551] Augmented reality technology is a technology that overlays digital information onto the real world environment, enabling users to experience digital information in a realistic context.

[0552] "User emotional state" refers to the user's psychological and emotional condition and is an important indicator for personalizing navigation information.

[0553] This invention is a multilingual real-time navigation system that smoothly supports travel in foreign lands and is equipped with emotion recognition capabilities. The system mainly consists of a terminal embedded in digital signage and a server connected to it.

[0554] The user speaks their destination to the voice input device on the digital signage. The voice input is converted into text in real time by a speech recognition mechanism utilizing the Google Cloud Speech-to-Text API and other tools. In addition, data indicating the user's emotional state is analyzed from the voice using Azure Cognitive Services and other tools. This allows the system to determine the user's psychological state, such as tension or relaxation.

[0555] The device sends this analyzed text data and sentiment information to the server. SSL / TLS protocol is used for communication to ensure data security. Upon receiving this data, the server uses the Alpaca API to extract keywords for destination information and the DeepL API to perform multilingual translation. Furthermore, by obtaining real-time travel status information via the Google Maps API, it generates the optimal travel route. This combination of data creates guidance information tailored to the user's emotional state. For example, if the user is in a hurry, the fastest route can be suggested; if they prefer a relaxing route, a scenic route can be proposed.

[0556] The terminal receives guidance information from the server and displays it visually on digital signage using augmented reality technology. This allows users to intuitively receive and easily understand the information.

[0557] This system aims to reduce stress and provide a richer travel experience by overcoming language barriers when traveling in foreign countries and offering optimal navigation tailored to the user's emotions. Examples of prompts generated using a generative AI model are as follows:

[0558] "When a user says they want to quickly move on to the next tourist spot, and the emotion engine detects a stressed state, please generate the optimal route to provide."

[0559] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0560] Step 1:

[0561] The user voice-inputs their destination into the microphone on the digital signage. The input data is the user's voice, which is then transmitted to the terminal. The terminal activates a speech recognition mechanism and, utilizing the Google Cloud Speech-to-Text API and other tools, converts the voice into text data in real time. As a result, the text information extracted from the voice input is output.

[0562] Step 2:

[0563] The device inputs the text information generated by speech recognition into the emotion engine. The emotion engine uses Azure Cognitive Services and other tools to analyze the emotional nuances contained in the speech. This process outputs emotional information such as whether the user is relaxed or tense.

[0564] Step 3:

[0565] The terminal sends text information and sentiment information together to the server. The input data consists of stringified destination information and sentiment data. The SSL / TLS protocol is used to securely transmit the data to the server. The output is the destination information and sentiment information received on the server side.

[0566] Step 4:

[0567] The server analyzes the received destination information and extracts relevant keywords using the Alpaca API. Based on this input data, it outputs processed key data. Next, it translates the destination information into multiple languages ​​using the DeepL API and outputs the translated data.

[0568] Step 5:

[0569] The server uses the Google Maps API to obtain real-time travel information. This input includes the user's destination information. Based on the acquired travel data, the server generates the optimal travel route, taking into account the user's emotional state. The output is personalized guidance information.

[0570] Step 6:

[0571] The terminal displays navigation information received from the server on a digital signage screen. The input is navigation information from the server, and the output is a navigation display enhanced using augmented reality technology. This allows the user to intuitively understand directions to their destination and travel comfortably.

[0572] (Application Example 2)

[0573] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0574] When traveling in a foreign country, there is a need to alleviate the language barrier and emotional stress that users face and provide a comfortable travel experience. Conventional navigation systems present routes in a single way without considering the user's emotional state, making it difficult to meet the diverse needs of users.

[0575] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0576] In this invention, the server includes a speech recognition module and means for converting the user's voice input into text information, means for identifying a destination based on the text information and determining the user's emotional state using emotion analysis technology, and means for creating guidance information by translating it into multiple languages ​​and incorporating feedback according to the user's emotional state. As a result, even in a foreign country, the user will be guided to the optimal travel route according to their emotional state, enabling a stress-free travel experience without feeling the language barrier.

[0577] A "speech recognition module" is a device that receives voice input from a user, analyzes it, and converts it into text information.

[0578] "Feedback tailored to the user's emotional state" refers to a function that analyzes the user's emotions and provides appropriate information and guidance based on that state.

[0579] "Emotional analysis technology" is a technology that analyzes the tone of a user's voice and the content of their speech to identify their emotions.

[0580] "Translating into multiple languages" is the process of converting one language into several other languages, providing information in a way that is understandable to users whose native languages ​​are different.

[0581] To realize this invention, it is necessary to build a system that effectively combines a speech recognition module, emotion analysis technology, a multilingual translation engine, and augmented reality technology. The server receives the user's voice input, analyzes it to identify the destination and determine the emotion. The voice input is converted into text information using a speech recognition engine (e.g., Google Speech-to-Text). This text information is then used to determine the user's emotional state through an emotion analysis engine. The emotion analysis takes into account the tone and content of the voice.

[0582] The device uses a multilingual translation engine to provide guidance information that can be understood in the user's native language. The generated guidance information includes the optimal travel route, taking real-time traffic information into account. This allows users to travel to their destination comfortably and efficiently without experiencing language barriers.

[0583] Augmented reality technology is employed for visual presentation. The terminal visually highlights and presents guidance information, allowing users to intuitively understand the information. This technology makes even complex information easy to grasp.

[0584] As a concrete example, if a native Japanese speaker uses voice input to say, "I want to get to the next tourist spot quickly" while at a tourist destination abroad, the server will analyze the slightly tense tone of voice and suggest the fastest route. Example prompt: "Translate 'How do I get to the Eiffel Tower?' into French, including sentiment analysis."

[0585] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0586] Step 1:

[0587] The user provides voice input, and the device receives the audio. The input is the user's voice data, which the device sends to a speech recognition engine (e.g., Google Speech-to-Text) to be converted into text. This converted text becomes the output of the next process.

[0588] Step 2:

[0589] The server receives text information and uses an emotion analysis engine to determine the user's emotional state. The input is text information, and emotion analysis technology is used to analyze the user's tone of voice and word choices to determine whether the user is stressed or relaxed. The output is the analyzed emotional state.

[0590] Step 3:

[0591] The server analyzes text information to identify the destination and extract location information. The input is text information, and natural language processing is used to identify the place name or facility name of the destination. The output is the location information of the destination.

[0592] Step 4:

[0593] The server uses a multilingual translation engine to translate the analysis results into a language the user can understand. The input consists of the guidance information and the user's native language, and the translation is performed using a language model. The output is the translated guidance information.

[0594] Step 5:

[0595] The server acquires real-time traffic information and calculates the optimal travel route based on the identified destination and the user's emotional state. The input consists of real-time traffic data and emotional analysis results. A route planning algorithm is used to identify the best route for the user. The output is the optimal travel route.

[0596] Step 6:

[0597] The device uses augmented reality technology to visually present guidance information to the user. Input consists of translated guidance information and the optimal travel route, presented in a visually enhanced format through the display device. Output is an intuitive visual display of guidance information for the user.

[0598] This entire process allows users to obtain a comfortable mode of transportation that suits their needs, without being hindered by language barriers.

[0599] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0600] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0601] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0602] [Fourth Embodiment]

[0603] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0604] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0605] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0606] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0607] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0608] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0609] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0610] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0611] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0612] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0613] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0614] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0615] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0616] This invention is a multilingual real-time navigation system designed to enable tourists and visitors to smoothly reach their destinations in foreign lands. Specifically, a terminal installed as part of digital signage accepts voice input from users.

[0617] Users can use their device's voice input function to communicate their destination in their native language. For example, they can give instructions such as, "Tell me how to get to the art museum."

[0618] The terminal uses a speech recognition engine to convert the user's voice into text. The converted text is then sent directly to the server.

[0619] The server identifies the destination from the received text information and performs the corresponding multilingual translation process. This translation utilizes generative AI to generate multilingual guidance information. Furthermore, real-time traffic information is incorporated to calculate the optimal route.

[0620] Navigation information generated by the server is transmitted to the terminal in a visually easy-to-understand format and provided to the user through a visual display on the terminal.

[0621] For example, if a user says "I want to go to Central Park" in front of a digital signage screen, the speech recognition engine converts the speech into text and sends it to the server. The server recognizes "Central Park" as the destination and generates multilingual directions. These directions include the best transportation options and recommended routes based on current traffic conditions.

[0622] Users can follow the displayed information and travel to their destination without stress. If augmented reality technology is used in this process, the guidance information is visually overlaid, allowing users to understand it intuitively.

[0623] The following describes the processing flow.

[0624] Step 1:

[0625] The user speaks their destination into the voice input device on the digital signage. At this time, the user can naturally specify the destination in their native language.

[0626] Step 2:

[0627] The device receives voice input and uses its built-in speech recognition engine to convert the voice data into text. The converted text contains information about the destination specified by the user.

[0628] Step 3:

[0629] The terminal constructs the converted character information as a data packet and sends it to the server. This packet contains the user's voice input recorded as a string of characters.

[0630] Step 4:

[0631] The server analyzes the received text information to identify the destination. It then compares it with a tourism database to retrieve detailed information related to the destination.

[0632] Step 5:

[0633] The server uses a generative AI model to translate the analyzed destination information into multiple languages. Based on the translation results, it constructs navigation information and makes it available in a displayable format.

[0634] Step 6:

[0635] The server acquires real-time traffic information and calculates the optimal travel route. This calculation takes into account traffic congestion and the operating status of public transportation.

[0636] Step 7:

[0637] The server sends the generated navigation information to the terminal. This data includes visual map information and translated text.

[0638] Step 8:

[0639] The terminal displays the received information on a visual display device. The displayed information may also include visual effects utilizing augmented reality technology.

[0640] Step 9:

[0641] The user begins their journey based on the displayed navigation information. Additional voice instructions may be given to help them understand the route to their destination.

[0642] (Example 1)

[0643] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0644] A challenge exists in that users who speak different languages ​​often find it difficult to reach their destinations accurately and smoothly in unfamiliar places. Therefore, there is a need for a system that provides real-time multilingual support and optimal routes. Furthermore, there is a need for a method of presenting guidance information in an intuitively understandable format.

[0645] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0646] In this invention, the server includes a voice conversion means that converts the user's voice input into text information, a means that analyzes the text information to identify the destination and uses a generative AI model to translate it into multiple languages ​​and create guidance information, and a means that acquires real-time information and calculates the optimal travel route based on the guidance information. This makes it possible for users from different language regions to reach their destination intuitively and quickly.

[0647] "Voice conversion means" refers to a technology or device for converting a user's voice input into text information.

[0648] A "generative AI model" is a generative artificial intelligence model intended for natural language processing and translation, and it performs information conversion between different languages.

[0649] "Real-time information" refers to information that is acquired and updated immediately in accordance with the current situation.

[0650] "Guidance information" refers to information that provides users with directions for travel, including routes to their destination and related information.

[0651] "Visual display means" refers to a device or technology for visually presenting guidance information to users.

[0652] Augmented reality technology is a technology that overlays digital information onto visual information from the real world.

[0653] A "prompt message" is a text message used to convey instructions or requests to a generative AI model.

[0654] This invention is a multilingual real-time navigation system for users who speak different languages ​​to reach their destination. This system provides guidance information to the user by combining voice conversion means, a generative AI model, multilingual translation, real-time information acquisition, and visual display means.

[0655] The user voice-inputs their destination in their native language into a digital display terminal. For example, they might say, "Please tell me how to get to the museum." This voice input is converted into text by the speech recognition engine built into the terminal. The speech recognition engine used here includes various speech recognition technologies as general software.

[0656] The terminal sends data to the server based on the converted text information. The server uses a generative AI model to analyze the received text information. The generative AI model incorporates technologies that enable diverse language conversion. The server uses this model to create multilingual guidance information. This information includes the optimal travel route to the destination and current traffic conditions.

[0657] The server utilizes a traffic information acquisition system to analyze traffic conditions using real-time data and calculate the optimal travel route. In this case, the route is adjusted to optimize it according to the traffic conditions.

[0658] The generated navigation information is transmitted to the terminal in a visually easy-to-understand format and displayed to the user by the terminal. The displayed content includes digital maps and directional icons so that users can understand it immediately. This allows users to reach their destination smoothly.

[0659] As a concrete example, consider a user prompt: "I want to go to Central Park." Based on this input, the server uses a generative AI model to create directions and provide the user with the most appropriate navigation.

[0660] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0661] Step 1:

[0662] The user voice-inputs their destination in their native language into a digital display terminal. The input voice is received by a "voice recognition engine," which then serves as initial data for processing in the next step.

[0663] Step 2:

[0664] The device activates its speech recognition engine and converts the received audio into text information in real time. This conversion process analyzes the audio signal and converts the spoken content into text format. The resulting text information is specific text data, such as "Please tell me how to get to the art museum." The main data calculation here is the conversion of the audio signal into a string of characters.

[0665] Step 3:

[0666] The terminal sends the converted character information to the server. The character information, as input, is transferred to the server using a secure communication protocol. The output becomes data for destination identification and translation in the next step.

[0667] Step 4:

[0668] The server analyzes the received text information. In this step, it uses text analysis technology to identify the destination based on the input text information. Subsequently, a generative AI model is used to translate the information into multiple languages. The output is the translated guidance information data. The data processing here involves text analysis and language conversion.

[0669] Step 5:

[0670] The server calls a traffic information acquisition system to obtain real-time traffic information. Using destination information as input, it calculates the optimal route considering traffic conditions. As a result of this step, navigation guidance incorporating traffic information is generated. The output is optimal travel route data.

[0671] Step 6:

[0672] The server sends the generated multilingual navigation information to the terminal. This transmission process provides information including the selected optimal route and directions to the destination. The data output is provided in a format suitable for visual display.

[0673] Step 7:

[0674] The terminal displays navigation information to the user using visual display means. Specifically, maps, transportation information, walking routes, etc., are displayed on the terminal's screen. The input is navigation data received from the server, and the output is visual information that the user can refer to in real time.

[0675] (Application Example 1)

[0676] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] Modern travelers and visitors face challenges in navigating foreign lands, including the need for multilingual support and real-time traffic information. This challenge stems from the need for fast and accurate navigation to ensure travelers with diverse language and cultural backgrounds reach their destinations smoothly. Furthermore, there is a lack of effective methods for integrating real-world and digital information and presenting it visually and clearly.

[0678] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0679] In this invention, the server includes means for converting the user's voice input into text information, means for identifying the destination and generating guidance information by translating it into multiple languages, and means for visually overlaying and presenting the guidance information using AR technology. As a result, even in a foreign country, users can receive appropriate guidance in an intuitively understandable form and reach their destination quickly and efficiently.

[0680] "Speech recognition means" refers to a technology that converts a user's voice into text information, and is a means that analyzes voice input as a string of characters.

[0681] "Means for identifying a destination" refers to methods for identifying a destination from information entered by the user and deriving information about related locations.

[0682] "Methods for translating into multiple languages" refer to technologies that convert information expressed in one language into different languages, making it understandable to users who speak those languages.

[0683] "Means for generating guidance information" refers to technologies for creating detailed route guidance and traffic information related to a destination, and is a means of constructing information to be provided to users.

[0684] "AR technology" refers to augmented reality technology, which is a technology that overlays digital information onto the real world environment.

[0685] "Visual overlay presentation methods" refer to technologies that integrate and display digital information within the user's field of vision in a way that is intuitively understandable.

[0686] The system for carrying out this invention consists of a voice recognition means, a destination identification means, a multi-language translation means, a guidance information generation means, and a visual presentation means that utilizes AR technology.

[0687] The user enters their destination by voice via their smart device. The device is equipped with a speech recognition engine and uses the Google Speech-to-Text API to convert the voice data into text. The converted text is then sent to the server.

[0688] The server analyzes the received text information to identify the destination. The Google Translate API is used to generate multilingual guidance information for the destination using AI generation. This makes it possible to provide guidance in different languages ​​selected by the user. Furthermore, the optimal travel route is calculated while considering real-time traffic information.

[0689] Unity and Vuforia are used to leverage augmented reality technology as a visual presentation tool. The generated navigation information is presented to the user through smart glasses or other visual devices. This allows users to intuitively receive navigation information in a form overlaid on the real world.

[0690] As a concrete example, if a user uses smart glasses and says, "I want to go to the museum," the voice is converted to text and sent to a server. The server generates the optimal route and multilingual navigation to the destination "museum" and displays it on the smart glasses using augmented reality. This allows the user to intuitively understand the route.

[0691] Examples of prompts for a generative AI model:

[0692] "Destination: Museum. Please generate multilingual navigation directions for the best route from my current location."

[0693] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0694] Step 1:

[0695] The user uses the microphone function of their smart device to input destination information by voice. The input voice data is received by the device's voice recognition engine.

[0696] Step 2:

[0697] The device uses the Google Speech-to-Text API to convert audio data into text. This data processing involves real-time analysis of the audio waveform to generate text, which in turn generates a string of characters. This string is then output and sent to the server.

[0698] Step 3:

[0699] The server analyzes the received string information to identify the destination. Natural language processing is used for the string data, extracting the destination and recognizing it as output. The recognized destination information is then passed on to the next step.

[0700] Step 4:

[0701] The server acquires real-time traffic information based on destination and current location information and calculates the optimal route. The output is generated as optimal route information. In this process, the best route is derived through data calculations performed by referencing an external traffic information database.

[0702] Step 5:

[0703] The server uses the Google Translate API to access a generation AI model and create multilingual navigation guidance information. It sends prompt text to the generation AI model, which then generates multilingual guidance information about the destination as output.

[0704] Step 6:

[0705] The device uses Unity and Vuforia to display the obtained multilingual guidance information as a visual presentation on smart glasses. In practice, AR technology is used to overlay digital information onto the real-world scenery, providing the user with intuitive navigation.

[0706] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0707] This invention is a multilingual real-time navigation system designed to make travel more comfortable for visitors in foreign countries, and it includes a function to recognize user emotions. The system's main components are a terminal embedded in digital signage and a server connected to it.

[0708] The user speaks their destination into the voice input on the digital signage. This voice input is converted into text by a speech recognition engine. The voice often contains hints indicating the user's emotions, and this information is also analyzed by an emotion engine.

[0709] The device converts speech to text while simultaneously using an emotion engine to analyze the user's emotions. This analysis determines the user's emotional state, such as whether they are relaxed or stressed. This allows for the incorporation of emotion-responsive feedback into navigation information.

[0710] The server analyzes text and emotional information received from the terminal to identify the destination. It then performs multilingual translation and generates navigation information. This generated navigation information incorporates real-time traffic data, suggesting the optimal travel route based on the user's emotional state.

[0711] For example, if a user enters "I want to get to the next tourist spot quickly" in a somewhat tense voice, the emotion engine will detect the stress level, and the server will generate guidance information that prioritizes the fastest route. On the other hand, if a user says in a relaxed voice, "Tell me a route with nice scenery," it is possible to incorporate scenic routes into the guidance information to provide a relaxing sightseeing experience.

[0712] The terminal displays information received from the server to the user on a visual display device. The displayed guidance information is visually enhanced using augmented reality technology, making it easier for the user to intuitively understand the navigation information.

[0713] Thus, the present invention aims to reduce the stress of travel and enrich the travel experience of visitors by providing a navigation system with emotion recognition capabilities.

[0714] The following describes the processing flow.

[0715] Step 1:

[0716] The user speaks their desired destination in their native language into the voice input device on the digital signage. In this process, the user's voice may naturally convey emotions.

[0717] Step 2:

[0718] The device receives voice input and converts the voice data into text using a speech recognition engine. This converted text contains information about the destination.

[0719] Step 3:

[0720] The device uses an emotion engine to analyze the emotional characteristics during voice input. The analysis identifies the user's emotional state, such as whether they are relaxed or tense.

[0721] Step 4:

[0722] The terminal packages the processed text information and emotional state data into a packet and sends it to the server. This packet contains both destination information and emotional data.

[0723] Step 5:

[0724] The server identifies the destination based on the received text information and compares it with a tourism database. Furthermore, it uses a generative AI model to translate destination-related information into multiple languages.

[0725] Step 6:

[0726] The server customizes navigation information based on the user's emotional state. For example, if the user is stressed, it prioritizes the fastest route; if they are relaxed, it suggests a route that includes tourist attractions.

[0727] Step 7:

[0728] The server calculates the optimal travel route, taking real-time traffic information into account. This calculation takes into account current travel conditions such as traffic congestion and public transport.

[0729] Step 8:

[0730] The server sends the generated, customized guidance information to the terminal. The data includes detailed maps and translated guidance text.

[0731] Step 9:

[0732] The terminal displays the received guidance information to the user on a visual display device. Using augmented reality technology, the guidance information is visually overlaid, allowing the user to visually confirm the guidance on the spot.

[0733] Step 10:

[0734] Based on the provided guidance information, users can comfortably begin their journey. The provision of emotionally considerate services allows them to enjoy a stress-free travel experience.

[0735] (Example 2)

[0736] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0737] When traveling in a foreign country, there is a need to provide personalized navigation information that responds to the user's emotional state while addressing language barriers and real-time, ever-changing traffic conditions. Furthermore, there is a lack of easily understandable visual methods, making it a challenge to provide a less stressful travel experience.

[0738] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0739] In this invention, the server includes means for recording the user's voice input and converting the voice into text information using a voice recognition mechanism; means including a mechanism for analyzing emotional information from the voice; and means for identifying a destination based on the text information and emotional information, translating it into multiple languages, and creating guidance information that corresponds to the user's emotional state. This makes it possible to provide users with intuitive, easy-to-understand, and emotionally sensitive real-time navigation information.

[0740] A "speech recognition mechanism" is a technology that converts a user's voice input into text data, and is a means of changing speech into a format that can be processed by a computer.

[0741] "Emotional information" is data that represents the user's psychological state by analyzing the emotional nuances extracted from the user's voice.

[0742] "Multilingual translation" refers to the process of converting information expressed in one language into multiple other languages, and is a technology that enables communication between people who speak different languages.

[0743] "Visual display devices" are devices used to visually present digital information to users, and include output devices such as screens and displays.

[0744] "Real-time travel information" refers to the latest travel data based on current traffic and road conditions, and is information used to respond to changing travel environments.

[0745] Augmented reality technology is a technology that overlays digital information onto the real world environment, enabling users to experience digital information in a realistic context.

[0746] "User emotional state" refers to the user's psychological and emotional condition and is an important indicator for personalizing navigation information.

[0747] This invention is a multilingual real-time navigation system that smoothly supports travel in foreign lands and is equipped with emotion recognition capabilities. The system mainly consists of a terminal embedded in digital signage and a server connected to it.

[0748] The user speaks their destination to the voice input device on the digital signage. The voice input is converted into text in real time by a speech recognition mechanism utilizing the Google Cloud Speech-to-Text API and other tools. In addition, data indicating the user's emotional state is analyzed from the voice using Azure Cognitive Services and other tools. This allows the system to determine the user's psychological state, such as tension or relaxation.

[0749] The device sends this analyzed text data and sentiment information to the server. SSL / TLS protocol is used for communication to ensure data security. Upon receiving this data, the server uses the Alpaca API to extract keywords for destination information and the DeepL API to perform multilingual translation. Furthermore, by obtaining real-time travel status information via the Google Maps API, it generates the optimal travel route. This combination of data creates guidance information tailored to the user's emotional state. For example, if the user is in a hurry, the fastest route can be suggested; if they prefer a relaxing route, a scenic route can be proposed.

[0750] The terminal receives guidance information from the server and displays it visually on digital signage using augmented reality technology. This allows users to intuitively receive and easily understand the information.

[0751] This system aims to reduce stress and provide a richer travel experience by overcoming language barriers when traveling in foreign countries and offering optimal navigation tailored to the user's emotions. Examples of prompts generated using a generative AI model are as follows:

[0752] "When a user says they want to quickly move on to the next tourist spot, and the emotion engine detects a stressed state, please generate the optimal route to provide."

[0753] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0754] Step 1:

[0755] The user voice-inputs their destination into the microphone on the digital signage. The input data is the user's voice, which is then transmitted to the terminal. The terminal activates a speech recognition mechanism and, utilizing the Google Cloud Speech-to-Text API and other tools, converts the voice into text data in real time. As a result, the text information extracted from the voice input is output.

[0756] Step 2:

[0757] The device inputs the text information generated by speech recognition into the emotion engine. The emotion engine uses Azure Cognitive Services and other tools to analyze the emotional nuances contained in the speech. This process outputs emotional information such as whether the user is relaxed or tense.

[0758] Step 3:

[0759] The terminal sends text information and sentiment information together to the server. The input data consists of stringified destination information and sentiment data. The SSL / TLS protocol is used to securely transmit the data to the server. The output is the destination information and sentiment information received on the server side.

[0760] Step 4:

[0761] The server analyzes the received destination information and extracts relevant keywords using the Alpaca API. Based on this input data, it outputs processed key data. Next, it translates the destination information into multiple languages ​​using the DeepL API and outputs the translated data.

[0762] Step 5:

[0763] The server uses the Google Maps API to obtain real-time travel information. This input includes the user's destination information. Based on the acquired travel data, the server generates the optimal travel route, taking into account the user's emotional state. The output is personalized guidance information.

[0764] Step 6:

[0765] The terminal displays navigation information received from the server on a digital signage screen. The input is navigation information from the server, and the output is a navigation display enhanced using augmented reality technology. This allows the user to intuitively understand directions to their destination and travel comfortably.

[0766] (Application Example 2)

[0767] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0768] When traveling in a foreign country, there is a need to alleviate the language barrier and emotional stress that users face and provide a comfortable travel experience. Conventional navigation systems present routes in a single way without considering the user's emotional state, making it difficult to meet the diverse needs of users.

[0769] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0770] In this invention, the server includes a speech recognition module and means for converting the user's voice input into text information, means for identifying a destination based on the text information and determining the user's emotional state using emotion analysis technology, and means for creating guidance information by translating it into multiple languages ​​and incorporating feedback according to the user's emotional state. As a result, even in a foreign country, the user will be guided to the optimal travel route according to their emotional state, enabling a stress-free travel experience without feeling the language barrier.

[0771] A "speech recognition module" is a device that receives voice input from a user, analyzes it, and converts it into text information.

[0772] "Feedback tailored to the user's emotional state" refers to a function that analyzes the user's emotions and provides appropriate information and guidance based on that state.

[0773] "Emotional analysis technology" is a technology that analyzes the tone of a user's voice and the content of their speech to identify their emotions.

[0774] "Translating into multiple languages" is the process of converting one language into several other languages, providing information in a way that is understandable to users whose native languages ​​are different.

[0775] To realize this invention, it is necessary to build a system that effectively combines a speech recognition module, emotion analysis technology, a multilingual translation engine, and augmented reality technology. The server receives the user's voice input, analyzes it to identify the destination and determine the emotion. The voice input is converted into text information using a speech recognition engine (e.g., Google Speech-to-Text). This text information is then used to determine the user's emotional state through an emotion analysis engine. The emotion analysis takes into account the tone and content of the voice.

[0776] The device uses a multilingual translation engine to provide guidance information that can be understood in the user's native language. The generated guidance information includes the optimal travel route, taking real-time traffic information into account. This allows users to travel to their destination comfortably and efficiently without experiencing language barriers.

[0777] Augmented reality technology is employed for visual presentation. The terminal visually highlights and presents guidance information, allowing users to intuitively understand the information. This technology makes even complex information easy to grasp.

[0778] As a concrete example, if a native Japanese speaker uses voice input to say, "I want to get to the next tourist spot quickly" while at a tourist destination abroad, the server will analyze the slightly tense tone of voice and suggest the fastest route. Example prompt: "Translate 'How do I get to the Eiffel Tower?' into French, including sentiment analysis."

[0779] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0780] Step 1:

[0781] The user provides voice input, and the device receives the audio. The input is the user's voice data, which the device sends to a speech recognition engine (e.g., Google Speech-to-Text) to be converted into text. This converted text becomes the output of the next process.

[0782] Step 2:

[0783] The server receives text information and uses an emotion analysis engine to determine the user's emotional state. The input is text information, and emotion analysis technology is used to analyze the user's tone of voice and word choices to determine whether the user is stressed or relaxed. The output is the analyzed emotional state.

[0784] Step 3:

[0785] The server analyzes text information to identify the destination and extract location information. The input is text information, and natural language processing is used to identify the place name or facility name of the destination. The output is the location information of the destination.

[0786] Step 4:

[0787] The server uses a multilingual translation engine to translate the analysis results into a language the user can understand. The input consists of the guidance information and the user's native language, and the translation is performed using a language model. The output is the translated guidance information.

[0788] Step 5:

[0789] The server acquires real-time traffic information and calculates the optimal travel route based on the identified destination and the user's emotional state. The input consists of real-time traffic data and emotional analysis results. A route planning algorithm is used to identify the best route for the user. The output is the optimal travel route.

[0790] Step 6:

[0791] The device uses augmented reality technology to visually present guidance information to the user. Input consists of translated guidance information and the optimal travel route, presented in a visually enhanced format through the display device. Output is an intuitive visual display of guidance information for the user.

[0792] This entire process allows users to obtain a comfortable mode of transportation that suits their needs, without being hindered by language barriers.

[0793] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0794] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0795] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0796] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0797] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0798] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0799] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0800] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0801] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0802] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0803] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0804] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0805] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0806] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0807] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0808] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0809] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0810] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0811] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0812] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0813] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0814] The following is further disclosed regarding the embodiments described above.

[0815] (Claim 1)

[0816] Equipped with a speech recognition engine, and a means for converting the user's voice input into text information,

[0817] A means for identifying a destination based on the aforementioned textual information and creating guidance information by translating it into multiple languages,

[0818] A means for presenting the aforementioned guidance information to the user via a visual display device,

[0819] A system that includes this.

[0820] (Claim 2)

[0821] The system according to claim 1, further comprising means for acquiring real-time traffic information and including in the guidance information an optimal travel route that takes traffic conditions into account.

[0822] (Claim 3)

[0823] The system according to claim 1, further comprising means for visually overlaying and presenting the guidance information displayed by the visual display device using augmented reality technology.

[0824] "Example 1"

[0825] (Claim 1)

[0826] A voice conversion means that converts the user's voice input into text information,

[0827] A means for analyzing the aforementioned textual information to identify the destination, and for creating guidance information by translating it into multiple languages ​​using a generative AI model,

[0828] A means for acquiring real-time information and calculating the optimal travel route based on the aforementioned guidance information,

[0829] Means for providing the aforementioned guidance information to the user via visual display means,

[0830] A device that includes this.

[0831] (Claim 2)

[0832] The apparatus according to claim 1, further comprising means for overlaying and presenting the guidance information displayed by the visual display means using augmented reality technology.

[0833] (Claim 3)

[0834] The apparatus according to claim 1, further comprising means for generating input prompt sentences for the generation AI model and optimizing the translation process.

[0835] "Application Example 1"

[0836] (Claim 1)

[0837] A speech recognition means for converting user voice input into text information,

[0838] A means for identifying a destination based on the aforementioned textual information and generating guidance information by translating it into multiple languages,

[0839] A means for presenting the aforementioned guidance information to the user via a display means,

[0840] A means of visually overlaying and presenting the aforementioned guidance information using AR technology,

[0841] A system that includes this.

[0842] (Claim 2)

[0843] The system according to claim 1, further comprising means for acquiring real-time traffic information and including an optimal route that takes traffic conditions into account in the guidance information, and means for displaying the information on a visual device that can be worn by the user.

[0844] (Claim 3)

[0845] The system according to claim 1, comprising means for sending a user's voice input to a server and creating user-specific multilingual navigation information using a generating AI.

[0846] "Example 2 of combining an emotion engine"

[0847] (Claim 1)

[0848] A means for recording the user's voice input and converting the voice into text information using a speech recognition mechanism,

[0849] A means including a mechanism for analyzing emotional information from the aforementioned audio,

[0850] A means for identifying a destination based on the aforementioned textual information and emotional information, translating it into multiple languages, and creating guidance information that corresponds to the user's emotional state,

[0851] A means for presenting the aforementioned guidance information to the user via a visual display device,

[0852] A system that includes this.

[0853] (Claim 2)

[0854] The system according to claim 1, further comprising means for acquiring real-time travel status information and including an optimal travel route that takes travel status into consideration in the guidance information.

[0855] (Claim 3)

[0856] The system according to claim 1, further comprising means for visually overlaying and presenting the guidance information displayed by the visual display device using augmented reality technology.

[0857] "Application example 2 when combining with an emotional engine"

[0858] (Claim 1)

[0859] Equipped with a speech recognition module, and a means for converting the user's voice input into text information,

[0860] A means for identifying a destination based on the aforementioned textual information and determining the user's emotional state using emotion analysis technology,

[0861] A means of creating guidance information by translating it into multiple languages ​​and incorporating feedback that responds to the user's emotional state,

[0862] A means for presenting the aforementioned guidance information to the user via a visual presentation means,

[0863] A system that includes this.

[0864] (Claim 2)

[0865] The system according to claim 1, further comprising means for acquiring real-time traffic information and including an optimal travel route based on traffic conditions, taking into account the emotional state of the user.

[0866] (Claim 3)

[0867] The system according to claim 1, further comprising means for utilizing augmented reality technology to visually overlay and present the guidance information displayed by the visual presentation means according to the user's emotional state. [Explanation of Symbols]

[0868] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Equipped with a speech recognition engine, and a means for converting the user's voice input into text information, A means for identifying a destination based on the aforementioned textual information and creating guidance information by translating it into multiple languages, A means for presenting the aforementioned guidance information to the user via a visual display device, A system that includes this.

2. The system according to claim 1, further comprising means for acquiring real-time traffic information and including in the guidance information an optimal travel route that takes traffic conditions into account.

3. The system according to claim 1, further comprising means for visually overlaying and presenting the guidance information displayed by the visual display device using augmented reality technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A