system
The integrated system addresses language and cultural barriers in travel by using a multimodal AI model for real-time communication and augmented reality, ensuring smooth interaction and information acquisition, thus improving the travel experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Travelers face challenges with smooth communication in different languages and cultural regions, leading to anxieties and misunderstandings that hinder a deep travel experience and increase stress and dissatisfaction.
An integrated system utilizing a multimodal AI model for real-time multilingual communication through speech recognition and machine translation, combined with image analysis and augmented reality to overlay local information onto the physical environment, and seamless data linkage with digital platforms for efficient service provision.
Enables travelers to communicate seamlessly across languages and cultures, intuitively acquire local information, and enjoy a rich travel experience without language or currency barriers, enhancing user satisfaction.
Smart Images

Figure 2026074912000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Many travelers have problems with smooth communication in different languages and cultural regions. There are also anxieties caused by lack of local information and misunderstandings. These problems make it difficult to obtain a deep travel experience and may increase stress and dissatisfaction during travel.
Means for Solving the Problems
[0005] This invention provides an integrated system utilizing a multimodal AI model, enabling real-time multilingual communication for travelers through speech recognition and machine translation. It also incorporates means for analyzing local visual information using image recognition technology and overlaying that information onto the physical environment. Furthermore, seamless data linkage with digital platforms enables the rapid and efficient provision of information and services needed by travelers, thereby realizing a rich travel experience that transcends language barriers.
[0006] "Audio data" refers to information collected as sound that has been digitized, and is a signal that is the target of speech recognition and analysis.
[0007] "Real-time" refers to a format where data processing is performed instantly the moment it occurs, allowing results to be obtained with virtually no delay.
[0008] "Recognition" refers to the process of analyzing input data and understanding its content, and is particularly relevant to processes involving sound and image data.
[0009] "Machine translation" is a technology that uses computer programs to automatically translate text between different languages.
[0010] "Image data" refers to a digital representation of visual information, and is a set of pixels that is the target of image recognition and analysis.
[0011] "Image analysis technology" refers to a series of processes that extract and understand useful information from image data.
[0012] "Information extraction" refers to the process of extracting specific useful components or meanings from data.
[0013] "Overlay display" is a technique that visually combines additional information or graphical elements on top of the original video or image.
[0014] "Communication technology" refers to all technologies, including hardware and software, for transmitting and receiving information.
[0015] "Digital platform" refers to a series of technical foundations for providing data and services online.
Brief Description of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] ]> Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a system that enables smooth communication and information acquisition when travelers visit regions with different languages and cultures. To achieve this, the system's components work together in the following manner.
[0038] First, the "terminal" captures the user's voice data through the microphone. This voice data is then encoded in a digital format and sent to the server. The "server" receives the voice data and converts it into text data using speech recognition technology. The converted text is translated into the specified language by a translation engine and sent back to the terminal. The terminal then converts this translation result into audio format using speech synthesis technology and plays it back to the user, allowing users who do not understand the local language to enjoy conversations in real time.
[0039] Furthermore, the "device" sends image data captured by the user to the "server," which then analyzes the image. Using image analysis technology, the server recognizes, for example, landmarks and the content of signs in tourist areas, and retrieves relevant information from a database. The collected information is sent back to the device, which then uses AR technology to overlay the information onto the camera image. This allows the user to intuitively understand detailed information about the location.
[0040] Furthermore, this system utilizes advanced communication technology to link data with multiple digital platforms. Users can access services using their digital accounts and, for example, make smooth purchases locally using electronic payment platforms. During this process, the server manages the communication data, ensuring that payment procedures are completed securely and efficiently. This provides a comfortable travel experience without language or currency barriers.
[0041] As a concrete example, consider a scenario where a user visits an international festival where many languages are spoken. The user can use the server's translation service to communicate smoothly in multiple languages, and use the image analysis service to obtain and understand information displayed at various booths. As a result, cross-cultural interaction increases, and the enjoyment of the event is significantly enhanced.
[0042] Thus, the present invention aims to provide travelers with a new travel experience that overcomes language barriers and information disparities.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user speaks into the device. The device acquires the voice data through the microphone and encodes it in digital format. This data is then sent to the server via a communication module.
[0046] Step 2:
[0047] The server receives the audio data from the terminal and passes it to the speech recognition engine, which converts the audio data into text format. The converted text is then fed into a translation engine to translate it into multiple languages.
[0048] Step 3:
[0049] The server retrieves the translation result and converts it back into speech data using a speech synthesis engine. This converted speech data is then sent to the terminal.
[0050] Step 4:
[0051] The terminal plays the audio data received from the server through its speaker and outputs the translation result to the user as audio.
[0052] Step 5:
[0053] The user takes a picture with the device's camera. The device uses a communication module to send the image data to the server.
[0054] Step 6:
[0055] The server passes the image data received from the terminal to the image analysis engine, which analyzes its contents. Based on the analyzed information, it searches the database for related data and extracts additional information.
[0056] Step 7:
[0057] The server sends the extracted information back to the terminal. The terminal then displays this information overlaid on the camera image, using augmented reality (AR) technology to provide it to the user.
[0058] Step 8:
[0059] Users visually review the information provided and use it to guide their actions on-site. For example, they can use the payment function to make payments at local stores.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] Travelers often face difficulties communicating smoothly in different linguistic and cultural environments, and in efficiently obtaining local information. Furthermore, language and currency differences create barriers that hinder a smooth travel experience. This project aims to address these challenges and provide travelers with a more comfortable and fulfilling experience.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into a different language, means for acquiring image information and extracting information using image analysis technology, means for overlaying the extracted information onto real-world images using augmented reality technology and displaying it, and means for exchanging information with multiple digital infrastructures using communication means. This makes it possible for travelers to communicate smoothly in a cross-cultural environment, intuitively acquire necessary information, and realize highly convenient travel regardless of language or currency.
[0065] "Audio information" refers to digital data of human voices and sounds captured through audio input devices such as microphones.
[0066] "Real-time recognition" refers to the process of instantly analyzing voice input and converting it into text data without delay.
[0067] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0068] "Image information" refers to digital images of visual data acquired through a camera or other imaging device.
[0069] "Image analysis technology" refers to computational techniques for analyzing image data and extracting specific features of objects, text, and scenes.
[0070] Augmented reality technology is a technique that overlays computer-generated information and images onto the real world's field of view.
[0071] "Communication methods" refer to technical techniques for sending and receiving data between multiple devices or systems.
[0072] "Digital infrastructure" is a general term for the systems and services that constitute information technology infrastructure and support the transmission and processing of data.
[0073] This invention is a system designed to facilitate smooth communication and information acquisition for travelers visiting regions with different language areas and cultural backgrounds. The system utilizes speech recognition, machine translation, image analysis, augmented reality technology, and electronic payment functionality.
[0074] The device captures the user's voice through the microphone and encodes the audio information into a digital format. The device then transmits this data to a server in real time. The server uses speech recognition technology to convert the audio information into text data, and then uses machine translation technology to translate that text into the specified language. Common cloud-based services can be used for the specific speech recognition and translation. The translation results are sent back from the server to the device, which uses speech synthesis technology to convert them into user-friendly speech and play it back. This allows the user to enjoy real-time conversations without being aware of language differences.
[0075] The device also uses its camera to send images captured by the user to a server. The server uses image analysis technology to analyze the content of the images and extract information such as landmarks and signs in tourist areas. This information is used to search for relevant information from multiple digital sources and is sent back to the device. The device then uses augmented reality technology to overlay this information onto the camera image. This makes it easier for the user to visually grasp detailed information about the location.
[0076] Regarding electronic payments, users can access their digital accounts via a terminal if they wish to make a purchase locally. The server exchanges data with the payment platform using secure and efficient communication methods to complete the transaction. As a result, users can enjoy a comfortable shopping experience even in environments with different languages and currencies.
[0077] As a concrete example, consider a scenario where a user visits an international festival where multiple languages are used. The user can communicate smoothly in multiple languages using the translation service provided by the server. They can also use the image analysis service to obtain and understand information displayed at various booths. Examples of prompts include: "I want to understand French using the translation service," "I want to know more about this landmark," and "I want to use electronic payment at the food stall."
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The terminal captures the user's voice through the microphone. The input is the user's spoken voice. This is encoded into a digital format and sent to the server. Specifically, the analog audio signal is converted into digital data and compressed as needed. Through this process, the audio information is delivered to the server as digital data.
[0081] Step 2:
[0082] The server converts received audio data into text data using speech recognition technology. The input is encoded audio data. The server applies a speech recognition algorithm to convert the audio data into text data. Specifically, it extracts the audio waveform data as features and converts them into strings by comparing them with a language model. The output is text in which the user's utterance is expressed as sentences.
[0083] Step 3:
[0084] The server inputs the converted text into a machine translation service and translates it into the specified language. The input is text data generated by speech recognition. The translation engine performs interlingual translation processing, converting the text into different languages. Specifically, it uses translation memory and neural network models to generate text in the target language corresponding to the source text. The output is the translated text data.
[0085] Step 4:
[0086] The server sends the translated text data back to the terminal. The input is the translated text data. The terminal uses speech synthesis technology to convert this text data into speech. Specifically, it uses a speech synthesis engine to convert the text into an audio signal. The output is the generated speech, which is played back to the user through the speaker.
[0087] Step 5:
[0088] The user takes an image using the device's camera, and the device sends the image data to the server. The input is the image taken by the user. The device prepares to transfer the image data to the server in the appropriate format. This process provides the server with the basic data necessary to obtain detailed information about the object.
[0089] Step 6:
[0090] The server analyzes the received image data using image analysis techniques. The input is image data sent by the user. The image analysis algorithm identifies objects within the image and extracts their features. Specifically, it uses a computer vision model to recognize landmarks and characters within the image and extract semantic information. The output is the information resulting from the analysis.
[0091] Step 7:
[0092] The server retrieves relevant information from a database based on the results of image analysis. The input is feature information extracted by image analysis. The server refers to a pre-existing database to identify relevant content. Specifically, it retrieves a dataset containing landmark names and descriptions, related event information, etc. The output is data as relevant information.
[0093] Step 8:
[0094] The server sends the collected relevant information to the terminal, which then displays it using augmented reality technology. The input is the relevant information sent from the server. The terminal uses an AR engine to overlay the information onto the real-world camera image. Specifically, it superimposes the information onto specific points within the image and presents it to the user's vision. The output is the visualized information displayed in AR format.
[0095] Step 9:
[0096] When electronic payment is required, the user initiates an online payment through the terminal. Inputs include the selection of items the user wishes to purchase and payment information. The server communicates with the electronic payment platform after a secure authentication process. Specifically, it transmits payment information using encryption technology and obtains transaction authorization. The output is a confirmation message of the completed payment, displayed on the terminal.
[0097] (Application Example 1)
[0098] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0099] In today's global consumer society, language barriers and a lack of information that travelers from different languages and cultures face when visiting physical stores hinder smooth purchasing and customer service. This often results in travelers not fully understanding product information and having difficulty communicating with store staff. To address this challenge, there is a need for methods that minimize the language and information gap between travelers and stores, providing a smooth and comfortable shopping experience.
[0100] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0101] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into multiple languages, and means for acquiring image information and extracting information using image analysis technology. As a result, travelers can intuitively understand product information using augmented reality technology at physical stores they visit, and conversations with store staff are translated in real time, allowing them to make purchases with peace of mind without feeling any language or information gaps.
[0102] "Real-time recognition of speech information" is a technology that instantly converts a user's speech into a digital format and analyzes it as text information.
[0103] "Machine translation" is the process by which a computer converts natural language into another language, and it is a technology that supports communication between multiple languages.
[0104] "Acquiring image information" is the process of collecting visual data of objects and landscapes using cameras and sensors.
[0105] "Image analysis technology" refers to the technology used to analyze and extract specific information from digitized images.
[0106] Augmented reality technology is a technology that overlays digital information onto the real world environment, providing users with a visually enhanced experience.
[0107] An "information platform" is an online or offline system for exchanging and managing data.
[0108] "Consumer-relevant information" refers to detailed data about products and services that can help customers make purchasing decisions.
[0109] This invention is a system designed to enhance the shopping experience for travelers in physical stores with diverse language and cultural backgrounds. This system operates collaboratively between users, servers, and terminals.
[0110] The server first acquires audio information transmitted from the device. This audio information is captured in real time using the device's microphone. The server converts this audio information into text format using speech recognition technology. Google Cloud Speech-to-Text API is often used for this. The text data is then translated into other languages via a translation engine. Google Cloud Translation API is commonly used for this translation. The converted text is sent back from the server to the device, which then converts the text into speech using speech synthesis technology and provides it to the user. Google Cloud Text-to-Speech API is useful for speech synthesis.
[0111] The device also uses its camera to photograph products in the store and sends the image information to a server. The server uses image analysis technologies such as Amazon Rekognition to recognize the products in the image and retrieves their detailed information from Firebase Cloud Firestore. The recognized information is then passed to the device, which uses augmented reality technologies such as Unity to display the product details in the user's field of view.
[0112] As a concrete example, consider a scenario where a user wants to purchase cosmetics at a store overseas. The user can use their smartphone camera to point at the cosmetics, and augmented reality technology will display their ingredients, price, and other consumers' reviews on the screen. Furthermore, questions and answers with store staff are translated in real time, allowing users to ask questions confidently despite language barriers.
[0113] Examples of prompts for a generative AI model:
[0114] "Translate voice input into a specified language in real time and output it as voice."
[0115] "Acquire product information from the camera and display it using augmented reality (AR)."
[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0117] Step 1:
[0118] The user takes a picture of the products in the store using their smartphone camera. The captured image is saved on the device and then sent to the server. The input is the image information acquired by the camera, and the output is the transmission of the image data to the server.
[0119] Step 2:
[0120] The server analyzes the received image data. During this process, it uses image analysis technologies such as Amazon Rekognition to identify products within the image and extract information associated with those products. The input is the image data sent to the server, and the output is the analyzed product ID and related information.
[0121] Step 3:
[0122] The server searches for and retrieves relevant detailed information (price, ingredients, user reviews, etc.) from Firebase Cloud Firestore based on the analyzed product information. The input is the product ID, and the output is detailed information about the product.
[0123] Step 4:
[0124] The server sends the acquired product information to the terminal. The input is detailed product information, and the output is the transmission of information to the terminal.
[0125] Step 5:
[0126] The device displays received product information in the user's field of view using augmented reality technology. It overlays the information onto the real-world camera screen using an AR framework such as Unity. The input is product information transmitted from the server, and the output is the information visually presented to the user.
[0127] Step 6:
[0128] The user begins a conversation with the store clerk using their smartphone's microphone. The audio data is captured by the device and sent to the server. The input is the voice spoken by the user, and the output is the transmission of audio data to the server.
[0129] Step 7:
[0130] The server converts the received audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data.
[0131] Step 8:
[0132] The server translates the converted text into the specified language using the Google Cloud Translation API. The input is text data, and the output is the translated text in the other language.
[0133] Step 9:
[0134] The translated text is sent from the server to the terminal. The input is the translated text, and the output is the transmission of the translated data to the terminal.
[0135] Step 10:
[0136] The device converts the received translation data into speech using the Google Cloud Text-to-Speech API and outputs it through the speaker. The input is translated text, and the output is audio.
[0137] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0138] This invention is a system designed to support smooth communication and information acquisition for travelers visiting different cultural regions, and in particular, by incorporating an emotion engine, it enables the provision of information that takes the user's emotions into consideration. This system personalizes the user's experience through voice data, image data, and emotion recognition.
[0139] The "terminal" first captures the user's voice data in real time, converts it to a digital format, and sends it to the "server." The server converts the voice data into text using a speech recognition engine, and then translates it into multiple languages using a translation engine. The translation results are converted back into speech by a speech synthesis engine and sent to the terminal. The terminal then plays this speech for the user to support communication in different languages.
[0140] Simultaneously, the "device" acquires emotional data from the user's voice tone and facial expressions. This data is sent to a server and analyzed by an emotion engine. The emotion engine identifies the user's emotions and adjusts the tone and presentation of information based on the results. Emotion-responsive feedback makes the user feel more comfortable.
[0141] Furthermore, the "terminal" transmits image data captured by the user using the camera to the server. The "server" uses image analysis technology to recognize objects and text within the image and extract relevant information. This information is overlaid on the camera image, allowing the user to understand it intuitively.
[0142] Furthermore, users can link their accounts through the system and digital platform to receive personalized experiences based on their emotions. This linking allows users to receive information and services customized according to their emotional data, enriching their travel experience.
[0143] As a concrete example, consider a scenario where a user visits a restaurant in a foreign country and can order without experiencing a language barrier. If the user is feeling anxious or nervous, the emotion engine detects this, and the device suggests reassuring and encouraging messages, as well as interesting local information. As a result, the user can enjoy the experience with peace of mind. Thus, the present invention aims to significantly improve user satisfaction during travel by providing information and support in a way that takes the user's emotions into consideration.
[0144] The following describes the processing flow.
[0145] Step 1:
[0146] To enable voice input, the device captures audio data using its built-in microphone. This data is digitized in real time and sent to the server.
[0147] Step 2:
[0148] The server uses a speech recognition engine to convert the received audio data into text. This text is then passed to a translation engine for translation into multiple languages.
[0149] Step 3:
[0150] The translated text is regenerated as audio data by a speech synthesis engine. The server sends this audio data back to the terminal.
[0151] Step 4:
[0152] The device plays the regenerated audio data to the user through its speaker, providing the translated content in audio.
[0153] Step 5:
[0154] The device simultaneously analyzes the user's voice tone and facial expressions to acquire emotional data. This data is then sent to a server.
[0155] Step 6:
[0156] The server uses an emotion engine to analyze emotional data and identify the user's emotions. This emotional information is used to adjust service delivery.
[0157] Step 7:
[0158] When a user takes a picture, the device acquires the image data and sends it to the server.
[0159] Step 8:
[0160] The server uses image analysis technology to analyze image data and extract relevant information. The extracted results are then sent to the terminal.
[0161] Step 9:
[0162] The device overlays the extracted information onto the real-world camera footage, presenting it visually to the user.
[0163] Step 10:
[0164] Users can view personalized information provided through digital platforms and utilize emotion-based services. In this process, emotional information is used to refine the service content.
[0165] (Example 2)
[0166] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0167] In today's travel environment, when travelers visit different cultural regions, language barriers and cultural gaps make smooth communication difficult. This can make it challenging for travelers to deeply understand different cultures or to enjoy their trips with peace of mind. Furthermore, gathering and understanding local information is not easy. Traditional systems fail to adequately provide emotionally resonant information or enable real-time multilingual communication.
[0168] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0169] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for outputting the translated voice data by speech synthesis. This enables travelers to communicate smoothly with speakers of different languages in real time. Furthermore, by including means for analyzing the user's voice tone and facial expressions to acquire emotional data, and means for analyzing the acquired emotional data to adjust the way information is presented, personalized information provision according to the traveler's emotions is realized, enabling a more user-friendly travel experience. In addition, by including means for acquiring image data and extracting information using image analysis technology, and means for overlaying the extracted information onto real-world images, travelers can intuitively understand local visual information. This allows travelers to understand different cultures more deeply and enjoy their trip with peace of mind.
[0170] "Audio data" refers to a digital representation of sound waveforms, and is used to record human speech and analyze its content.
[0171] "Real-time recognition" refers to a method that processes audio and information instantly on the spot, allowing for immediate results.
[0172] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0173] "Multilingualization" is the process of translating information expressed in one language into multiple different languages and providing it to the public.
[0174] "Speech synthesis" is a technology that imitates human speech by artificially generating voices based on text data.
[0175] "Analyzing voice tone and facial expressions" refers to a method of analyzing collected audio tones and images to infer emotions and intentions.
[0176] "Emotional data" refers to information that indicates a person's emotional state and is used to identify specific emotions from voice and facial expressions.
[0177] "Image data" refers to a collection of visual information expressed in digital format, including graphic information such as photographs and drawings.
[0178] "Image analysis technology" refers to techniques for processing image data and extracting or recognizing useful information such as objects and text from it.
[0179] "Overlaying digital information onto real-world images" means adding digital information to the actual field of view to visually integrate it and improve visibility.
[0180] "Communication technology" refers to technologies for sending and receiving data and exchanging information between multiple devices and platforms.
[0181] "Electronic devices" refer to all fundamental devices that process information, including computers and their peripherals.
[0182] This system is designed to facilitate smooth communication across language barriers and personalize travel experiences for travelers visiting different cultural regions. The following describes a specific implementation of this system.
[0183] First, regarding the processing of audio data, the terminal captures the user's voice in real time. Hardware-wise, this involves using a mobile device with a built-in microphone or a dedicated audio capture device. The captured audio is then transmitted to the server using digital signal processing technology.
[0184] Next, the server uses a speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio data into text. The text is then translated into multiple languages by a translation engine (e.g., a general-purpose translation service). The converted text data is then re-converted into speech using a speech synthesis engine (e.g., general-purpose speech synthesis technology) and sent to the terminal. This allows the user to understand languages other than their native language.
[0185] The device also acquires emotional data from the user's voice tone and facial expressions. It analyzes the user's emotional state in real time using a combination of a microphone and camera, and transmits this data to a server. The server uses an emotion engine to analyze the emotional data and sends instructions to the device to provide the user with appropriate information.
[0186] Furthermore, the terminal transmits image data acquired by the user's camera to a server. The server uses image analysis technology to extract objects and text from the image data and provide the user with useful information. This allows the user to instantly obtain additional information about the images they have taken.
[0187] The system assists users in understanding menus and ordering without difficulty at restaurants in foreign countries. If the user feels anxious, the system suggests encouraging messages and interesting local information to alleviate their anxiety, ensuring reassuring communication. An example of a prompt for the generative AI model is: "The user needs help understanding the menu at the restaurant. If the user's tone indicates anxiety, what positive message should you offer?"
[0188] Thus, the present invention enables the provision of information based on the user's emotions and actively supports the travel experience.
[0189] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0190] Step 1:
[0191] The terminal captures the user's speech in real time using a microphone. The input audio signal is acquired in analog format. The terminal performs digital signal processing to convert the analog audio into a digital signal, generating digital audio data. This digital audio data is output and sent to the next processing step.
[0192] Step 2:
[0193] The server receives audio data transmitted from the terminal and converts it into text using a speech recognition engine. The input is digital audio data, which is converted into text format by speech recognition. The converted text data is output and used as input for translation processing.
[0194] Step 3:
[0195] The server translates the converted text data into multiple languages using a translation engine. Here, the input text data is converted into the specified target language. The translation engine understands the vocabulary and context to perform the optimal translation. The output is the translated text data, which is used for speech synthesis processing.
[0196] Step 4:
[0197] The server converts translated text data into speech using a speech synthesis engine. The input is translated text data, and the speech synthesis engine generates synthesized speech. The output is synthesized speech data (digital speech data), which is then prepared for transmission to the terminal.
[0198] Step 5:
[0199] The device receives audio data transmitted from the server and plays it back through its speaker so that the user can understand it. The input is synthesized speech data, which is then output as a physical audio signal. This output is ultimately delivered to the user, enabling communication in different languages.
[0200] Step 6:
[0201] The device captures the user's voice tone and facial expressions using its camera and microphone to acquire emotion data. The input is the user's current voice tone and facial expression. The device uses an analysis algorithm to identify emotions and outputs them as emotion data.
[0202] Step 7:
[0203] The server receives emotion data sent from the terminal and analyzes it using the emotion engine. The input is emotion data, which the emotion engine analyzes to identify the user's emotional state. The output is the analyzed emotion information, which is used to optimize information delivery.
[0204] Step 8:
[0205] The server selects and outputs information to the user based on the analyzed emotional information. For example, if the user is feeling anxious, it selects information that provides reassurance. The output information is sent to the terminal and displayed or played back.
[0206] (Application Example 2)
[0207] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0208] When travelers use autonomous vehicles, it is necessary to alleviate anxieties caused by language barriers and cultural differences and provide a more comfortable and personalized travel experience. However, current systems have challenges in adequately providing flexible information based on emotions and optimizing the in-vehicle environment.
[0209] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0210] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for analyzing emotions in real time and adjusting the information and tone of voice based on the results. This enables the provision of information tailored to the traveler's emotions and the optimization of the in-vehicle environment.
[0211] "Methods for recognizing audio data in real time" refer to technologies that instantly convert audio into a digital format and understand its content.
[0212] "Means of machine-translating audio data and converting it into multiple languages" refers to an automated process for converting recognized audio content into different languages.
[0213] "Means of acquiring image data and extracting information using image analysis techniques" refers to techniques for extracting specific information from images acquired by cameras or sensors.
[0214] "A means of overlaying extracted information onto images of the real world" refers to a technology for integrating and displaying analyzed information within the actual visual environment.
[0215] "Means of exchanging data with multiple digital platforms using communication technology" refers to technologies for exchanging information between different digital systems.
[0216] "A means of analyzing emotions in real time and adjusting the tone of information and voice based on the results" refers to a technology that instantly grasps the user's emotional state and changes the expression of the information and voice provided according to those emotions.
[0217] "Means of optimizing the in-vehicle environment according to the passengers' emotions" refers to technology that adjusts the music, lighting, and temperature inside the vehicle according to the passengers' emotions to provide a comfortable space.
[0218] The system implementing this invention performs speech recognition, emotion analysis, translation, image analysis, and optimization of the in-vehicle environment in real time. The server recognizes speech data in real time and enables multilingual communication through real-time machine translation. For example, when a user converses with a passenger who speaks a different language, the speech data is instantly translated and output in the target language using speech synthesis technology. The terminal also detects emotions from the user's voice tone and facial expressions and sends emotion data to the server. The emotion engine analyzes this data and adjusts the way information is presented and the tone of voice according to the emotion.
[0219] Furthermore, the device sends image data acquired using its camera to a server, where relevant information is extracted using image analysis technology. This information is overlaid on the real-world scenery and presented intuitively to the user. When a user visits a specific tourist spot, the device can provide relevant historical and cultural information on the spot.
[0220] Furthermore, based on passenger emotional data, the server adjusts and optimizes the in-car environment, such as music and lighting, through the vehicle's environmental control system. For example, if a user is feeling anxious, the system provides a sense of security by selecting calming music and warm lighting. In this way, users can enjoy a more comfortable and personalized travel experience.
[0221] A concrete example of a prompt might be, "Translate the audio data immediately and adjust the tone of voice to match the passenger's emotions before outputting it." This would improve the user experience and support a more comfortable journey.
[0222] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0223] Step 1:
[0224] The device captures the user's voice using a microphone. This audio data is converted to a digital format and sent to the server. The input is the user's raw voice data, and the output is the digitally converted audio data. At this stage, a noise reduction filter is used to improve the audio quality.
[0225] Step 2:
[0226] The server converts the received digital audio data into text using a speech recognition engine. The input is the digitally converted audio data, and the output is text data in which the audio has been replaced with characters. Here, natural language processing techniques are applied to generate accurate textual information from the audio data.
[0227] Step 3:
[0228] The server translates text data generated from speech into different languages using a translation engine. The input is text data obtained by speech recognition, and the output is translated text data. In this step, a generative AI model is used to perform highly accurate translations.
[0229] Step 4:
[0230] The server converts the translated text into speech using a speech synthesis engine. The input is the translated text data, and the output is the synthesized speech data. At this point, the tone and speed of the speech are adjusted to make it sound natural and friendly.
[0231] Step 5:
[0232] The device outputs synthesized speech to the user through its speaker. The input is synthesized speech data, and the output is spoken language audible to the user. In this step, check and adjust whether the volume and sound quality are appropriate.
[0233] Step 6:
[0234] The device acquires emotional data from the user's facial expressions and voice tone and sends it to the server. The input is the user's voice tone and facial expression data, and the output is analyzable emotional data. This process combines camera images and voice analysis technology to determine emotions.
[0235] Step 7:
[0236] The server analyzes emotional data using an emotion engine and returns feedback to the terminal that adjusts the information and tone of voice based on the results. The input is emotional data, and the output is adjusted information and tone of voice. A generative AI model is used to make the adjustments that are most appropriate for the situation.
[0237] Step 8:
[0238] The terminal sends image data captured by the user's camera to the server. The input is raw image data, and the output is the data sent to the server. In this step, the image data is compressed and optimized to improve transmission efficiency.
[0239] Step 9:
[0240] The server uses image analysis technology to recognize objects and characters within images and extract relevant information. The input is image data, and the output is informational data based on the recognized content. A generative AI model is used to improve contextual understanding.
[0241] Step 10:
[0242] The server sends commands to optimize the in-vehicle environment settings based on the analyzed information. The input is recognition and emotion data, and the output is commands to adjust the environment settings. The vehicle's music, lighting, and temperature settings are adjusted based on specific feedback.
[0243] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0244] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0245] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0246] [Second Embodiment]
[0247] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0248] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0249] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0250] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0251] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0252] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0253] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0254] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0255] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0256] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0257] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0258] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0259] This invention is a system that enables smooth communication and information acquisition when travelers visit regions with different languages and cultures. To achieve this, the system's components work together in the following manner.
[0260] First, the "terminal" captures the user's voice data through the microphone. This voice data is then encoded in a digital format and sent to the server. The "server" receives the voice data and converts it into text data using speech recognition technology. The converted text is translated into the specified language by a translation engine and sent back to the terminal. The terminal then converts this translation result into audio format using speech synthesis technology and plays it back to the user, allowing users who do not understand the local language to enjoy conversations in real time.
[0261] Furthermore, the "device" sends image data captured by the user to the "server," which then analyzes the image. Using image analysis technology, the server recognizes, for example, landmarks and the content of signs in tourist areas, and retrieves relevant information from a database. The collected information is sent back to the device, which then uses AR technology to overlay the information onto the camera image. This allows the user to intuitively understand detailed information about the location.
[0262] Furthermore, this system utilizes advanced communication technology to link data with multiple digital platforms. Users can access services using their digital accounts and, for example, make smooth purchases locally using electronic payment platforms. During this process, the server manages the communication data, ensuring that payment procedures are completed securely and efficiently. This provides a comfortable travel experience without language or currency barriers.
[0263] As a concrete example, consider a scenario where a user visits an international festival where many languages are spoken. The user can use the server's translation service to communicate smoothly in multiple languages, and use the image analysis service to obtain and understand information displayed at various booths. As a result, cross-cultural interaction increases, and the enjoyment of the event is significantly enhanced.
[0264] Thus, the present invention aims to provide travelers with a new travel experience that overcomes language barriers and information disparities.
[0265] The following describes the processing flow.
[0266] Step 1:
[0267] The user speaks into the device. The device acquires the voice data through the microphone and encodes it in digital format. This data is then sent to the server via a communication module.
[0268] Step 2:
[0269] The server receives the audio data from the terminal and passes it to the speech recognition engine, which converts the audio data into text format. The converted text is then fed into a translation engine to translate it into multiple languages.
[0270] Step 3:
[0271] The server retrieves the translation result and converts it back into speech data using a speech synthesis engine. This converted speech data is then sent to the terminal.
[0272] Step 4:
[0273] The terminal plays the audio data received from the server through its speaker and outputs the translation result to the user as audio.
[0274] Step 5:
[0275] The user takes a picture with the device's camera. The device uses a communication module to send the image data to the server.
[0276] Step 6:
[0277] The server passes the image data received from the terminal to the image analysis engine, which analyzes its contents. Based on the analyzed information, it searches the database for related data and extracts additional information.
[0278] Step 7:
[0279] The server sends back the extracted information to the terminal. The terminal utilizes AR technology to provide this information to the user for overlay display on the camera video.
[0280] Step 8:
[0281] The user visually confirms the provided information and uses it for actions in the actual local area. For example, use the payment function to make payments at local stores.
[0282] (Example 1)
[0283] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0284] It is difficult for travelers to communicate smoothly in different language and cultural environments and to efficiently obtain local information. Also, barriers due to language and currency differences are factors hindering a smooth travel experience. The aim is to solve these problems and provide a more comfortable and fulfilling experience for travelers.
[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0286] In this invention, the server includes means for real-time recognition of voice information, means for machine translation of the recognized voice information into different languages, means for acquiring image information and extracting information using image analysis technology, means for overlaying and displaying the extracted information on the video of the real world using augmented reality technology, and means for exchanging information with multiple digital bases using communication means. Thereby, it becomes possible for travelers to communicate smoothly in a different cultural environment, intuitively obtain necessary information, and realize a highly convenient travel without being affected by language or currency.
[0287] "Audio information" refers to digital data of human voices and sounds captured through audio input devices such as microphones.
[0288] "Real-time recognition" refers to the process of instantly analyzing voice input and converting it into text data without delay.
[0289] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0290] "Image information" refers to digital images of visual data acquired through a camera or other imaging device.
[0291] "Image analysis technology" refers to computational techniques for analyzing image data and extracting specific features of objects, text, and scenes.
[0292] Augmented reality technology is a technique that overlays computer-generated information and images onto the real world's field of view.
[0293] "Communication methods" refer to technical techniques for sending and receiving data between multiple devices or systems.
[0294] "Digital infrastructure" is a general term for the systems and services that constitute information technology infrastructure and support the transmission and processing of data.
[0295] This invention is a system designed to facilitate smooth communication and information acquisition for travelers visiting regions with different language areas and cultural backgrounds. The system utilizes speech recognition, machine translation, image analysis, augmented reality technology, and electronic payment functionality.
[0296] The device captures the user's voice through the microphone and encodes the audio information into a digital format. The device then transmits this data to a server in real time. The server uses speech recognition technology to convert the audio information into text data, and then uses machine translation technology to translate that text into the specified language. Common cloud-based services can be used for the specific speech recognition and translation. The translation results are sent back from the server to the device, which uses speech synthesis technology to convert them into user-friendly speech and play it back. This allows the user to enjoy real-time conversations without being aware of language differences.
[0297] The device also uses its camera to send images captured by the user to a server. The server uses image analysis technology to analyze the content of the images and extract information such as landmarks and signs in tourist areas. This information is used to search for relevant information from multiple digital sources and is sent back to the device. The device then uses augmented reality technology to overlay this information onto the camera image. This makes it easier for the user to visually grasp detailed information about the location.
[0298] Regarding electronic payments, users can access their digital accounts via a terminal if they wish to make a purchase locally. The server exchanges data with the payment platform using secure and efficient communication methods to complete the transaction. As a result, users can enjoy a comfortable shopping experience even in environments with different languages and currencies.
[0299] As a concrete example, consider a scenario where a user visits an international festival where multiple languages are used. The user can communicate smoothly in multiple languages using the translation service provided by the server. They can also use the image analysis service to obtain and understand information displayed at various booths. Examples of prompts include: "I want to understand French using the translation service," "I want to know more about this landmark," and "I want to use electronic payment at the food stall."
[0300] The flow of the specific process in Example 1 will be described with reference to FIG. 11.
[0301] Step 1:
[0302] The terminal captures the user's voice through the microphone. The input is the voice spoken by the user. This is encoded into a digital format and transmitted to the server. Specifically, the analog voice signal is converted into digital data, and compression processing is performed as necessary. Through this process, the voice information is passed to the server as digital data.
[0303] Step 2:
[0304] The server converts the received voice data into text data using voice recognition technology. The input is the encoded voice data. The server applies a voice recognition algorithm to convert the voice data into text data. As a specific operation, the waveform data of the voice is extracted as a feature amount, and it is compared with a language model and converted into a character string. The output is text in which the content of the user's utterance is expressed as a sentence.
[0305] Step 3:
[0306] The server inputs the converted text into a machine translation service and translates it into the specified language. The input is the text data generated by voice recognition. The translation engine performs translation processing between languages and converts the text into a different language. As a specific operation, using a translation memory or a neural network model, text in the target language corresponding to the original text is generated. The output is the translated text data.
[0307] Step 4:
[0308] The server sends the translated text data back to the terminal. The input is the translated text data. The terminal uses speech synthesis technology to convert this text data into speech. Specifically, it uses a speech synthesis engine to convert the text into an audio signal. The output is the generated speech, which is played back to the user through the speaker.
[0309] Step 5:
[0310] The user takes an image using the device's camera, and the device sends the image data to the server. The input is the image taken by the user. The device prepares to transfer the image data to the server in the appropriate format. This process provides the server with the basic data necessary to obtain detailed information about the object.
[0311] Step 6:
[0312] The server analyzes the received image data using image analysis techniques. The input is image data sent by the user. The image analysis algorithm identifies objects within the image and extracts their features. Specifically, it uses a computer vision model to recognize landmarks and characters within the image and extract semantic information. The output is the information resulting from the analysis.
[0313] Step 7:
[0314] The server retrieves relevant information from a database based on the results of image analysis. The input is feature information extracted by image analysis. The server refers to a pre-existing database to identify relevant content. Specifically, it retrieves a dataset containing landmark names and descriptions, related event information, etc. The output is data as relevant information.
[0315] Step 8:
[0316] The server sends the collected relevant information to the terminal, which then displays it using augmented reality technology. The input is the relevant information sent from the server. The terminal uses an AR engine to overlay the information onto the real-world camera image. Specifically, it superimposes the information onto specific points within the image and presents it to the user's vision. The output is the visualized information displayed in AR format.
[0317] Step 9:
[0318] When electronic payment is required, the user initiates an online payment through the terminal. Inputs include the selection of items the user wishes to purchase and payment information. The server communicates with the electronic payment platform after a secure authentication process. Specifically, it transmits payment information using encryption technology and obtains transaction authorization. The output is a confirmation message of the completed payment, displayed on the terminal.
[0319] (Application Example 1)
[0320] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0321] In today's global consumer society, language barriers and a lack of information that travelers from different languages and cultures face when visiting physical stores hinder smooth purchasing and customer service. This often results in travelers not fully understanding product information and having difficulty communicating with store staff. To address this challenge, there is a need for methods that minimize the language and information gap between travelers and stores, providing a smooth and comfortable shopping experience.
[0322] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0323] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into multiple languages, and means for acquiring image information and extracting information using image analysis technology. As a result, travelers can intuitively understand product information using augmented reality technology at physical stores they visit, and conversations with store staff are translated in real time, allowing them to make purchases with peace of mind without feeling any language or information gaps.
[0324] "Real-time recognition of speech information" is a technology that instantly converts a user's speech into a digital format and analyzes it as text information.
[0325] "Machine translation" is the process by which a computer converts natural language into another language, and it is a technology that supports communication between multiple languages.
[0326] "Acquiring image information" is the process of collecting visual data of objects and landscapes using cameras and sensors.
[0327] "Image analysis technology" refers to the technology used to analyze and extract specific information from digitized images.
[0328] Augmented reality technology is a technology that overlays digital information onto the real world environment, providing users with a visually enhanced experience.
[0329] An "information platform" is an online or offline system for exchanging and managing data.
[0330] "Consumer-relevant information" refers to detailed data about products and services that can help customers make purchasing decisions.
[0331] This invention is a system designed to enhance the shopping experience for travelers in physical stores with diverse language and cultural backgrounds. This system operates collaboratively between users, servers, and terminals.
[0332] The server first acquires audio information transmitted from the device. This audio information is captured in real time using the device's microphone. The server converts this audio information into text format using speech recognition technology. The Google Cloud Speech-to-Text API is often used for this. The text data is then translated into other languages via a translation engine. The Google Cloud Translation API is commonly used for this translation. The converted text is sent back from the server to the device, which then converts the text into speech using speech synthesis technology and provides it to the user. The Google Cloud Text-to-Speech API is useful for speech synthesis.
[0333] The device also uses its camera to photograph products in the store and sends the image information to a server. The server uses image analysis technologies such as Amazon Rekognition to recognize the products in the image and retrieves their detailed information from Firebase Cloud Firestore. The recognized information is then passed to the device, which uses augmented reality technologies such as Unity to display the product details in the user's field of view.
[0334] As a concrete example, consider a scenario where a user wants to purchase cosmetics at a store overseas. The user can use their smartphone camera to point at the cosmetics, and augmented reality technology will display their ingredients, price, and other consumers' reviews on the screen. Furthermore, questions and answers with store staff are translated in real time, allowing users to ask questions confidently despite language barriers.
[0335] Examples of prompts for a generative AI model:
[0336] "Translate voice input into a specified language in real time and output it as voice."
[0337] "Acquire product information from the camera and display it using augmented reality (AR)."
[0338] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0339] Step 1:
[0340] The user takes a picture of the products in the store using their smartphone camera. The captured image is saved on the device and then sent to the server. The input is the image information acquired by the camera, and the output is the transmission of the image data to the server.
[0341] Step 2:
[0342] The server analyzes the received image data. During this process, it uses image analysis technologies such as Amazon Rekognition to identify products within the image and extract information associated with those products. The input is the image data sent to the server, and the output is the analyzed product ID and related information.
[0343] Step 3:
[0344] The server searches for and retrieves relevant detailed information (price, ingredients, user reviews, etc.) from Firebase Cloud Firestore based on the analyzed product information. The input is the product ID, and the output is detailed information about the product.
[0345] Step 4:
[0346] The server sends the acquired product information to the terminal. The input is detailed product information, and the output is the transmission of information to the terminal.
[0347] Step 5:
[0348] The device displays received product information in the user's field of view using augmented reality technology. It overlays the information onto the real-world camera screen using an AR framework such as Unity. The input is product information transmitted from the server, and the output is the information visually presented to the user.
[0349] Step 6:
[0350] The user begins a conversation with the store clerk using their smartphone's microphone. The audio data is captured by the device and sent to the server. The input is the voice spoken by the user, and the output is the transmission of audio data to the server.
[0351] Step 7:
[0352] The server converts the received audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data.
[0353] Step 8:
[0354] The server translates the converted text into the specified language using the Google Cloud Translation API. The input is text data, and the output is the translated text in the other language.
[0355] Step 9:
[0356] The translated text is sent from the server to the terminal. The input is the translated text, and the output is the transmission of the translated data to the terminal.
[0357] Step 10:
[0358] The device converts the received translation data into speech using the Google Cloud Text-to-Speech API and outputs it through the speaker. The input is translated text, and the output is audio.
[0359] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0360] This invention is a system designed to support smooth communication and information acquisition for travelers visiting different cultural regions, and in particular, by incorporating an emotion engine, it enables the provision of information that takes the user's emotions into consideration. This system personalizes the user's experience through voice data, image data, and emotion recognition.
[0361] The "terminal" first captures the user's voice data in real time, converts it to a digital format, and sends it to the "server." The server converts the voice data into text using a speech recognition engine, and then translates it into multiple languages using a translation engine. The translation results are converted back into speech by a speech synthesis engine and sent to the terminal. The terminal then plays this speech for the user to support communication in different languages.
[0362] Simultaneously, the "device" acquires emotional data from the user's voice tone and facial expressions. This data is sent to a server and analyzed by an emotion engine. The emotion engine identifies the user's emotions and adjusts the tone and presentation of information based on the results. Emotion-responsive feedback makes the user feel more comfortable.
[0363] Furthermore, the "terminal" transmits image data captured by the user using the camera to the server. The "server" uses image analysis technology to recognize objects and text within the image and extract relevant information. This information is overlaid on the camera image, allowing the user to understand it intuitively.
[0364] Furthermore, users can link their accounts through the system and digital platform to receive personalized experiences based on their emotions. This linking allows users to receive information and services customized according to their emotional data, enriching their travel experience.
[0365] As a concrete example, consider a scenario where a user visits a restaurant in a foreign country and can order without experiencing a language barrier. If the user is feeling anxious or nervous, the emotion engine detects this, and the device suggests reassuring and encouraging messages, as well as interesting local information. As a result, the user can enjoy the experience with peace of mind. Thus, the present invention aims to significantly improve user satisfaction during travel by providing information and support in a way that takes the user's emotions into consideration.
[0366] The following describes the processing flow.
[0367] Step 1:
[0368] To enable voice input, the device captures audio data using its built-in microphone. This data is digitized in real time and sent to the server.
[0369] Step 2:
[0370] The server uses a speech recognition engine to convert the received audio data into text. This text is then passed to a translation engine for translation into multiple languages.
[0371] Step 3:
[0372] The translated text is regenerated as audio data by a speech synthesis engine. The server sends this audio data back to the terminal.
[0373] Step 4:
[0374] The device plays the regenerated audio data to the user through its speaker, providing the translated content in audio.
[0375] Step 5:
[0376] The device simultaneously analyzes the user's voice tone and facial expressions to acquire emotional data. This data is then sent to a server.
[0377] Step 6:
[0378] The server uses an emotion engine to analyze emotional data and identify the user's emotions. This emotional information is used to adjust service delivery.
[0379] Step 7:
[0380] When a user takes a picture, the device acquires the image data and sends it to the server.
[0381] Step 8:
[0382] The server uses image analysis technology to analyze image data and extract relevant information. The extracted results are then sent to the terminal.
[0383] Step 9:
[0384] The device overlays the extracted information onto the real-world camera footage, presenting it visually to the user.
[0385] Step 10:
[0386] Users can view personalized information provided through digital platforms and utilize emotion-based services. In this process, emotional information is used to refine the service content.
[0387] (Example 2)
[0388] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0389] In today's travel environment, when travelers visit different cultural regions, language barriers and cultural gaps make smooth communication difficult. This can make it challenging for travelers to deeply understand different cultures or to enjoy their trips with peace of mind. Furthermore, gathering and understanding local information is not easy. Traditional systems fail to adequately provide emotionally resonant information or enable real-time multilingual communication.
[0390] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0391] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for outputting the translated voice data by speech synthesis. This enables travelers to communicate smoothly with speakers of different languages in real time. Furthermore, by including means for analyzing the user's voice tone and facial expressions to acquire emotional data, and means for analyzing the acquired emotional data to adjust the way information is presented, personalized information provision according to the traveler's emotions is realized, enabling a more user-friendly travel experience. In addition, by including means for acquiring image data and extracting information using image analysis technology, and means for overlaying the extracted information onto real-world images, travelers can intuitively understand local visual information. This allows travelers to understand different cultures more deeply and enjoy their trip with peace of mind.
[0392] "Audio data" refers to a digital representation of sound waveforms, and is used to record human speech and analyze its content.
[0393] "Real-time recognition" refers to a method that processes audio and information instantly on the spot, allowing for immediate results.
[0394] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0395] "Multilingualization" is the process of translating information expressed in one language into multiple different languages and providing it to the public.
[0396] "Speech synthesis" is a technology that imitates human speech by artificially generating voices based on text data.
[0397] "Analyzing voice tone and facial expressions" refers to a method of analyzing collected audio tones and images to infer emotions and intentions.
[0398] "Emotional data" refers to information that indicates a person's emotional state and is used to identify specific emotions from voice and facial expressions.
[0399] "Image data" refers to a collection of visual information expressed in digital format, including graphic information such as photographs and drawings.
[0400] "Image analysis technology" refers to techniques for processing image data and extracting or recognizing useful information such as objects and text from it.
[0401] "Overlaying digital information onto real-world images" means adding digital information to the actual field of view to visually integrate it and improve visibility.
[0402] "Communication technology" refers to technologies for sending and receiving data and exchanging information between multiple devices and platforms.
[0403] "Electronic devices" refer to all fundamental devices that process information, including computers and their peripherals.
[0404] This system is designed to facilitate smooth communication across language barriers and personalize travel experiences for travelers visiting different cultural regions. The following describes a specific implementation of this system.
[0405] First, regarding the processing of audio data, the terminal captures the user's voice in real time. Hardware-wise, this involves using a mobile device with a built-in microphone or a dedicated audio capture device. The captured audio is then transmitted to the server using digital signal processing technology.
[0406] Next, the server uses a speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio data into text. The text is then translated into multiple languages by a translation engine (e.g., a general-purpose translation service). The converted text data is then re-converted into speech using a speech synthesis engine (e.g., general-purpose speech synthesis technology) and sent to the terminal. This allows the user to understand languages other than their native language.
[0407] The device also acquires emotional data from the user's voice tone and facial expressions. It analyzes the user's emotional state in real time using a combination of a microphone and camera, and transmits this data to a server. The server uses an emotion engine to analyze the emotional data and sends instructions to the device to provide the user with appropriate information.
[0408] Furthermore, the terminal transmits image data acquired by the user's camera to a server. The server uses image analysis technology to extract objects and text from the image data and provide the user with useful information. This allows the user to instantly obtain additional information about the images they have taken.
[0409] The system assists users in understanding menus and ordering without difficulty at restaurants in foreign countries. If the user feels anxious, the system suggests encouraging messages and interesting local information to alleviate their anxiety, ensuring reassuring communication. An example of a prompt for the generative AI model is: "The user needs help understanding the menu at the restaurant. If the user's tone indicates anxiety, what positive message should you offer?"
[0410] Thus, the present invention enables the provision of information based on the user's emotions and actively supports the travel experience.
[0411] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0412] Step 1:
[0413] The terminal captures the user's speech in real time using a microphone. The input audio signal is acquired in analog format. The terminal performs digital signal processing to convert the analog audio into a digital signal, generating digital audio data. This digital audio data is output and sent to the next processing step.
[0414] Step 2:
[0415] The server receives audio data transmitted from the terminal and converts it into text using a speech recognition engine. The input is digital audio data, which is converted into text format by speech recognition. The converted text data is output and used as input for translation processing.
[0416] Step 3:
[0417] The server translates the converted text data into multiple languages using a translation engine. Here, the input text data is converted into the specified target language. The translation engine understands the vocabulary and context to perform the optimal translation. The output is the translated text data, which is used for speech synthesis processing.
[0418] Step 4:
[0419] The server converts translated text data into speech using a speech synthesis engine. The input is translated text data, and the speech synthesis engine generates synthesized speech. The output is synthesized speech data (digital speech data), which is then prepared for transmission to the terminal.
[0420] Step 5:
[0421] The device receives audio data transmitted from the server and plays it back through its speaker so that the user can understand it. The input is synthesized speech data, which is then output as a physical audio signal. This output is ultimately delivered to the user, enabling communication in different languages.
[0422] Step 6:
[0423] The device captures the user's voice tone and facial expressions using its camera and microphone to acquire emotion data. The input is the user's current voice tone and facial expression. The device uses an analysis algorithm to identify emotions and outputs them as emotion data.
[0424] Step 7:
[0425] The server receives emotion data sent from the terminal and analyzes it using the emotion engine. The input is emotion data, which the emotion engine analyzes to identify the user's emotional state. The output is the analyzed emotion information, which is used to optimize information delivery.
[0426] Step 8:
[0427] The server selects and outputs information to the user based on the analyzed emotional information. For example, if the user is feeling anxious, it selects information that provides reassurance. The output information is sent to the terminal and displayed or played back.
[0428] (Application Example 2)
[0429] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0430] When travelers use autonomous vehicles, it is necessary to alleviate anxieties caused by language barriers and cultural differences and provide a more comfortable and personalized travel experience. However, current systems have challenges in adequately providing flexible information based on emotions and optimizing the in-vehicle environment.
[0431] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0432] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for analyzing emotions in real time and adjusting the information and tone of voice based on the results. This enables the provision of information tailored to the traveler's emotions and the optimization of the in-vehicle environment.
[0433] "Methods for recognizing audio data in real time" refer to technologies that instantly convert audio into a digital format and understand its content.
[0434] "Means of machine-translating audio data and converting it into multiple languages" refers to an automated process for converting recognized audio content into different languages.
[0435] "Means of acquiring image data and extracting information using image analysis techniques" refers to techniques for extracting specific information from images acquired by cameras or sensors.
[0436] "A means of overlaying extracted information onto images of the real world" refers to a technology for integrating and displaying analyzed information within the actual visual environment.
[0437] "Means of exchanging data with multiple digital platforms using communication technology" refers to technologies for exchanging information between different digital systems.
[0438] "A means of analyzing emotions in real time and adjusting the tone of information and voice based on the results" refers to a technology that instantly grasps the user's emotional state and changes the expression of the information and voice provided according to those emotions.
[0439] "Means of optimizing the in-vehicle environment according to the passengers' emotions" refers to technology that adjusts the music, lighting, and temperature inside the vehicle according to the passengers' emotions to provide a comfortable space.
[0440] The system implementing this invention performs speech recognition, emotion analysis, translation, image analysis, and optimization of the in-vehicle environment in real time. The server recognizes speech data in real time and enables multilingual communication through real-time machine translation. For example, when a user converses with a passenger who speaks a different language, the speech data is instantly translated and output in the target language using speech synthesis technology. The terminal also detects emotions from the user's voice tone and facial expressions and sends emotion data to the server. The emotion engine analyzes this data and adjusts the way information is presented and the tone of voice according to the emotion.
[0441] Furthermore, the device sends image data acquired using its camera to a server, where relevant information is extracted using image analysis technology. This information is overlaid on the real-world scenery and presented intuitively to the user. When a user visits a specific tourist spot, the device can provide relevant historical and cultural information on the spot.
[0442] Furthermore, based on passenger emotional data, the server adjusts and optimizes the in-car environment, such as music and lighting, through the vehicle's environmental control system. For example, if a user is feeling anxious, the system provides a sense of security by selecting calming music and warm lighting. In this way, users can enjoy a more comfortable and personalized travel experience.
[0443] A concrete example of a prompt might be, "Translate the audio data immediately and adjust the tone of voice to match the passenger's emotions before outputting it." This would improve the user experience and support a more comfortable journey.
[0444] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0445] Step 1:
[0446] The device captures the user's voice using a microphone. This audio data is converted to a digital format and sent to the server. The input is the user's raw voice data, and the output is the digitally converted audio data. At this stage, a noise reduction filter is used to improve the audio quality.
[0447] Step 2:
[0448] The server converts the received digital audio data into text using a speech recognition engine. The input is the digitally converted audio data, and the output is text data in which the audio has been replaced with characters. Here, natural language processing techniques are applied to generate accurate textual information from the audio data.
[0449] Step 3:
[0450] The server translates text data generated from speech into different languages using a translation engine. The input is text data obtained by speech recognition, and the output is translated text data. In this step, a generative AI model is used to perform highly accurate translations.
[0451] Step 4:
[0452] The server converts the translated text into speech using a speech synthesis engine. The input is the translated text data, and the output is the synthesized speech data. At this point, the tone and speed of the speech are adjusted to make it sound natural and friendly.
[0453] Step 5:
[0454] The device outputs synthesized speech to the user through its speaker. The input is synthesized speech data, and the output is spoken language audible to the user. In this step, check and adjust whether the volume and sound quality are appropriate.
[0455] Step 6:
[0456] The device acquires emotional data from the user's facial expressions and voice tone and sends it to the server. The input is the user's voice tone and facial expression data, and the output is analyzable emotional data. This process combines camera images and voice analysis technology to determine emotions.
[0457] Step 7:
[0458] The server analyzes emotional data using an emotion engine and returns feedback to the terminal that adjusts the information and tone of voice based on the results. The input is emotional data, and the output is adjusted information and tone of voice. A generative AI model is used to make the adjustments that are most appropriate for the situation.
[0459] Step 8:
[0460] The terminal sends image data captured by the user's camera to the server. The input is raw image data, and the output is the data sent to the server. In this step, the image data is compressed and optimized to improve transmission efficiency.
[0461] Step 9:
[0462] The server uses image analysis technology to recognize objects and characters within images and extract relevant information. The input is image data, and the output is informational data based on the recognized content. A generative AI model is used to improve contextual understanding.
[0463] Step 10:
[0464] The server sends commands to optimize the in-vehicle environment settings based on the analyzed information. The input is recognition and emotion data, and the output is commands to adjust the environment settings. The vehicle's music, lighting, and temperature settings are adjusted based on specific feedback.
[0465] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0466] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0467] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0468] [Third Embodiment]
[0469] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0470] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0471] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0472] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0473] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0474] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0475] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0476] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0477] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0478] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0479] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0480] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0481] This invention is a system that enables smooth communication and information acquisition when travelers visit regions with different languages and cultures. To achieve this, the system's components work together in the following manner.
[0482] First, the "terminal" captures the user's voice data through the microphone. This voice data is then encoded in a digital format and sent to the server. The "server" receives the voice data and converts it into text data using speech recognition technology. The converted text is translated into the specified language by a translation engine and sent back to the terminal. The terminal then converts this translation result into audio format using speech synthesis technology and plays it back to the user, allowing users who do not understand the local language to enjoy conversations in real time.
[0483] Furthermore, the "device" sends image data captured by the user to the "server," which then analyzes the image. Using image analysis technology, the server recognizes, for example, landmarks and the content of signs in tourist areas, and retrieves relevant information from a database. The collected information is sent back to the device, which then uses AR technology to overlay the information onto the camera image. This allows the user to intuitively understand detailed information about the location.
[0484] Furthermore, this system utilizes advanced communication technology to link data with multiple digital platforms. Users can access services using their digital accounts and, for example, make smooth purchases locally using electronic payment platforms. During this process, the server manages the communication data, ensuring that payment procedures are completed securely and efficiently. This provides a comfortable travel experience without language or currency barriers.
[0485] As a concrete example, consider a scenario where a user visits an international festival where many languages are spoken. The user can use the server's translation service to communicate smoothly in multiple languages, and use the image analysis service to obtain and understand information displayed at various booths. As a result, cross-cultural interaction increases, and the enjoyment of the event is significantly enhanced.
[0486] Thus, the present invention aims to provide travelers with a new travel experience that overcomes language barriers and information disparities.
[0487] The following describes the processing flow.
[0488] Step 1:
[0489] The user speaks into the device. The device acquires the voice data through the microphone and encodes it in digital format. This data is then sent to the server via a communication module.
[0490] Step 2:
[0491] The server receives the audio data from the terminal and passes it to the speech recognition engine, which converts the audio data into text format. The converted text is then fed into a translation engine to translate it into multiple languages.
[0492] Step 3:
[0493] The server retrieves the translation result and converts it back into speech data using a speech synthesis engine. This converted speech data is then sent to the terminal.
[0494] Step 4:
[0495] The terminal plays the audio data received from the server through its speaker and outputs the translation result to the user as audio.
[0496] Step 5:
[0497] The user takes a picture with the device's camera. The device uses a communication module to send the image data to the server.
[0498] Step 6:
[0499] The server passes the image data received from the terminal to the image analysis engine, which analyzes its contents. Based on the analyzed information, it searches the database for related data and extracts additional information.
[0500] Step 7:
[0501] The server sends the extracted information back to the terminal. The terminal then displays this information overlaid on the camera image, using augmented reality (AR) technology to provide it to the user.
[0502] Step 8:
[0503] Users visually review the information provided and use it to guide their actions on-site. For example, they can use the payment function to make payments at local stores.
[0504] (Example 1)
[0505] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0506] Travelers often face difficulties communicating smoothly in different linguistic and cultural environments, and in efficiently obtaining local information. Furthermore, language and currency differences create barriers that hinder a smooth travel experience. This project aims to address these challenges and provide travelers with a more comfortable and fulfilling experience.
[0507] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0508] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into a different language, means for acquiring image information and extracting information using image analysis technology, means for overlaying the extracted information onto real-world images using augmented reality technology and displaying it, and means for exchanging information with multiple digital infrastructures using communication means. This makes it possible for travelers to communicate smoothly in a cross-cultural environment, intuitively acquire necessary information, and realize highly convenient travel regardless of language or currency.
[0509] "Audio information" refers to digital data of human voices and sounds captured through audio input devices such as microphones.
[0510] "Real-time recognition" refers to the process of instantly analyzing voice input and converting it into text data without delay.
[0511] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0512] "Image information" refers to digital images of visual data acquired through a camera or other imaging device.
[0513] "Image analysis technology" refers to computational techniques for analyzing image data and extracting specific features of objects, text, and scenes.
[0514] Augmented reality technology is a technique that overlays computer-generated information and images onto the real world's field of view.
[0515] "Communication methods" refer to technical techniques for sending and receiving data between multiple devices or systems.
[0516] "Digital infrastructure" is a general term for the systems and services that constitute information technology infrastructure and support the transmission and processing of data.
[0517] This invention is a system designed to facilitate smooth communication and information acquisition for travelers visiting regions with different language areas and cultural backgrounds. The system utilizes speech recognition, machine translation, image analysis, augmented reality technology, and electronic payment functionality.
[0518] The device captures the user's voice through the microphone and encodes the audio information into a digital format. The device then transmits this data to a server in real time. The server uses speech recognition technology to convert the audio information into text data, and then uses machine translation technology to translate that text into the specified language. Common cloud-based services can be used for the specific speech recognition and translation. The translation results are sent back from the server to the device, which uses speech synthesis technology to convert them into user-friendly speech and play it back. This allows the user to enjoy real-time conversations without being aware of language differences.
[0519] The device also uses its camera to send images captured by the user to a server. The server uses image analysis technology to analyze the content of the images and extract information such as landmarks and signs in tourist areas. This information is used to search for relevant information from multiple digital sources and is sent back to the device. The device then uses augmented reality technology to overlay this information onto the camera image. This makes it easier for the user to visually grasp detailed information about the location.
[0520] Regarding electronic payments, users can access their digital accounts via a terminal if they wish to make a purchase locally. The server exchanges data with the payment platform using secure and efficient communication methods to complete the transaction. As a result, users can enjoy a comfortable shopping experience even in environments with different languages and currencies.
[0521] As a concrete example, consider a scenario where a user visits an international festival where multiple languages are used. The user can communicate smoothly in multiple languages using the translation service provided by the server. They can also use the image analysis service to obtain and understand information displayed at various booths. Examples of prompts include: "I want to understand French using the translation service," "I want to know more about this landmark," and "I want to use electronic payment at the food stall."
[0522] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0523] Step 1:
[0524] The terminal captures the user's voice through the microphone. The input is the user's spoken voice. This is encoded into a digital format and sent to the server. Specifically, the analog audio signal is converted into digital data and compressed as needed. Through this process, the audio information is delivered to the server as digital data.
[0525] Step 2:
[0526] The server converts received audio data into text data using speech recognition technology. The input is encoded audio data. The server applies a speech recognition algorithm to convert the audio data into text data. Specifically, it extracts the audio waveform data as features and converts them into strings by comparing them with a language model. The output is text in which the user's utterance is expressed as sentences.
[0527] Step 3:
[0528] The server inputs the converted text into a machine translation service and translates it into the specified language. The input is text data generated by speech recognition. The translation engine performs interlingual translation processing, converting the text into different languages. Specifically, it uses translation memory and neural network models to generate text in the target language corresponding to the source text. The output is the translated text data.
[0529] Step 4:
[0530] The server sends the translated text data back to the terminal. The input is the translated text data. The terminal uses speech synthesis technology to convert this text data into speech. Specifically, it uses a speech synthesis engine to convert the text into an audio signal. The output is the generated speech, which is played back to the user through the speaker.
[0531] Step 5:
[0532] The user takes an image using the device's camera, and the device sends the image data to the server. The input is the image taken by the user. The device prepares to transfer the image data to the server in the appropriate format. This process provides the server with the basic data necessary to obtain detailed information about the object.
[0533] Step 6:
[0534] The server analyzes the received image data using image analysis techniques. The input is image data sent by the user. The image analysis algorithm identifies objects within the image and extracts their features. Specifically, it uses a computer vision model to recognize landmarks and characters within the image and extract semantic information. The output is the information resulting from the analysis.
[0535] Step 7:
[0536] The server retrieves relevant information from a database based on the results of image analysis. The input is feature information extracted by image analysis. The server refers to a pre-existing database to identify relevant content. Specifically, it retrieves a dataset containing landmark names and descriptions, related event information, etc. The output is data as relevant information.
[0537] Step 8:
[0538] The server sends the collected relevant information to the terminal, which then displays it using augmented reality technology. The input is the relevant information sent from the server. The terminal uses an AR engine to overlay the information onto the real-world camera image. Specifically, it superimposes the information onto specific points within the image and presents it to the user's vision. The output is the visualized information displayed in AR format.
[0539] Step 9:
[0540] When electronic payment is required, the user initiates an online payment through the terminal. Inputs include the selection of items the user wishes to purchase and payment information. The server communicates with the electronic payment platform after a secure authentication process. Specifically, it transmits payment information using encryption technology and obtains transaction authorization. The output is a confirmation message of the completed payment, displayed on the terminal.
[0541] (Application Example 1)
[0542] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0543] In today's global consumer society, language barriers and a lack of information that travelers from different languages and cultures face when visiting physical stores hinder smooth purchasing and customer service. This often results in travelers not fully understanding product information and having difficulty communicating with store staff. To address this challenge, there is a need for methods that minimize the language and information gap between travelers and stores, providing a smooth and comfortable shopping experience.
[0544] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0545] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into multiple languages, and means for acquiring image information and extracting information using image analysis technology. As a result, travelers can intuitively understand product information using augmented reality technology at physical stores they visit, and conversations with store staff are translated in real time, allowing them to make purchases with peace of mind without feeling any language or information gaps.
[0546] "Real-time recognition of speech information" is a technology that instantly converts a user's speech into a digital format and analyzes it as text information.
[0547] "Machine translation" is the process by which a computer converts natural language into another language, and it is a technology that supports communication between multiple languages.
[0548] "Acquiring image information" is the process of collecting visual data of objects and landscapes using cameras and sensors.
[0549] "Image analysis technology" refers to the technology used to analyze and extract specific information from digitized images.
[0550] Augmented reality technology is a technology that overlays digital information onto the real world environment, providing users with a visually enhanced experience.
[0551] An "information platform" is an online or offline system for exchanging and managing data.
[0552] "Consumer-relevant information" refers to detailed data about products and services that can help customers make purchasing decisions.
[0553] This invention is a system designed to enhance the shopping experience for travelers in physical stores with diverse language and cultural backgrounds. This system operates collaboratively between users, servers, and terminals.
[0554] The server first acquires audio information transmitted from the device. This audio information is captured in real time using the device's microphone. The server converts this audio information into text format using speech recognition technology. The Google Cloud Speech-to-Text API is often used for this. The text data is then translated into other languages via a translation engine. The Google Cloud Translation API is commonly used for this translation. The converted text is sent back from the server to the device, which then converts the text into speech using speech synthesis technology and provides it to the user. The Google Cloud Text-to-Speech API is useful for speech synthesis.
[0555] The device also uses its camera to photograph products in the store and sends the image information to a server. The server uses image analysis technologies such as Amazon Rekognition to recognize the products in the image and retrieves their detailed information from Firebase Cloud Firestore. The recognized information is then passed to the device, which uses augmented reality technologies such as Unity to display the product details in the user's field of view.
[0556] As a concrete example, consider a scenario where a user wants to purchase cosmetics at a store overseas. The user can use their smartphone camera to point at the cosmetics, and augmented reality technology will display their ingredients, price, and other consumers' reviews on the screen. Furthermore, questions and answers with store staff are translated in real time, allowing users to ask questions confidently despite language barriers.
[0557] Examples of prompts for a generative AI model:
[0558] "Translate voice input into a specified language in real time and output it as voice."
[0559] "Acquire product information from the camera and display it using augmented reality (AR)."
[0560] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0561] Step 1:
[0562] The user takes a picture of the products in the store using their smartphone camera. The captured image is saved on the device and then sent to the server. The input is the image information acquired by the camera, and the output is the transmission of the image data to the server.
[0563] Step 2:
[0564] The server analyzes the received image data. During this process, it uses image analysis technologies such as Amazon Rekognition to identify products within the image and extract information associated with those products. The input is the image data sent to the server, and the output is the analyzed product ID and related information.
[0565] Step 3:
[0566] The server searches for and retrieves relevant detailed information (price, ingredients, user reviews, etc.) from Firebase Cloud Firestore based on the analyzed product information. The input is the product ID, and the output is detailed information about the product.
[0567] Step 4:
[0568] The server sends the acquired product information to the terminal. The input is detailed product information, and the output is the transmission of information to the terminal.
[0569] Step 5:
[0570] The device displays received product information in the user's field of view using augmented reality technology. It overlays the information onto the real-world camera screen using an AR framework such as Unity. The input is product information transmitted from the server, and the output is the information visually presented to the user.
[0571] Step 6:
[0572] The user begins a conversation with the store clerk using their smartphone's microphone. The audio data is captured by the device and sent to the server. The input is the voice spoken by the user, and the output is the transmission of audio data to the server.
[0573] Step 7:
[0574] The server converts the received audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data.
[0575] Step 8:
[0576] The server translates the converted text into the specified language using the Google Cloud Translation API. The input is text data, and the output is the translated text in the other language.
[0577] Step 9:
[0578] The translated text is sent from the server to the terminal. The input is the translated text, and the output is the transmission of the translated data to the terminal.
[0579] Step 10:
[0580] The device converts the received translation data into speech using the Google Cloud Text-to-Speech API and outputs it through the speaker. The input is translated text, and the output is audio.
[0581] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0582] This invention is a system designed to support smooth communication and information acquisition for travelers visiting different cultural regions, and in particular, by incorporating an emotion engine, it enables the provision of information that takes the user's emotions into consideration. This system personalizes the user's experience through voice data, image data, and emotion recognition.
[0583] The "terminal" first captures the user's voice data in real time, converts it to a digital format, and sends it to the "server." The server converts the voice data into text using a speech recognition engine, and then translates it into multiple languages using a translation engine. The translation results are converted back into speech by a speech synthesis engine and sent to the terminal. The terminal then plays this speech for the user to support communication in different languages.
[0584] Simultaneously, the "device" acquires emotional data from the user's voice tone and facial expressions. This data is sent to a server and analyzed by an emotion engine. The emotion engine identifies the user's emotions and adjusts the tone and presentation of information based on the results. Emotion-responsive feedback makes the user feel more comfortable.
[0585] Furthermore, the "terminal" transmits image data captured by the user using the camera to the server. The "server" uses image analysis technology to recognize objects and text within the image and extract relevant information. This information is overlaid on the camera image, allowing the user to understand it intuitively.
[0586] Furthermore, users can link their accounts through the system and digital platform to receive personalized experiences based on their emotions. This linking allows users to receive information and services customized according to their emotional data, enriching their travel experience.
[0587] As a concrete example, consider a scenario where a user visits a restaurant in a foreign country and can order without experiencing a language barrier. If the user is feeling anxious or nervous, the emotion engine detects this, and the device suggests reassuring and encouraging messages, as well as interesting local information. As a result, the user can enjoy the experience with peace of mind. Thus, the present invention aims to significantly improve user satisfaction during travel by providing information and support in a way that takes the user's emotions into consideration.
[0588] The following describes the processing flow.
[0589] Step 1:
[0590] To enable voice input, the device captures audio data using its built-in microphone. This data is digitized in real time and sent to the server.
[0591] Step 2:
[0592] The server uses a speech recognition engine to convert the received audio data into text. This text is then passed to a translation engine for translation into multiple languages.
[0593] Step 3:
[0594] The translated text is regenerated as audio data by a speech synthesis engine. The server sends this audio data back to the terminal.
[0595] Step 4:
[0596] The device plays the regenerated audio data to the user through its speaker, providing the translated content in audio.
[0597] Step 5:
[0598] The device simultaneously analyzes the user's voice tone and facial expressions to acquire emotional data. This data is then sent to a server.
[0599] Step 6:
[0600] The server uses an emotion engine to analyze emotional data and identify the user's emotions. This emotional information is used to adjust service delivery.
[0601] Step 7:
[0602] When a user takes a picture, the device acquires the image data and sends it to the server.
[0603] Step 8:
[0604] The server uses image analysis technology to analyze image data and extract relevant information. The extracted results are then sent to the terminal.
[0605] Step 9:
[0606] The device overlays the extracted information onto the real-world camera footage, presenting it visually to the user.
[0607] Step 10:
[0608] Users can view personalized information provided through digital platforms and utilize emotion-based services. In this process, emotional information is used to refine the service content.
[0609] (Example 2)
[0610] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0611] In today's travel environment, when travelers visit different cultural regions, language barriers and cultural gaps make smooth communication difficult. This can make it challenging for travelers to deeply understand different cultures or to enjoy their trips with peace of mind. Furthermore, gathering and understanding local information is not easy. Traditional systems fail to adequately provide emotionally resonant information or enable real-time multilingual communication.
[0612] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0613] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for outputting the translated voice data by speech synthesis. This enables travelers to communicate smoothly with speakers of different languages in real time. Furthermore, by including means for analyzing the user's voice tone and facial expressions to acquire emotional data, and means for analyzing the acquired emotional data to adjust the way information is presented, personalized information provision according to the traveler's emotions is realized, enabling a more user-friendly travel experience. In addition, by including means for acquiring image data and extracting information using image analysis technology, and means for overlaying the extracted information onto real-world images, travelers can intuitively understand local visual information. This allows travelers to understand different cultures more deeply and enjoy their trip with peace of mind.
[0614] "Audio data" refers to a digital representation of sound waveforms, and is used to record human speech and analyze its content.
[0615] "Real-time recognition" refers to a method that processes audio and information instantly on the spot, allowing for immediate results.
[0616] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0617] "Multilingualization" is the process of translating information expressed in one language into multiple different languages and providing it to the public.
[0618] "Speech synthesis" is a technology that imitates human speech by artificially generating voices based on text data.
[0619] "Analyzing voice tone and facial expressions" refers to a method of analyzing collected audio tones and images to infer emotions and intentions.
[0620] "Emotional data" refers to information that indicates a person's emotional state and is used to identify specific emotions from voice and facial expressions.
[0621] "Image data" refers to a collection of visual information expressed in digital format, including graphic information such as photographs and drawings.
[0622] "Image analysis technology" refers to techniques for processing image data and extracting or recognizing useful information such as objects and text from it.
[0623] "Overlaying digital information onto real-world images" means adding digital information to the actual field of view to visually integrate it and improve visibility.
[0624] "Communication technology" refers to technologies for sending and receiving data and exchanging information between multiple devices and platforms.
[0625] "Electronic devices" refer to all fundamental devices that process information, including computers and their peripherals.
[0626] This system is designed to facilitate smooth communication across language barriers and personalize travel experiences for travelers visiting different cultural regions. The following describes a specific implementation of this system.
[0627] First, regarding the processing of audio data, the terminal captures the user's voice in real time. Hardware-wise, this involves using a mobile device with a built-in microphone or a dedicated audio capture device. The captured audio is then transmitted to the server using digital signal processing technology.
[0628] Next, the server uses a speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio data into text. The text is then translated into multiple languages by a translation engine (e.g., a general-purpose translation service). The converted text data is then re-converted into speech using a speech synthesis engine (e.g., general-purpose speech synthesis technology) and sent to the terminal. This allows the user to understand languages other than their native language.
[0629] The device also acquires emotional data from the user's voice tone and facial expressions. It analyzes the user's emotional state in real time using a combination of a microphone and camera, and transmits this data to a server. The server uses an emotion engine to analyze the emotional data and sends instructions to the device to provide the user with appropriate information.
[0630] Furthermore, the terminal transmits image data acquired by the user's camera to a server. The server uses image analysis technology to extract objects and text from the image data and provide the user with useful information. This allows the user to instantly obtain additional information about the images they have taken.
[0631] The system assists users in understanding menus and ordering without difficulty at restaurants in foreign countries. If the user feels anxious, the system suggests encouraging messages and interesting local information to alleviate their anxiety, ensuring reassuring communication. An example of a prompt for the generative AI model is: "The user needs help understanding the menu at the restaurant. If the user's tone indicates anxiety, what positive message should you offer?"
[0632] Thus, the present invention enables the provision of information based on the user's emotions and actively supports the travel experience.
[0633] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0634] Step 1:
[0635] The terminal captures the user's speech in real time using a microphone. The input audio signal is acquired in analog format. The terminal performs digital signal processing to convert the analog audio into a digital signal, generating digital audio data. This digital audio data is output and sent to the next processing step.
[0636] Step 2:
[0637] The server receives audio data transmitted from the terminal and converts it into text using a speech recognition engine. The input is digital audio data, which is converted into text format by speech recognition. The converted text data is output and used as input for translation processing.
[0638] Step 3:
[0639] The server translates the converted text data into multiple languages using a translation engine. Here, the input text data is converted into the specified target language. The translation engine understands the vocabulary and context to perform the optimal translation. The output is the translated text data, which is used for speech synthesis processing.
[0640] Step 4:
[0641] The server converts translated text data into speech using a speech synthesis engine. The input is translated text data, and the speech synthesis engine generates synthesized speech. The output is synthesized speech data (digital speech data), which is then prepared for transmission to the terminal.
[0642] Step 5:
[0643] The device receives audio data transmitted from the server and plays it back through its speaker so that the user can understand it. The input is synthesized speech data, which is then output as a physical audio signal. This output is ultimately delivered to the user, enabling communication in different languages.
[0644] Step 6:
[0645] The device captures the user's voice tone and facial expressions using its camera and microphone to acquire emotion data. The input is the user's current voice tone and facial expression. The device uses an analysis algorithm to identify emotions and outputs them as emotion data.
[0646] Step 7:
[0647] The server receives emotion data sent from the terminal and analyzes it using the emotion engine. The input is emotion data, which the emotion engine analyzes to identify the user's emotional state. The output is the analyzed emotion information, which is used to optimize information delivery.
[0648] Step 8:
[0649] The server selects and outputs information to the user based on the analyzed emotional information. For example, if the user is feeling anxious, it selects information that provides reassurance. The output information is sent to the terminal and displayed or played back.
[0650] (Application Example 2)
[0651] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0652] When travelers use autonomous vehicles, it is necessary to alleviate anxieties caused by language barriers and cultural differences and provide a more comfortable and personalized travel experience. However, current systems have challenges in adequately providing flexible information based on emotions and optimizing the in-vehicle environment.
[0653] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0654] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for analyzing emotions in real time and adjusting the information and tone of voice based on the results. This enables the provision of information tailored to the traveler's emotions and the optimization of the in-vehicle environment.
[0655] "Methods for recognizing audio data in real time" refer to technologies that instantly convert audio into a digital format and understand its content.
[0656] "Means of machine-translating audio data and converting it into multiple languages" refers to an automated process for converting recognized audio content into different languages.
[0657] "Means of acquiring image data and extracting information using image analysis techniques" refers to techniques for extracting specific information from images acquired by cameras or sensors.
[0658] "A means of overlaying extracted information onto images of the real world" refers to a technology for integrating and displaying analyzed information within the actual visual environment.
[0659] "Means of exchanging data with multiple digital platforms using communication technology" refers to technologies for exchanging information between different digital systems.
[0660] "A means of analyzing emotions in real time and adjusting the tone of information and voice based on the results" refers to a technology that instantly grasps the user's emotional state and changes the expression of the information and voice provided according to those emotions.
[0661] "Means of optimizing the in-vehicle environment according to the passengers' emotions" refers to technology that adjusts the music, lighting, and temperature inside the vehicle according to the passengers' emotions to provide a comfortable space.
[0662] The system implementing this invention performs speech recognition, emotion analysis, translation, image analysis, and optimization of the in-vehicle environment in real time. The server recognizes speech data in real time and enables multilingual communication through real-time machine translation. For example, when a user converses with a passenger who speaks a different language, the speech data is instantly translated and output in the target language using speech synthesis technology. The terminal also detects emotions from the user's voice tone and facial expressions and sends emotion data to the server. The emotion engine analyzes this data and adjusts the way information is presented and the tone of voice according to the emotion.
[0663] Furthermore, the device sends image data acquired using its camera to a server, where relevant information is extracted using image analysis technology. This information is overlaid on the real-world scenery and presented intuitively to the user. When a user visits a specific tourist spot, the device can provide relevant historical and cultural information on the spot.
[0664] Furthermore, based on passenger emotional data, the server adjusts and optimizes the in-car environment, such as music and lighting, through the vehicle's environmental control system. For example, if a user is feeling anxious, the system provides a sense of security by selecting calming music and warm lighting. In this way, users can enjoy a more comfortable and personalized travel experience.
[0665] A concrete example of a prompt might be, "Translate the audio data immediately and adjust the tone of voice to match the passenger's emotions before outputting it." This would improve the user experience and support a more comfortable journey.
[0666] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0667] Step 1:
[0668] The device captures the user's voice using a microphone. This audio data is converted to a digital format and sent to the server. The input is the user's raw voice data, and the output is the digitally converted audio data. At this stage, a noise reduction filter is used to improve the audio quality.
[0669] Step 2:
[0670] The server converts the received digital audio data into text using a speech recognition engine. The input is the digitally converted audio data, and the output is text data in which the audio has been replaced with characters. Here, natural language processing techniques are applied to generate accurate textual information from the audio data.
[0671] Step 3:
[0672] The server translates text data generated from speech into different languages using a translation engine. The input is text data obtained by speech recognition, and the output is translated text data. In this step, a generative AI model is used to perform highly accurate translations.
[0673] Step 4:
[0674] The server converts the translated text into speech using a speech synthesis engine. The input is the translated text data, and the output is the synthesized speech data. At this point, the tone and speed of the speech are adjusted to make it sound natural and friendly.
[0675] Step 5:
[0676] The device outputs synthesized speech to the user through its speaker. The input is synthesized speech data, and the output is spoken language audible to the user. In this step, check and adjust whether the volume and sound quality are appropriate.
[0677] Step 6:
[0678] The device acquires emotional data from the user's facial expressions and voice tone and sends it to the server. The input is the user's voice tone and facial expression data, and the output is analyzable emotional data. This process combines camera images and voice analysis technology to determine emotions.
[0679] Step 7:
[0680] The server analyzes emotional data using an emotion engine and returns feedback to the terminal that adjusts the information and tone of voice based on the results. The input is emotional data, and the output is adjusted information and tone of voice. A generative AI model is used to make the adjustments that are most appropriate for the situation.
[0681] Step 8:
[0682] The terminal sends image data captured by the user's camera to the server. The input is raw image data, and the output is the data sent to the server. In this step, the image data is compressed and optimized to improve transmission efficiency.
[0683] Step 9:
[0684] The server uses image analysis technology to recognize objects and characters within images and extract relevant information. The input is image data, and the output is informational data based on the recognized content. A generative AI model is used to improve contextual understanding.
[0685] Step 10:
[0686] The server sends commands to optimize the in-vehicle environment settings based on the analyzed information. The input is recognition and emotion data, and the output is commands to adjust the environment settings. The vehicle's music, lighting, and temperature settings are adjusted based on specific feedback.
[0687] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0688] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0689] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0690] [Fourth Embodiment]
[0691] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0692] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0693] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0694] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0695] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0696] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0697] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0698] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0699] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0700] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0701] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0702] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0703] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0704] This invention is a system that enables smooth communication and information acquisition when travelers visit regions with different languages and cultures. To achieve this, the system's components work together in the following manner.
[0705] First, the "terminal" captures the user's voice data through the microphone. This voice data is then encoded in a digital format and sent to the server. The "server" receives the voice data and converts it into text data using speech recognition technology. The converted text is translated into the specified language by a translation engine and sent back to the terminal. The terminal then converts this translation result into audio format using speech synthesis technology and plays it back to the user, allowing users who do not understand the local language to enjoy conversations in real time.
[0706] Furthermore, the "device" sends image data captured by the user to the "server," which then analyzes the image. Using image analysis technology, the server recognizes, for example, landmarks and the content of signs in tourist areas, and retrieves relevant information from a database. The collected information is sent back to the device, which then uses AR technology to overlay the information onto the camera image. This allows the user to intuitively understand detailed information about the location.
[0707] Furthermore, this system utilizes advanced communication technology to link data with multiple digital platforms. Users can access services using their digital accounts and, for example, make smooth purchases locally using electronic payment platforms. During this process, the server manages the communication data, ensuring that payment procedures are completed securely and efficiently. This provides a comfortable travel experience without language or currency barriers.
[0708] As a concrete example, consider a scenario where a user visits an international festival where many languages are spoken. The user can use the server's translation service to communicate smoothly in multiple languages, and use the image analysis service to obtain and understand information displayed at various booths. As a result, cross-cultural interaction increases, and the enjoyment of the event is significantly enhanced.
[0709] Thus, the present invention aims to provide travelers with a new travel experience that overcomes language barriers and information disparities.
[0710] The following describes the processing flow.
[0711] Step 1:
[0712] The user speaks into the device. The device acquires the voice data through the microphone and encodes it in digital format. This data is then sent to the server via a communication module.
[0713] Step 2:
[0714] The server receives the audio data from the terminal and passes it to the speech recognition engine, which converts the audio data into text format. The converted text is then fed into a translation engine to translate it into multiple languages.
[0715] Step 3:
[0716] The server retrieves the translation result and converts it back into speech data using a speech synthesis engine. This converted speech data is then sent to the terminal.
[0717] Step 4:
[0718] The terminal plays the audio data received from the server through its speaker and outputs the translation result to the user as audio.
[0719] Step 5:
[0720] The user takes a picture with the device's camera. The device uses a communication module to send the image data to the server.
[0721] Step 6:
[0722] The server passes the image data received from the terminal to the image analysis engine, which analyzes its contents. Based on the analyzed information, it searches the database for related data and extracts additional information.
[0723] Step 7:
[0724] The server sends the extracted information back to the terminal. The terminal then displays this information overlaid on the camera image, using augmented reality (AR) technology to provide it to the user.
[0725] Step 8:
[0726] Users visually review the information provided and use it to guide their actions on-site. For example, they can use the payment function to make payments at local stores.
[0727] (Example 1)
[0728] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0729] Travelers often face difficulties communicating smoothly in different linguistic and cultural environments, and in efficiently obtaining local information. Furthermore, language and currency differences create barriers that hinder a smooth travel experience. This project aims to address these challenges and provide travelers with a more comfortable and fulfilling experience.
[0730] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0731] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into a different language, means for acquiring image information and extracting information using image analysis technology, means for overlaying the extracted information onto real-world images using augmented reality technology and displaying it, and means for exchanging information with multiple digital infrastructures using communication means. This makes it possible for travelers to communicate smoothly in a cross-cultural environment, intuitively acquire necessary information, and realize highly convenient travel regardless of language or currency.
[0732] "Audio information" refers to digital data of human voices and sounds captured through audio input devices such as microphones.
[0733] "Real-time recognition" refers to the process of instantly analyzing voice input and converting it into text data without delay.
[0734] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0735] "Image information" refers to digital images of visual data acquired through a camera or other imaging device.
[0736] "Image analysis technology" refers to computational techniques for analyzing image data and extracting specific features of objects, text, and scenes.
[0737] Augmented reality technology is a technique that overlays computer-generated information and images onto the real world's field of view.
[0738] "Communication methods" refer to technical techniques for sending and receiving data between multiple devices or systems.
[0739] "Digital infrastructure" is a general term for the systems and services that constitute information technology infrastructure and support the transmission and processing of data.
[0740] This invention is a system designed to facilitate smooth communication and information acquisition for travelers visiting regions with different language areas and cultural backgrounds. The system utilizes speech recognition, machine translation, image analysis, augmented reality technology, and electronic payment functionality.
[0741] The device captures the user's voice through the microphone and encodes the audio information into a digital format. The device then transmits this data to a server in real time. The server uses speech recognition technology to convert the audio information into text data, and then uses machine translation technology to translate that text into the specified language. Common cloud-based services can be used for the specific speech recognition and translation. The translation results are sent back from the server to the device, which uses speech synthesis technology to convert them into user-friendly speech and play it back. This allows the user to enjoy real-time conversations without being aware of language differences.
[0742] The device also uses its camera to send images captured by the user to a server. The server uses image analysis technology to analyze the content of the images and extract information such as landmarks and signs in tourist areas. This information is used to search for relevant information from multiple digital sources and is sent back to the device. The device then uses augmented reality technology to overlay this information onto the camera image. This makes it easier for the user to visually grasp detailed information about the location.
[0743] Regarding electronic payments, users can access their digital accounts via a terminal if they wish to make a purchase locally. The server exchanges data with the payment platform using secure and efficient communication methods to complete the transaction. As a result, users can enjoy a comfortable shopping experience even in environments with different languages and currencies.
[0744] As a concrete example, consider a scenario where a user visits an international festival where multiple languages are used. The user can communicate smoothly in multiple languages using the translation service provided by the server. They can also use the image analysis service to obtain and understand information displayed at various booths. Examples of prompts include: "I want to understand French using the translation service," "I want to know more about this landmark," and "I want to use electronic payment at the food stall."
[0745] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0746] Step 1:
[0747] The terminal captures the user's voice through the microphone. The input is the user's spoken voice. This is encoded into a digital format and sent to the server. Specifically, the analog audio signal is converted into digital data and compressed as needed. Through this process, the audio information is delivered to the server as digital data.
[0748] Step 2:
[0749] The server converts received audio data into text data using speech recognition technology. The input is encoded audio data. The server applies a speech recognition algorithm to convert the audio data into text data. Specifically, it extracts the audio waveform data as features and converts them into strings by comparing them with a language model. The output is text in which the user's utterance is expressed as sentences.
[0750] Step 3:
[0751] The server inputs the converted text into a machine translation service and translates it into the specified language. The input is text data generated by speech recognition. The translation engine performs interlingual translation processing, converting the text into different languages. Specifically, it uses translation memory and neural network models to generate text in the target language corresponding to the source text. The output is the translated text data.
[0752] Step 4:
[0753] The server sends the translated text data back to the terminal. The input is the translated text data. The terminal uses speech synthesis technology to convert this text data into speech. Specifically, it uses a speech synthesis engine to convert the text into an audio signal. The output is the generated speech, which is played back to the user through the speaker.
[0754] Step 5:
[0755] The user takes an image using the device's camera, and the device sends the image data to the server. The input is the image taken by the user. The device prepares to transfer the image data to the server in the appropriate format. This process provides the server with the basic data necessary to obtain detailed information about the object.
[0756] Step 6:
[0757] The server analyzes the received image data using image analysis techniques. The input is image data sent by the user. The image analysis algorithm identifies objects within the image and extracts their features. Specifically, it uses a computer vision model to recognize landmarks and characters within the image and extract semantic information. The output is the information resulting from the analysis.
[0758] Step 7:
[0759] The server retrieves relevant information from a database based on the results of image analysis. The input is feature information extracted by image analysis. The server refers to a pre-existing database to identify relevant content. Specifically, it retrieves a dataset containing landmark names and descriptions, related event information, etc. The output is data as relevant information.
[0760] Step 8:
[0761] The server sends the collected relevant information to the terminal, which then displays it using augmented reality technology. The input is the relevant information sent from the server. The terminal uses an AR engine to overlay the information onto the real-world camera image. Specifically, it superimposes the information onto specific points within the image and presents it to the user's vision. The output is the visualized information displayed in AR format.
[0762] Step 9:
[0763] When electronic payment is required, the user initiates an online payment through the terminal. Inputs include the selection of items the user wishes to purchase and payment information. The server communicates with the electronic payment platform after a secure authentication process. Specifically, it transmits payment information using encryption technology and obtains transaction authorization. The output is a confirmation message of the completed payment, displayed on the terminal.
[0764] (Application Example 1)
[0765] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0766] In today's global consumer society, language barriers and a lack of information that travelers from different languages and cultures face when visiting physical stores hinder smooth purchasing and customer service. This often results in travelers not fully understanding product information and having difficulty communicating with store staff. To address this challenge, there is a need for methods that minimize the language and information gap between travelers and stores, providing a smooth and comfortable shopping experience.
[0767] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0768] In this invention, the server includes means for recognizing voice information in real time, means for machine translating the recognized voice information and converting it into multiple languages, and means for acquiring image information and extracting information using image analysis technology. As a result, travelers can intuitively understand product information using augmented reality technology at physical stores they visit, and conversations with store staff are translated in real time, allowing them to make purchases with peace of mind without feeling any language or information gaps.
[0769] "Real-time recognition of speech information" is a technology that instantly converts a user's speech into a digital format and analyzes it as text information.
[0770] "Machine translation" is the process by which a computer converts natural language into another language, and it is a technology that supports communication between multiple languages.
[0771] "Acquiring image information" is the process of collecting visual data of objects and landscapes using cameras and sensors.
[0772] "Image analysis technology" refers to the technology used to analyze and extract specific information from digitized images.
[0773] Augmented reality technology is a technology that overlays digital information onto the real world environment, providing users with a visually enhanced experience.
[0774] An "information platform" is an online or offline system for exchanging and managing data.
[0775] "Consumer-relevant information" refers to detailed data about products and services that can help customers make purchasing decisions.
[0776] This invention is a system designed to enhance the shopping experience for travelers in physical stores with diverse language and cultural backgrounds. This system operates collaboratively between users, servers, and terminals.
[0777] The server first acquires audio information transmitted from the device. This audio information is captured in real time using the device's microphone. The server converts this audio information into text format using speech recognition technology. The Google Cloud Speech-to-Text API is often used for this. The text data is then translated into other languages via a translation engine. The Google Cloud Translation API is commonly used for this translation. The converted text is sent back from the server to the device, which then converts the text into speech using speech synthesis technology and provides it to the user. The Google Cloud Text-to-Speech API is useful for speech synthesis.
[0778] The device also uses its camera to photograph products in the store and sends the image information to a server. The server uses image analysis technologies such as Amazon Rekognition to recognize the products in the image and retrieves their detailed information from Firebase Cloud Firestore. The recognized information is then passed to the device, which uses augmented reality technologies such as Unity to display the product details in the user's field of view.
[0779] As a concrete example, consider a scenario where a user wants to purchase cosmetics at a store overseas. The user can use their smartphone camera to point at the cosmetics, and augmented reality technology will display their ingredients, price, and other consumers' reviews on the screen. Furthermore, questions and answers with store staff are translated in real time, allowing users to ask questions confidently despite language barriers.
[0780] Examples of prompts for a generative AI model:
[0781] "Translate voice input into a specified language in real time and output it as voice."
[0782] "Acquire product information from the camera and display it using augmented reality (AR)."
[0783] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0784] Step 1:
[0785] The user takes a picture of the products in the store using their smartphone camera. The captured image is saved on the device and then sent to the server. The input is the image information acquired by the camera, and the output is the transmission of the image data to the server.
[0786] Step 2:
[0787] The server analyzes the received image data. During this process, it uses image analysis technologies such as Amazon Rekognition to identify products within the image and extract information associated with those products. The input is the image data sent to the server, and the output is the analyzed product ID and related information.
[0788] Step 3:
[0789] The server searches for and retrieves relevant detailed information (price, ingredients, user reviews, etc.) from Firebase Cloud Firestore based on the analyzed product information. The input is the product ID, and the output is detailed information about the product.
[0790] Step 4:
[0791] The server sends the acquired product information to the terminal. The input is detailed product information, and the output is the transmission of information to the terminal.
[0792] Step 5:
[0793] The device displays received product information in the user's field of view using augmented reality technology. It overlays the information onto the real-world camera screen using an AR framework such as Unity. The input is product information transmitted from the server, and the output is the information visually presented to the user.
[0794] Step 6:
[0795] The user begins a conversation with the store clerk using their smartphone's microphone. The audio data is captured by the device and sent to the server. The input is the voice spoken by the user, and the output is the transmission of audio data to the server.
[0796] Step 7:
[0797] The server converts the received audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data.
[0798] Step 8:
[0799] The server translates the converted text into the specified language using the Google Cloud Translation API. The input is text data, and the output is the translated text in the other language.
[0800] Step 9:
[0801] The translated text is sent from the server to the terminal. The input is the translated text, and the output is the transmission of the translated data to the terminal.
[0802] Step 10:
[0803] The device converts the received translation data into speech using the Google Cloud Text-to-Speech API and outputs it through the speaker. The input is translated text, and the output is audio.
[0804] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0805] This invention is a system designed to support smooth communication and information acquisition for travelers visiting different cultural regions, and in particular, by incorporating an emotion engine, it enables the provision of information that takes the user's emotions into consideration. This system personalizes the user's experience through voice data, image data, and emotion recognition.
[0806] The "terminal" first captures the user's voice data in real time, converts it to a digital format, and sends it to the "server." The server converts the voice data into text using a speech recognition engine, and then translates it into multiple languages using a translation engine. The translation results are converted back into speech by a speech synthesis engine and sent to the terminal. The terminal then plays this speech for the user to support communication in different languages.
[0807] Simultaneously, the "device" acquires emotional data from the user's voice tone and facial expressions. This data is sent to a server and analyzed by an emotion engine. The emotion engine identifies the user's emotions and adjusts the tone and presentation of information based on the results. Emotion-responsive feedback makes the user feel more comfortable.
[0808] Furthermore, the "terminal" transmits image data captured by the user using the camera to the server. The "server" uses image analysis technology to recognize objects and text within the image and extract relevant information. This information is overlaid on the camera image, allowing the user to understand it intuitively.
[0809] Furthermore, users can link their accounts through the system and digital platform to receive personalized experiences based on their emotions. This linking allows users to receive information and services customized according to their emotional data, enriching their travel experience.
[0810] As a concrete example, consider a scenario where a user visits a restaurant in a foreign country and can order without experiencing a language barrier. If the user is feeling anxious or nervous, the emotion engine detects this, and the device suggests reassuring and encouraging messages, as well as interesting local information. As a result, the user can enjoy the experience with peace of mind. Thus, the present invention aims to significantly improve user satisfaction during travel by providing information and support in a way that takes the user's emotions into consideration.
[0811] The following describes the processing flow.
[0812] Step 1:
[0813] To enable voice input, the device captures audio data using its built-in microphone. This data is digitized in real time and sent to the server.
[0814] Step 2:
[0815] The server uses a speech recognition engine to convert the received audio data into text. This text is then passed to a translation engine for translation into multiple languages.
[0816] Step 3:
[0817] The translated text is regenerated as audio data by a speech synthesis engine. The server sends this audio data back to the terminal.
[0818] Step 4:
[0819] The device plays the regenerated audio data to the user through its speaker, providing the translated content in audio.
[0820] Step 5:
[0821] The device simultaneously analyzes the user's voice tone and facial expressions to acquire emotional data. This data is then sent to a server.
[0822] Step 6:
[0823] The server uses an emotion engine to analyze emotional data and identify the user's emotions. This emotional information is used to adjust service delivery.
[0824] Step 7:
[0825] When a user takes a picture, the device acquires the image data and sends it to the server.
[0826] Step 8:
[0827] The server uses image analysis technology to analyze image data and extract relevant information. The extracted results are then sent to the terminal.
[0828] Step 9:
[0829] The device overlays the extracted information onto the real-world camera footage, presenting it visually to the user.
[0830] Step 10:
[0831] Users can view personalized information provided through digital platforms and utilize emotion-based services. In this process, emotional information is used to refine the service content.
[0832] (Example 2)
[0833] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0834] In today's travel environment, when travelers visit different cultural regions, language barriers and cultural gaps make smooth communication difficult. This can make it challenging for travelers to deeply understand different cultures or to enjoy their trips with peace of mind. Furthermore, gathering and understanding local information is not easy. Traditional systems fail to adequately provide emotionally resonant information or enable real-time multilingual communication.
[0835] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0836] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for outputting the translated voice data by speech synthesis. This enables travelers to communicate smoothly with speakers of different languages in real time. Furthermore, by including means for analyzing the user's voice tone and facial expressions to acquire emotional data, and means for analyzing the acquired emotional data to adjust the way information is presented, personalized information provision according to the traveler's emotions is realized, enabling a more user-friendly travel experience. In addition, by including means for acquiring image data and extracting information using image analysis technology, and means for overlaying the extracted information onto real-world images, travelers can intuitively understand local visual information. This allows travelers to understand different cultures more deeply and enjoy their trip with peace of mind.
[0837] "Audio data" refers to a digital representation of sound waveforms, and is used to record human speech and analyze its content.
[0838] "Real-time recognition" refers to a method that processes audio and information instantly on the spot, allowing for immediate results.
[0839] "Machine translation" is a technology that uses computer algorithms to automatically convert text from one language to another.
[0840] "Multilingualization" is the process of translating information expressed in one language into multiple different languages and providing it to the public.
[0841] "Speech synthesis" is a technology that imitates human speech by artificially generating voices based on text data.
[0842] "Analyzing voice tone and facial expressions" refers to a method of analyzing collected audio tones and images to infer emotions and intentions.
[0843] "Emotional data" refers to information that indicates a person's emotional state and is used to identify specific emotions from voice and facial expressions.
[0844] "Image data" refers to a collection of visual information expressed in digital format, including graphic information such as photographs and drawings.
[0845] "Image analysis technology" refers to techniques for processing image data and extracting or recognizing useful information such as objects and text from it.
[0846] "Overlaying digital information onto real-world images" means adding digital information to the actual field of view to visually integrate it and improve visibility.
[0847] "Communication technology" refers to technologies for sending and receiving data and exchanging information between multiple devices and platforms.
[0848] "Electronic devices" refer to all fundamental devices that process information, including computers and their peripherals.
[0849] This system is designed to facilitate smooth communication across language barriers and personalize travel experiences for travelers visiting different cultural regions. The following describes a specific implementation of this system.
[0850] First, regarding the processing of audio data, the terminal captures the user's voice in real time. Hardware-wise, this involves using a mobile device with a built-in microphone or a dedicated audio capture device. The captured audio is then transmitted to the server using digital signal processing technology.
[0851] Next, the server uses a speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio data into text. The text is then translated into multiple languages by a translation engine (e.g., a general-purpose translation service). The converted text data is then re-converted into speech using a speech synthesis engine (e.g., general-purpose speech synthesis technology) and sent to the terminal. This allows the user to understand languages other than their native language.
[0852] The device also acquires emotional data from the user's voice tone and facial expressions. It analyzes the user's emotional state in real time using a combination of a microphone and camera, and transmits this data to a server. The server uses an emotion engine to analyze the emotional data and sends instructions to the device to provide the user with appropriate information.
[0853] Furthermore, the terminal transmits image data acquired by the user's camera to a server. The server uses image analysis technology to extract objects and text from the image data and provide the user with useful information. This allows the user to instantly obtain additional information about the images they have taken.
[0854] The system assists users in understanding menus and ordering without difficulty at restaurants in foreign countries. If the user feels anxious, the system suggests encouraging messages and interesting local information to alleviate their anxiety, ensuring reassuring communication. An example of a prompt for the generative AI model is: "The user needs help understanding the menu at the restaurant. If the user's tone indicates anxiety, what positive message should you offer?"
[0855] Thus, the present invention enables the provision of information based on the user's emotions and actively supports the travel experience.
[0856] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0857] Step 1:
[0858] The terminal captures the user's speech in real time using a microphone. The input audio signal is acquired in analog format. The terminal performs digital signal processing to convert the analog audio into a digital signal, generating digital audio data. This digital audio data is output and sent to the next processing step.
[0859] Step 2:
[0860] The server receives audio data transmitted from the terminal and converts it into text using a speech recognition engine. The input is digital audio data, which is converted into text format by speech recognition. The converted text data is output and used as input for translation processing.
[0861] Step 3:
[0862] The server translates the converted text data into multiple languages using a translation engine. Here, the input text data is converted into the specified target language. The translation engine understands the vocabulary and context to perform the optimal translation. The output is the translated text data, which is used for speech synthesis processing.
[0863] Step 4:
[0864] The server converts translated text data into speech using a speech synthesis engine. The input is translated text data, and the speech synthesis engine generates synthesized speech. The output is synthesized speech data (digital speech data), which is then prepared for transmission to the terminal.
[0865] Step 5:
[0866] The device receives audio data transmitted from the server and plays it back through its speaker so that the user can understand it. The input is synthesized speech data, which is then output as a physical audio signal. This output is ultimately delivered to the user, enabling communication in different languages.
[0867] Step 6:
[0868] The device captures the user's voice tone and facial expressions using its camera and microphone to acquire emotion data. The input is the user's current voice tone and facial expression. The device uses an analysis algorithm to identify emotions and outputs them as emotion data.
[0869] Step 7:
[0870] The server receives emotion data sent from the terminal and analyzes it using the emotion engine. The input is emotion data, which the emotion engine analyzes to identify the user's emotional state. The output is the analyzed emotion information, which is used to optimize information delivery.
[0871] Step 8:
[0872] The server selects and outputs information to the user based on the analyzed emotional information. For example, if the user is feeling anxious, it selects information that provides reassurance. The output information is sent to the terminal and displayed or played back.
[0873] (Application Example 2)
[0874] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0875] When travelers use autonomous vehicles, it is necessary to alleviate anxieties caused by language barriers and cultural differences and provide a more comfortable and personalized travel experience. However, current systems have challenges in adequately providing flexible information based on emotions and optimizing the in-vehicle environment.
[0876] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0877] In this invention, the server includes means for recognizing voice data in real time, means for machine translating the recognized voice data and converting it into multiple languages, and means for analyzing emotions in real time and adjusting the information and tone of voice based on the results. This enables the provision of information tailored to the traveler's emotions and the optimization of the in-vehicle environment.
[0878] "Methods for recognizing audio data in real time" refer to technologies that instantly convert audio into a digital format and understand its content.
[0879] "Means of machine-translating audio data and converting it into multiple languages" refers to an automated process for converting recognized audio content into different languages.
[0880] "Means of acquiring image data and extracting information using image analysis techniques" refers to techniques for extracting specific information from images acquired by cameras or sensors.
[0881] "A means of overlaying extracted information onto images of the real world" refers to a technology for integrating and displaying analyzed information within the actual visual environment.
[0882] "Means of exchanging data with multiple digital platforms using communication technology" refers to technologies for exchanging information between different digital systems.
[0883] "A means of analyzing emotions in real time and adjusting the tone of information and voice based on the results" refers to a technology that instantly grasps the user's emotional state and changes the expression of the information and voice provided according to those emotions.
[0884] "Means of optimizing the in-vehicle environment according to the passengers' emotions" refers to technology that adjusts the music, lighting, and temperature inside the vehicle according to the passengers' emotions to provide a comfortable space.
[0885] The system implementing this invention performs speech recognition, emotion analysis, translation, image analysis, and optimization of the in-vehicle environment in real time. The server recognizes speech data in real time and enables multilingual communication through real-time machine translation. For example, when a user converses with a passenger who speaks a different language, the speech data is instantly translated and output in the target language using speech synthesis technology. The terminal also detects emotions from the user's voice tone and facial expressions and sends emotion data to the server. The emotion engine analyzes this data and adjusts the way information is presented and the tone of voice according to the emotion.
[0886] Furthermore, the device sends image data acquired using its camera to a server, where relevant information is extracted using image analysis technology. This information is overlaid on the real-world scenery and presented intuitively to the user. When a user visits a specific tourist spot, the device can provide relevant historical and cultural information on the spot.
[0887] Furthermore, based on passenger emotional data, the server adjusts and optimizes the in-car environment, such as music and lighting, through the vehicle's environmental control system. For example, if a user is feeling anxious, the system provides a sense of security by selecting calming music and warm lighting. In this way, users can enjoy a more comfortable and personalized travel experience.
[0888] A concrete example of a prompt might be, "Translate the audio data immediately and adjust the tone of voice to match the passenger's emotions before outputting it." This would improve the user experience and support a more comfortable journey.
[0889] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0890] Step 1:
[0891] The device captures the user's voice using a microphone. This audio data is converted to a digital format and sent to the server. The input is the user's raw voice data, and the output is the digitally converted audio data. At this stage, a noise reduction filter is used to improve the audio quality.
[0892] Step 2:
[0893] The server converts the received digital audio data into text using a speech recognition engine. The input is the digitally converted audio data, and the output is text data in which the audio has been replaced with characters. Here, natural language processing techniques are applied to generate accurate textual information from the audio data.
[0894] Step 3:
[0895] The server translates text data generated from speech into different languages using a translation engine. The input is text data obtained by speech recognition, and the output is translated text data. In this step, a generative AI model is used to perform highly accurate translations.
[0896] Step 4:
[0897] The server converts the translated text into speech using a speech synthesis engine. The input is the translated text data, and the output is the synthesized speech data. At this point, the tone and speed of the speech are adjusted to make it sound natural and friendly.
[0898] Step 5:
[0899] The device outputs synthesized speech to the user through its speaker. The input is synthesized speech data, and the output is spoken language audible to the user. In this step, check and adjust whether the volume and sound quality are appropriate.
[0900] Step 6:
[0901] The device acquires emotional data from the user's facial expressions and voice tone and sends it to the server. The input is the user's voice tone and facial expression data, and the output is analyzable emotional data. This process combines camera images and voice analysis technology to determine emotions.
[0902] Step 7:
[0903] The server analyzes emotional data using an emotion engine and returns feedback to the terminal that adjusts the information and tone of voice based on the results. The input is emotional data, and the output is adjusted information and tone of voice. A generative AI model is used to make the adjustments that are most appropriate for the situation.
[0904] Step 8:
[0905] The terminal sends image data captured by the user's camera to the server. The input is raw image data, and the output is the data sent to the server. In this step, the image data is compressed and optimized to improve transmission efficiency.
[0906] Step 9:
[0907] The server uses image analysis technology to recognize objects and characters within images and extract relevant information. The input is image data, and the output is informational data based on the recognized content. A generative AI model is used to improve contextual understanding.
[0908] Step 10:
[0909] The server sends commands to optimize the in-vehicle environment settings based on the analyzed information. The input is recognition and emotion data, and the output is commands to adjust the environment settings. The vehicle's music, lighting, and temperature settings are adjusted based on specific feedback.
[0910] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0911] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0912] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0913] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0914] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0915] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0916] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0917] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0918] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0919] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0920] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0921] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0922] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0923] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0924] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0925] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0926] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0927] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0928] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0929] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0930] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0931] The following is further disclosed regarding the embodiments described above.
[0932] (Claim 1)
[0933] A means of recognizing audio data in real time,
[0934] A means of machine-translating recognized speech data and converting it into multiple languages,
[0935] A means of acquiring image data and extracting information using image analysis techniques,
[0936] A means of displaying extracted information by overlaying it onto images of the real world,
[0937] A means of exchanging data with multiple digital platforms using communication technology,
[0938] A system that includes this.
[0939] (Claim 2)
[0940] The system according to claim 1, further comprising means for outputting a translation result generated from audio data by speech synthesis.
[0941] (Claim 3)
[0942] The system according to claim 1, further comprising means for automatically searching for and providing relevant information for travelers based on the results of image data analysis.
[0943] "Example 1"
[0944] (Claim 1)
[0945] A means of recognizing voice information in real time,
[0946] A means of machine-translating recognized speech information and converting it into a different language,
[0947] A means of acquiring image information and extracting information using image analysis technology,
[0948] A means of overlaying extracted information onto real-world images using augmented reality technology,
[0949] A means of exchanging information with multiple digital infrastructures using communication means,
[0950] A system that includes this.
[0951] (Claim 2)
[0952] The system according to claim 1, further comprising means for outputting translated audio using speech synthesis technology.
[0953] (Claim 3)
[0954] The system according to claim 1, further comprising means for automatically searching for and providing relevant information for travelers based on the results of image information analysis.
[0955] "Application Example 1"
[0956] (Claim 1)
[0957] A means of recognizing voice information in real time,
[0958] A means of machine-translating recognized speech information and converting it into multiple languages,
[0959] A means for acquiring image information and extracting information using image analysis technology,
[0960] A method for presenting extracted information by overlaying it onto images of the real world,
[0961] A means of exchanging data with multiple information platforms using communication technology,
[0962] A means of using augmented reality technology to present product information,
[0963] A system that includes this.
[0964] (Claim 2)
[0965] The system according to claim 1, further comprising means for providing a translation result generated from audio information by speech generation.
[0966] (Claim 3)
[0967] The system according to claim 1, further comprising means for automatically searching for and providing relevant consumer information based on the results of image information analysis.
[0968] "Example 2 of combining an emotion engine"
[0969] (Claim 1)
[0970] A means of recognizing audio data in real time,
[0971] A means of machine-translating recognized speech data and converting it into multiple languages,
[0972] A means for outputting translated audio data by speech synthesis,
[0973] A means for analyzing the user's voice tone and facial expressions to obtain emotional data,
[0974] A means of analyzing acquired emotional data and adjusting the way information is presented,
[0975] A means of acquiring image data and extracting information using image analysis techniques,
[0976] A means of displaying extracted information by overlaying it onto images of the real world,
[0977] A means of exchanging information with multiple electronic devices using communication technology,
[0978] A system that includes this.
[0979] (Claim 2)
[0980] The system according to claim 1, further comprising means for generating customized information based on the user's emotions.
[0981] (Claim 3)
[0982] The system according to claim 1, further comprising means for automatically searching for and providing relevant information for travelers based on the results of image data analysis.
[0983] "Application example 2 of combining emotional engines"
[0984] (Claim 1)
[0985] A means of recognizing audio data in real time,
[0986] A means of machine-translating recognized speech data and converting it into multiple languages,
[0987] A means of acquiring image data and extracting information using image analysis techniques,
[0988] A means of displaying extracted information by overlaying it onto images of the real world,
[0989] A means of exchanging data with multiple digital platforms using communication technology,
[0990] A means of analyzing emotions in real time and adjusting information and tone of voice based on the results,
[0991] A means of optimizing the in-vehicle environment according to the emotions of the passengers,
[0992] A system that includes this.
[0993] (Claim 2)
[0994] The system according to claim 1, further comprising means for outputting a translation result generated from audio data by speech synthesis.
[0995] (Claim 3)
[0996] The system according to claim 1, further comprising means for automatically searching for and providing relevant information for travelers based on the results of image data analysis. [Explanation of symbols]
[0997] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of recognizing audio data in real time, A means of machine-translating recognized speech data and converting it into multiple languages, A means of acquiring image data and extracting information using image analysis techniques, A means of displaying extracted information by overlaying it onto images of the real world, A means of exchanging data with multiple digital platforms using communication technology, A system that includes this.
2. The system according to claim 1, further comprising means for outputting a translation result generated from audio data by speech synthesis.
3. The system according to claim 1, further comprising means for automatically searching for and providing relevant information for travelers based on the results of image data analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A