System

The system addresses the challenge of customizable user interfaces by receiving natural language input, analyzing user requests, and dynamically updating the interface to provide personalized information in real-time, enhancing usability and accessibility for all users.

JP2026018052APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119113
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

Smart Images

  • Figure 2026018052000001_ABST
    Figure 2026018052000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving a natural language input from a user; natural language processing means for parsing a request of the user; generating means for generating relevant information based on the parsed request; and means for dynamically updating a user interface to display the generated information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, advances in digital technology have led to the emergence of many applications, but most of them use a fixed user interface (UI). This makes it difficult to customize them to meet individual user needs, and elderly users and those unfamiliar with technology often find it difficult to use the apps. Furthermore, complex UIs can lead to information overload, which can degrade the user experience. There is a need for systems that can resolve these issues and enable users to quickly and intuitively obtain the information they need. [Means for solving the problem]

[0005] The present invention provides a system including a means for receiving natural language input from a user, a natural language processing means for analyzing the user's request, a generating means for generating related information based on the analyzed request, and a means for dynamically updating a user interface to display the generated information. The system also includes a means for acquiring user location information and providing it to the generating means, and a means having a database for searching related information based on the user's location information. Furthermore, the user interface for displaying the generated information supports voice input and text input, and is also compatible with users with visual and hearing impairments, making the system accessible to a wide range of users. This configuration provides the information desired by the user in real time, preventing information overload and improving usability.

[0006] "Means for receiving natural language input from a user" refers to an interface or device for capturing input made by a user in natural language.

[0007] "Natural language processing means for analyzing user requests" refers to algorithms or programs that analyze the user's natural language input and extract their intentions and requests.

[0008] "Generation means for generating relevant information based on analyzed requests" refers to a system or software that dynamically generates the information or content desired by the user based on requests analyzed by natural language processing means.

[0009] "Means for dynamically updating the user interface to display the generated information" refers to the function of displaying the information generated by the generating means in real time and dynamically changing the interface in response to user operations or requests.

[0010] "Means for obtaining the user's location information and providing it to the generation means" refers to a mechanism for obtaining the user's current location via GPS or a network and passing that information to the generation means.

[0011] "Means having a database that searches for relevant information based on the user's location information" refers to a database that searches for optimal information based on location information and provides it to the user, and a method of using the database.

[0012] "Means supporting voice and text input" refers to interfaces and technologies that allow users to input information by voice or text.

[0013] "Means to accommodate users with visual and hearing impairments" refers to assistive technologies and designs that accommodate a diverse range of users, such as audio output for the visually impaired and text display for the hearing impaired. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] overview

[0036] The present invention provides a system that accepts natural language input from users, analyzes their requests, and displays and updates the generated information in real time. The system acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0037] System configuration

[0038] The system consists of the following main components:

[0039] 1. User Interface (UI)

[0040] An interface that allows users to enter questions or requests in natural language.

[0041] 2. Natural Language Processing (NLP) Engine

[0042] Algorithms for parsing user input and extracting its intent.

[0043] 3. Generation AI

[0044] Generates appropriate information based on the request analyzed by the NLP engine.

[0045] 4. Data Management Module

[0046] The user's location information and past usage history are stored and provided to the generating AI.

[0047] 5. Display module

[0048] Dynamically update the UI to display the generated information.

[0049] Program processing

[0050] Getting User Input

[0051] The device receives natural language input from the user (e.g., "What are the nearest restaurants?"). Voice input or text input is possible through the user interface. The user's input is stored in text form and passed on to the next processing step.

[0052] Natural Language Processing

[0053] The device passes user input to a natural language processing engine to parse the request. The NLP engine extracts meaningful keywords and phrases from the input text to identify the user's intent. This process reveals a request such as "I want to find nearby restaurants."

[0054] information generation

[0055] The server sends these requests to the generation AI, which generates the most appropriate information in real time based on the user's location information obtained from the data management module. The generated information searches a database related to the user's location information to provide the most useful information for the user.

[0056] Dynamic UI Updates

[0057] The device passes the information received from the generation AI to the display module, which dynamically updates the user interface, allowing the user to instantly see the information they were looking for (e.g., a list of nearby restaurants). The same process can be repeated to provide real-time responses when the user asks for more information or other options.

[0058] Specific examples

[0059] Example 1: Finding a restaurant

[0060] user

[0061] A user types, "What restaurants are nearby?"

[0062] Terminal

[0063] It takes input at the user interface and sends it to a natural language processing engine.

[0064] NLP Engine

[0065] Parse the input to determine the intent to find "nearby restaurants."

[0066] Generation AI

[0067] Generate a list of nearby restaurants based on the user's location.

[0068] Terminal

[0069] The list is displayed in the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0070] Example 2: Restaurant details

[0071] user

[0072] User types, "Give me details about Restaurant A."

[0073] Terminal

[0074] It takes input and sends it to a natural language processing engine.

[0075] NLP Engine

[0076] Analyze the input and determine the intent to ask for "details about Restaurant A."

[0077] Generation AI

[0078] Generate detailed information about Restaurant A, such as its address, opening hours, and menu.

[0079] Terminal

[0080] Display detailed information in the user interface.

[0081] In this way, the system of the present invention smoothly carries out a series of processes, starting with the user's natural language input, through analysis, information generation, and display, thereby intuitively and efficiently providing the information the user desires.

[0082] The processing flow will be explained below.

[0083] Step 1:

[0084] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0085] Step 2:

[0086] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0087] Step 3:

[0088] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0089] Step 4:

[0090] The device sends the request content and location information analyzed by the NLP engine to the generation AI. The location information is obtained from the device's location information service.

[0091] Step 5:

[0092] The server provides the request details and location information to the AI, which then prepares to generate appropriate restaurant information based on this information.

[0093] Step 6:

[0094] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[0095] Step 7:

[0096] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0097] Step 8:

[0098] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to see a list of nearby restaurants and their details.

[0099] Step 9:

[0100] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[0101] Step 10:

[0102] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[0103] As described above, each step works in conjunction to provide users with the information they want in real time, creating an intuitive and efficient user experience.

[0104] Example 1

[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0106] Conventional information provision systems have had the problem of making it difficult to quickly and accurately obtain the information users want. It has also been difficult to effectively utilize users' location information and provide information tailored to individual user needs. Furthermore, there has been a lack of effective means of providing information to users with visual and hearing impairments.

[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0108] In this invention, the server includes a terminal means for receiving natural language input from a user, a natural language processing means for analyzing the user's request, a generation AI means for generating relevant information based on the analyzed request, and a display means for dynamically updating a user interface to display the generated information. This allows users to quickly and accurately obtain the information they desire, and makes it possible to effectively utilize the user's location information and provide personalized information based on their location. Furthermore, by supporting voice input and text input, information can be provided effectively to users with visual and hearing impairments.

[0109] "Terminal Means" refers to a device or interface for receiving natural language input from a user.

[0110] "Natural language processing" refers to an algorithm or system that analyzes a user's input and determines their intent.

[0111] "Generative AI means" refers to an artificial intelligence model that generates relevant information based on requests analyzed by natural language processing means.

[0112] "Display Means" means a system for visually and audibly presenting the information generated by the Generating AI Means to the user and dynamically updating the user interface.

[0113] "Data management means" refers to a system that has the function of acquiring user location information and providing it to the generating AI means.

[0114] "Database" refers to a collection of information used to search for relevant information based on a user's location.

[0115] "User interface" refers to the screen and input devices that allow a user to interact with a system.

[0116] "Voice and text input" refers to the methods by which a user inputs information into a system in the form of voice or text.

[0117] "Visually and hearing impaired users" refers to users who have difficulty obtaining information through normal means due to visual or hearing limitations.

[0118] This invention is a system that receives natural language input from users, analyzes their requests, generates relevant information, and dynamically updates the user interface. The system features the ability to obtain the user's location information and provide personalized information based on their usage history. It also supports both voice and text input to accommodate users with visual and hearing impairments.

[0119] The main components of the system are:

[0120] 1. Terminal means:

[0121] Accepts natural language input from the user. In the case of voice input, converts speech to text using speech recognition technology (e.g., Google Speech-to-Text API). Also accepts text input through the user interface.

[0122] 2. Natural Language Processing Tools:

[0123] Analyze user input and determine its intent, specifically by using a natural language processing engine (e.g., spaCy) to parse the text and extract meaningful keywords and phrases.

[0124] 3. Generation AI means:

[0125] Relevant information is generated based on the analysis results. The generation AI generates appropriate information by referencing the user's location information and usage history. OpenAI GPT-4 and other models can be used as generation AI models.

[0126] 4. Display means:

[0127] The generated information is displayed in a user interface that is dynamically updated based on the data received from the generating AI, providing the user with visual and auditory information.

[0128] As a concrete example, the following shows what happens when a user types "What restaurants are nearby?" The process obtains the user's location information and generates and displays a list of nearby restaurants. The device uses the Google Speech-to-Text API to convert speech to text, and sends this text to a natural language processing engine (spaCy) to extract keywords. The generative AI (OpenAI GPT-4) generates a list based on the user's location information and displays it on the screen.

[0129] Example prompt sentence:

[0130] Example 1: Finding a restaurant

[0131] "User is looking for nearby restaurants. GPS information is ____. Please provide a list of the nearest restaurants."

[0132] Example 2: Restaurant details

[0133] "A user is looking for more information about Restaurant A. Please provide details such as Restaurant A's address, opening hours, and menu."

[0134] This process allows users to quickly and accurately obtain the information they need, and the system can provide personalized information based on location information and is also compatible with users with visual and hearing impairments.

[0135] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0136] Step 1: Getting User Input

[0137] The device receives natural language input from the user. The user can enter voice or text through the device's user interface. For voice input, the device uses speech recognition software (e.g., a speech recognition API) to convert the speech to text. The converted text is stored in memory and passed on to the next processing step.

[0138] Specific behavior:

[0139] The user speaks into the device's microphone, "What restaurants are nearby?"

[0140] The device uses a speech recognition API to convert the speech into text, generating the text "What restaurants are nearby?"

[0141] Input and Output:

[0142] Input: User voice or text input

[0143] Output: User request in text format

[0144] Step 2: Natural Language Processing

[0145] The device sends the text input received from the user to a natural language processing engine (e.g., NLP), which analyzes the text and extracts meaningful keywords and phrases. The extracted intents and keywords are passed on to the next step.

[0146] Specific behavior:

[0147] The device sends the text "What restaurants are nearby?" to a natural language processing engine.

[0148] A natural language processing engine analyzes the text and extracts keywords such as "nearby" and "restaurant."

[0149] Input and Output:

[0150] Input: User request in text format

[0151] Output: Extracted keywords and intent

[0152] Step 3: Information generation

[0153] The server sends the user's intent, analyzed by the natural language processing engine, to the generative AI (e.g., generative AI model), which generates relevant information. It also obtains the user's location information from the data management module and provides this location information to the generative AI. Based on this information, the generative AI generates information appropriate to the user's request in real time.

[0154] Specific behavior:

[0155] The server obtains the user's location information (e.g., latitude 35.6895, longitude 139.6917) from the data management module.

[0156] The intent and location information obtained from the natural language processing engine is sent to the generation AI.

[0157] The generation AI generates a list of restaurants such as "Restaurant A (500m)" and "Restaurant B (700m)."

[0158] Input and Output:

[0159] Input: Extracted keywords and intent, user location

[0160] Output: Generated related information (e.g., a list of restaurants)

[0161] Step 4: Dynamic UI Updates

[0162] The terminal passes the output data of the generated AI received from the server to the display module, which dynamically updates the user interface based on this data and provides the user with visual and auditory information.

[0163] Specific behavior:

[0164] The terminal passes the list of restaurants received from the server to the display module.

[0165] The display module updates the user interface to show information such as "Restaurant A (Distance: 500m)" and "Restaurant B (Distance: 700m)."

[0166] Users can view a list of nearby restaurants on their device screen.

[0167] Input and Output:

[0168] Input: Generated related information (e.g., a list of restaurants)

[0169] Output: Dynamically updated user interface and visual / auditory information

[0170] (Application example 1)

[0171] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0172] Conventional food delivery systems face challenges in efficiently obtaining the information users require and making it difficult to check order and delivery status in real time. They also lack an interface that is easy for visually and hearing impaired users to use. This reduces the quality of the user experience and limits the number of users.

[0173] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0174] In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, means for providing related information based on the user's location information, and means for receiving natural language input in voice or text format. This makes it possible to provide a system that allows users to efficiently obtain information using natural language and check order and delivery status in real time. It also makes it possible to provide an easy-to-use interface that is compatible with users with visual and hearing impairments.

[0175] A "means for receiving natural language input from a user" is any device or software for obtaining natural language input from a user in the form of speech or text.

[0176] "Natural language processing means for analyzing user requests" refers to technologies and algorithms that analyze the natural language entered by the user and understand their intentions and requests.

[0177] The "means for generating related information based on the analyzed request" refers to a method or technology for generating necessary information based on the results of analyzing the user's request.

[0178] "Means for dynamically updating the user interface to display generated information" refers to technologies and systems that change and update the displayed content in real time to present generated information to the user on screen or via audio.

[0179] "Means for providing relevant information based on a user's location information" refers to technologies or methods for utilizing a user's current location information to provide information related to that location.

[0180] A "means for receiving natural language input in voice or text form" is a device or software that recognizes and captures natural language input by a user through voice or text.

[0181] The "means for acquiring user location information and providing it to the generating means" refers to a technology or system for acquiring user location information and passing it to the means for generating related information.

[0182] "Means having a database for searching related information" refers to a database for searching related information based on the user's location information and a method for using the database.

[0183] "Means for tracking and displaying order and delivery status in real time" refers to technologies and systems that allow users to check the fulfillment status of their orders and the progress of deliveries in real time and display them to users.

[0184] "Means having an interface to assist in ordering procedures" refers to devices or software that have an interface that provides guidance and input assistance to enable users to place orders smoothly.

[0185] As an embodiment of this invention, we present a system that allows users to smoothly use food delivery services using natural language. The system is composed of components such as a user interface, a natural language processing engine, a generative AI, a data management module, and a display module.

[0186] Hardware and software used

[0187] Hardware:

[0188] Head-mounted displays and smartphones

[0189] microphone

[0190] software:

[0191] Python

[0192] geopy library: Getting location information

[0193] requests library: API calls

[0194] transformers library: natural language processing

[0195] speech_recognition library: speech recognition

[0196] Program processing explanation

[0197] User Interface

[0198] The user interface supports both voice and text input: users can say things like "I'd like to order a pizza," and the input is captured by a microphone and converted into text.

[0199] Natural Language Processing

[0200] User requests, entered via voice or text, are analyzed by a natural language processing engine. The transformers library is used to identify the user's intent. For example, a request such as "I want to order a pizza" can be interpreted as "Find nearby restaurants that serve pizza."

[0201] Information Generation and Display

[0202] The parsed request is sent to the generation AI, which generates relevant information based on the user's location information obtained from the data management module. For example, it generates a list of pizza restaurants nearest to the user's current location. The generated information is then passed to the display module, which dynamically updates the user interface and presents it to the user.

[0203] Order and delivery status tracking

[0204] The interface supports the user with menu information and ordering procedures for the restaurant they select, and also allows them to track and display the delivery status in real time after ordering, providing users with information such as the order progress and estimated arrival time.

[0205] Examples and prompts

[0206] For example, if a user says "I'd like to order pizza" to the head-mounted display, the system will analyze the request and display a list of nearby restaurants that serve pizza. The menu of the restaurant selected from the list will also be displayed, allowing the user to immediately proceed with the ordering process. After placing an order, the user can also check the current delivery status by asking "What's the status of my order?"

[0207] Example prompt sentence:

[0208] User: I want to order a pizza.

[0209] System: Finding pizza restaurants near you...

[0210] System: Choose from Restaurant A (distance: 500m), Restaurant B (distance: 700m), or Restaurant C (distance: 900m).

[0211] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0212] Step 1:

[0213] The user provides input via voice or text. The input is in the form of natural language, such as "I'd like to order a pizza." The device takes this input and, if it's voice, converts it to text using speech recognition software. The input data is passed on to the next processing step in natural language.

[0214] Step 2:

[0215] The device passes the captured input data to a natural language processing engine, which uses the transformers library to parse the input text and determine the user's intent, which is then converted into a clear instruction: "Find nearby pizza restaurants."

[0216] Step 3:

[0217] The server sends the parsed request to the generation AI, which generates a list of pizza restaurants based on the user's location information obtained from the data management module. The generation AI uses the user's location information as coordinates to extract data on pizza restaurants in the vicinity from a database. The output data is generated as a list of pizza restaurants and is passed to the next processing step.

[0218] Step 4:

[0219] The server sends the generated list of pizza restaurants to the display module. The display module dynamically updates the user interface and presents the list of pizza restaurants to the user. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed.

[0220] Step 5:

[0221] The user selects a restaurant from the list. The user's selection information is acquired by the terminal and sent to the server. The server then retrieves the detailed menu of the selected restaurant from the database and uses a generation AI to generate interface information to assist with the ordering process.

[0222] Step 6:

[0223] The server sends the generated interface information to the display module, which displays the menu of the selected restaurant to the user and allows the user to complete the ordering process. Once the user confirms the order, the information is sent back to the server.

[0224] Step 7:

[0225] Once an order is confirmed, the server stores the order information in a database and generates delivery information. The delivery status is updated in real time.

[0226] Step 8:

[0227] When the user asks, "What's the status of my order?", the device again sends this new input to the natural language processing engine, which analyzes it and determines that it's a request to check the delivery status.

[0228] Step 9:

[0229] The server retrieves delivery status information from the database and updates the current delivery status. The information is passed to the display module, which displays the delivery status to the user. For example, statuses such as "Currently cooking," "Delivering," and "Estimated arrival time: 20 minutes later" are displayed in real time.

[0230] In this way, a system is provided that allows users to efficiently use food delivery services using natural language input.

[0231] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0232] overview

[0233] The present invention provides a system that receives natural language input from a user, analyzes the user's request, and displays and updates the generated information in real time. It also combines an emotion engine that recognizes the user's emotions. The system also acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0234] System configuration

[0235] The system consists of the following main components:

[0236] 1. User Interface (UI)

[0237] An interface that allows users to enter questions or requests in natural language.

[0238] 2. Natural Language Processing (NLP) Engine

[0239] Algorithms for parsing user input and extracting its intent.

[0240] 3. Generation AI

[0241] Generates appropriate information based on the request analyzed by the NLP engine.

[0242] 4. Data Management Module

[0243] The user's location information and past usage history are stored and provided to the generating AI.

[0244] 5. Display module

[0245] Dynamically update the UI to display the generated information.

[0246] 6. Emotion Engine

[0247] An engine that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[0248] Program processing

[0249] Getting User Input

[0250] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0251] Natural Language Processing

[0252] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0253] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0254] Emotion recognition

[0255] The device's emotion engine recognizes the user's emotions from natural language input, voice, and facial expressions. For example, it can identify the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[0256] information generation

[0257] The server sends the request content, emotion information, and location information analyzed by the NLP engine and emotion engine to the generation AI. The location information is obtained from the device's location information service.

[0258] The generative AI uses this information to generate appropriate restaurant information in real time. For example, if the user is feeling stressed, it will suggest quiet and relaxing restaurants.

[0259] Dynamic UI Updates

[0260] The device analyzes the information it receives from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants) and repeat the same process to provide real-time responses for further inquiries about details or other options.

[0261] Specific examples

[0262] Example 1: Finding a restaurant

[0263] 1. User types "What restaurants are near me?"

[0264] 2. The device sends the input to a natural language processing engine for analysis.

[0265] 3. The NLP engine identifies the intent: "Looking for nearby restaurants."

[0266] 4. The emotion engine identifies stress from the user's voice.

[0267] 5. Generative AI will generate a list of quiet and relaxing restaurants based on the user's location and emotional information.

[0268] 6. The device displays the listed information on the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0269] Example 2: Restaurant details

[0270] 1. The user types, "Tell me the details about Restaurant A."

[0271] 2. The device sends the input to a natural language processing engine for analysis.

[0272] 3. The NLP engine identifies the intent as "I want to know more about Restaurant A."

[0273] 4. The emotion engine identifies that the user is relaxed in their voice.

[0274] 5. The generation AI generates detailed information about Restaurant A, such as its address, opening hours, and menu.

[0275] 6. The device displays the detailed information in the user interface.

[0276] In this way, the system of the present invention provides the information the user wants in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[0277] The processing flow will be explained below.

[0278] Step 1:

[0279] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0280] Step 2:

[0281] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0282] Step 3:

[0283] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0284] Step 4:

[0285] The device's emotion engine recognizes the user's emotions from natural language input and voice. For example, it identifies the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[0286] Step 5:

[0287] The device sends the request content analyzed by the NLP engine, the emotional information recognized by the emotion engine, and the user's location information to the generation AI. The location information is obtained from the device's location information service.

[0288] Step 6:

[0289] The server's AI generates optimal restaurant information based on the request, emotional information, and location information received. If the user is feeling stressed, it will prioritize quiet and relaxing restaurants.

[0290] Step 7:

[0291] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[0292] Step 8:

[0293] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0294] Step 9:

[0295] The device analyzes the information it receives from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants).

[0296] Step 10:

[0297] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[0298] Step 11:

[0299] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[0300] Through these steps, the system provides users with the information they want in real time, realizing an intuitive and efficient user experience, allowing users to obtain the most appropriate information according to their emotions and circumstances.

[0301] Example 2

[0302] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0303] Conventional natural language processing systems provide only simple information without considering the user's emotions or location, making it difficult to provide appropriate information that meets the user's needs.In addition, the information provided to users with visual and hearing impairments is insufficient.

[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0305] In this invention, the server includes means for recognizing a user's emotion from the user's natural language input, voice, and facial expression, means for analyzing the emotion information, and means for acquiring the user's location information and providing it to the generation means. This makes it possible to provide appropriate information based on the user's emotion and location information, and realizes a system that can also accommodate users with visual and hearing impairments.

[0306] "Natural language input" is a method by which users enter questions or requests in everyday language.

[0307] "Natural language processing" is a technology that analyzes user input and extracts their intent.

[0308] The "generator" is a function that generates related information based on the analyzed request.

[0309] "User interface" refers to a screen or operating interface that is dynamically updated to display generated information.

[0310] An "emotion engine" is a system that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[0311] "Emotional information" is data about the user's emotional state as recognized using the emotion engine.

[0312] "Location information" is data that indicates a user's current location.

[0313] "Dynamic update" refers to changing screens and data in real time as needed.

[0314] A "database" is a system for storing and searching information.

[0315] "Visually and hearing impaired users" refers to people who have limited vision or hearing, and includes special assistive devices to accommodate these users.

[0316] The present invention combines a system that receives natural language input from users, analyzes their requests, and displays and updates the generated information in real time with an emotion engine that recognizes the user's emotions. The system also has the ability to acquire the user's location information and provide relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0317] System configuration

[0318] The system consists of the following main components:

[0319] 1. User Interface (UI)

[0320] An interface that allows users to input questions or requests in natural language, supports both speech and text input, and is designed to be accessible to users with visual or hearing impairments.

[0321] 2. Natural Language Processing (NLP) Engine

[0322] It is an algorithm that analyzes user input and extracts their intent. For example, if a user types "What restaurants are nearby?", an NLP engine analyzes this input and identifies the intent as "looking for a nearby restaurant."

[0323] 3. Generation AI

[0324] It generates appropriate information based on requests analyzed by the NLP engine, and also references the emotion engine and location information to provide information that is best suited to the user's situation.

[0325] 4. Data Management Module

[0326] The user's location information and past usage history are stored and provided to the AI ​​generator. In particular, location information is obtained from the device's location information service.

[0327] 5. Display module

[0328] The UI for displaying the generated information is dynamically updated, allowing users to instantly see the information they are looking for.

[0329] 6. Emotion Engine

[0330] This engine recognizes the user's emotions from natural language input, voice, and facial expressions. The emotion engine identifies emotions such as stress or relaxation from the tone of voice and the content of the input.

[0331] Specific operation example

[0332] Example 1: Finding a restaurant

[0333] When a user types "What restaurants are nearby?", the device's voice recognition system converts the speech into text, which is then sent to a natural language processing engine. The NLP engine analyzes the input and identifies the user's intent: "I'm looking for a nearby restaurant." At the same time, the emotion engine identifies stress from the user's voice, and this information is sent to the server's generation AI. The generation AI then creates a list of "quiet, relaxing restaurants nearby" and sends the information to the device. The device dynamically updates the user interface, providing information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[0334] Example 2: Restaurant details

[0335] When a user types "Tell me more about Restaurant A," the device sends the input to a natural language processing engine for analysis. The NLP engine identifies the intent as "I want to know more about Restaurant A," and the emotion engine identifies a relaxed tone from the user's voice. The server then requests detailed information about Restaurant A, such as its address, opening hours, and menu, from the generation AI, and the generated details are sent to the device and displayed in the user interface.

[0336] Prompt Sentence Examples

[0337] "Can you recommend a quiet, relaxing restaurant nearby?"

[0338] "Please tell me more information about Restaurant A."

[0339] In this way, the system of the present invention provides the information the user desires in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[0340] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0341] Step 1: Getting User Input

[0342] A user inputs a question or request through the application's user interface. For example, the user may input "What restaurants are nearby?" by voice or text. The input may be the user's voice or text data. In the case of voice, the device's speech recognition system converts the voice to text. The converted text is "What restaurants are nearby?"

[0343] Specific behavior:

[0344] The user presses the microphone button and says, "What restaurants are nearby?"

[0345] The device's voice recognition system converts the speech into text.

[0346] The text data "What are the nearby restaurants?" is saved.

[0347] Step 2: Natural Language Processing

[0348] The device sends user input to a natural language processing (NLP) engine. The input is text data such as "What restaurants are nearby?" The NLP engine analyzes the user's input and extracts its intent. The output of the NLP engine is the analysis result, which identifies the user's intent as "looking for a nearby restaurant."

[0349] Specific behavior:

[0350] The text data "What restaurants are nearby?" is sent to the NLP engine.

[0351] An NLP engine analyzes text data and identifies intent.

[0352] The analysis result, "Looking for a nearby restaurant," is returned to the device.

[0353] Step 3: Recognize emotions

[0354] The device's emotion engine recognizes the user's emotions from their natural language input, voice, and facial expressions. The input is the user's voice and natural language input. The emotion engine identifies emotions from the tone of the voice and the content of the input. The output is emotional information that the user is "feeling stressed."

[0355] Specific behavior:

[0356] The user's voice is sent to the emotion engine.

[0357] The emotion engine analyzes the tone and pace of the voice.

[0358] Emotional information "user is feeling stressed" is stored on the device.

[0359] Step 4: Information Generation

[0360] The server sends the analyzed request content (user intent), emotional information, and location information to the generation AI. The inputs are the analysis result "looking for a nearby restaurant," emotional information "stress," and location information. The generation AI generates appropriate information based on this information. The output is the information "quiet, relaxing restaurants nearby."

[0361] Specific behavior:

[0362] The analysis result, "Looking for a nearby restaurant," emotional information, "stress," and location information are sent to the generating AI.

[0363] The generative AI creates a prompt and sends it to the AI ​​model.

[0364] The AI ​​model generates appropriate restaurant information and returns a list such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[0365] Step 5: Dynamic UI Updates

[0366] The device analyzes the information received from the generation AI and dynamically updates the user interface. The input is the generated restaurant list. The output is specific information displayed on the user interface (e.g., "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)").

[0367] Specific behavior:

[0368] The generated restaurant list is sent to the terminal.

[0369] The device's UI module parses the list and updates the user interface.

[0370] As a result, the interface will display information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[0371] (Application example 2)

[0372] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0373] In recent years, an increasing number of systems analyze natural language input from users and provide relevant information to improve user experience. However, conventional systems often do not fully utilize emotion recognition or location information, making it difficult to provide services that respond to individual users' situations and emotions. Therefore, there is a need to develop systems that take user emotions and location information into account and provide more personalized information in real time.

[0374] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, generation means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, emotion recognition means for recognizing emotions from the user's voice input and facial expressions, and means for customizing suggestions based on the emotions recognized by the emotion recognition means. This makes it possible to provide more personalized information based on the user's real-time emotions and location information.

[0375] - "Natural language input" refers to users asking questions or giving instructions in the language they normally use.

[0376] "Natural language processing means" refers to technology that analyzes the natural language entered by the user and extracts its meaning and intent.

[0377] "Generation means" refers to a technology that creates relevant information based on analyzed user requirements.

[0378] "User interface" refers to the display device and operating environment through which a user receives information.

[0379] "Dynamic updating" refers to instantly changing the displayed content in response to new user input or circumstances.

[0380] "Emotion recognition means" refers to technology that identifies emotions from the user's voice, facial expressions, etc.

[0381] "Customizing recommendations" refers to providing more personalized information and recommendations based on perceived emotions.

[0382] "Location Information" refers to geographic data about a user's current location.

[0383] A "database" refers to a system that systematically stores and searches information.

[0384] "Voice input" refers to a method in which a user gives instructions to a system through a voice input device such as a microphone.

[0385] "Text input" refers to the method of entering characters using a keyboard or touch panel.

[0386] "Accessible to the visually impaired" refers to a design that allows users with visual impairments to use the system.

[0387] "Hearing-impaired" refers to a design that allows users with hearing impairments to use the system.

[0388] System Overview

[0389] The system of this invention receives natural language input from users, analyzes their requests, and provides relevant information in real time. The system combines an emotion recognition unit that recognizes the user's emotions with a unit that acquires the user's location information to provide more personalized information. Furthermore, the system supports both voice and text input, and is designed to accommodate users with visual and hearing impairments.

[0390] Hardware and Software Configuration

[0391] Hardware:

[0392] Smart glasses: To display product information

[0393] Camera: To recognize product IDs (e.g., QR codes)

[0394] Microphone: To recognize your voice input

[0395] software:

[0396] OpenCV: To capture the camera stream and recognize product IDs

[0397] dlib: Face recognition and eye tracking

[0398] Transformers (Hugging Face): Natural Language Processing and Emotion Recognition

[0399] Geopy: Getting location information

[0400] SpeechRecognition: Recognizing voice input

[0401] Data processing and calculation

[0402] The server performs the following data processing and calculations.

[0403] 1. Natural Language Processing: Receiving and analyzing voice and text input from the user. This is done using the Transformers library.

[0404] 2. Product ID recognition: Analyze the input from the camera using OpenCV and recognize the product ID.

[0405] 3. Emotion Recognition: Based on the user's voice or input data, emotion recognition is used to identify emotions. This process is also performed using the Transformers library.

[0406] 4. Information generation: Generative AI generates optimal information based on analyzed requests, emotional information, and location information.

[0407] 5. Dynamic UI Updates: Update the user interface in real time to display the appropriate information.

[0408] Specific Examples

[0409] Example 1: Finding a restaurant

[0410] User: "What restaurants are nearby?" speaks through the smart glasses.

[0411] Device: The input voice data is converted to text using SpeechRecognition. The text data is then analyzed using Transformers' NLP engine. The analysis results identify that the user is searching for a restaurant.

[0412] Emotion recognition: Recognizes stress from the user's tone of voice and language.

[0413] Generative AI: Lists quiet and relaxing restaurants based on the user's location (obtained through Geopy) and emotional information.

[0414] User interface: The smart glasses display shows information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0415] Example 2: Retrieving product information

[0416] User: Browses a shelf of red wines in a brick-and-mortar store and asks aloud through smart glasses, "Tell me about this wine."

[0417] Camera: The camera in the smart glasses recognizes the product ID (e.g., QR code) using OpenCV.

[0418] Generative AI: Generates information about red wine based on the recognized product ID.

[0419] Emotion recognition: Recognizes when the user is relaxed and determines that no additional suggestions are necessary.

[0420] User Interface: Smart glasses display the information "Red wine, price: \2500, delicious red wine."

[0421] Prompt Sentence Examples

[0422] Product ID recognition

[0423] "Tell me about this product"

[0424] emotion recognition

[0425] "tired..."

[0426] As described above, the system of the present invention can provide personalized information to users in real time based on the user's natural language input, emotions, and location information.

[0427] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0428] Step 1:

[0429] The user inputs natural language through the smart glasses. In this case, the user asks questions such as "What are some restaurants nearby?" The input voice data is converted into text data by the voice input means. Specifically, voice data is acquired from the microphone and converted into text using the SpeechRecognition library. Input data: voice data, output data: text data.

[0430] Step 2:

[0431] The device uses a natural language processing engine (NLP engine) to extract intent from the user's natural language input. Specifically, it uses the Transformers library to analyze the input text data and extract intents such as "looking for a restaurant." Input data: text data, Output data: extracted intent.

[0432] Step 3:

[0433] The device sends data to a generation means for generating related information based on the analyzed intent. The generation means then sends the related information to a generative AI model (Transformers) based on the extracted intent and location information. Input data: extracted intent, location information; output data: related information.

[0434] Step 4:

[0435] The device uses the camera to recognize the IDs of surrounding products. It analyzes the camera image using the OpenCV library and obtains the product ID by reading the QR code or barcode. Input data: camera image, output data: product ID.

[0436] Step 5:

[0437] The device uses emotion recognition to customize the information generated by the generative AI model according to the user's emotions. Specifically, it uses the Transformers library to identify emotions from the user's voice tone and vocabulary, and adjusts the recommendations accordingly. Input data: voice data and text data. Output data: recognized emotions.

[0438] Step 6:

[0439] The server customizes the generated related information based on the emotional information obtained by the emotion recognition means. The generative AI model provides optimal information based on the user's emotions, such as suggesting quiet and relaxing restaurants. Input data: emotional information, related information. Output data: customized suggestions.

[0440] Step 7:

[0441] The device dynamically updates the user interface and displays information created by the generative AI model on the smart glasses display. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed in real time and provided to the user. Input data: customized suggestions, output data: dynamically updated user interface.

[0442] This series of processing steps allows users to obtain personalized information in real time, enabling them to use services more efficiently.

[0443] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0444] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0445] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0446] [Second embodiment]

[0447] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0448] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0449] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0450] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0451] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0453] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0454] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0455] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0456] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0457] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0458] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0459] overview

[0460] The present invention provides a system that accepts natural language input from users, analyzes their requests, and displays and updates the generated information in real time. The system acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0461] System configuration

[0462] The system consists of the following main components:

[0463] 1. User Interface (UI)

[0464] An interface that allows users to enter questions or requests in natural language.

[0465] 2. Natural Language Processing (NLP) Engine

[0466] Algorithms for parsing user input and extracting its intent.

[0467] 3. Generation AI

[0468] Generates appropriate information based on the request analyzed by the NLP engine.

[0469] 4. Data Management Module

[0470] The user's location information and past usage history are stored and provided to the generating AI.

[0471] 5. Display module

[0472] Dynamically update the UI to display the generated information.

[0473] Program processing

[0474] Getting User Input

[0475] The device receives natural language input from the user (e.g., "What are the nearest restaurants?"). Voice input or text input is possible through the user interface. The user's input is stored in text form and passed on to the next processing step.

[0476] Natural Language Processing

[0477] The device passes user input to a natural language processing engine to parse the request. The NLP engine extracts meaningful keywords and phrases from the input text to identify the user's intent. This process reveals a request such as "I want to find nearby restaurants."

[0478] information generation

[0479] The server sends these requests to the generation AI, which generates the most appropriate information in real time based on the user's location information obtained from the data management module. The generated information searches a database related to the user's location information to provide the most useful information for the user.

[0480] Dynamic UI Updates

[0481] The device passes the information received from the generation AI to the display module, which dynamically updates the user interface, allowing the user to instantly see the information they were looking for (e.g., a list of nearby restaurants). The same process can be repeated to provide real-time responses when the user asks for more information or other options.

[0482] Specific examples

[0483] Example 1: Finding a restaurant

[0484] user

[0485] A user types, "What restaurants are nearby?"

[0486] Terminal

[0487] It takes input at the user interface and sends it to a natural language processing engine.

[0488] NLP Engine

[0489] Parse the input to determine the intent to find "nearby restaurants."

[0490] Generation AI

[0491] Generate a list of nearby restaurants based on the user's location.

[0492] Terminal

[0493] The list is displayed in the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0494] Example 2: Restaurant details

[0495] user

[0496] User types, "Give me details about Restaurant A."

[0497] Terminal

[0498] It takes input and sends it to a natural language processing engine.

[0499] NLP Engine

[0500] Analyze the input and determine the intent to ask for "details about Restaurant A."

[0501] Generation AI

[0502] Generate detailed information about Restaurant A, such as its address, opening hours, and menu.

[0503] Terminal

[0504] Display detailed information in the user interface.

[0505] In this way, the system of the present invention smoothly carries out a series of processes, starting with the user's natural language input, through analysis, information generation, and display, thereby intuitively and efficiently providing the information the user desires.

[0506] The processing flow will be explained below.

[0507] Step 1:

[0508] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0509] Step 2:

[0510] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0511] Step 3:

[0512] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0513] Step 4:

[0514] The device sends the request content and location information analyzed by the NLP engine to the generation AI. The location information is obtained from the device's location information service.

[0515] Step 5:

[0516] The server provides the request details and location information to the AI, which then prepares to generate appropriate restaurant information based on this information.

[0517] Step 6:

[0518] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[0519] Step 7:

[0520] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0521] Step 8:

[0522] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to see a list of nearby restaurants and their details.

[0523] Step 9:

[0524] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[0525] Step 10:

[0526] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[0527] As described above, each step works in conjunction to provide users with the information they want in real time, creating an intuitive and efficient user experience.

[0528] Example 1

[0529] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0530] Conventional information provision systems have had the problem of making it difficult to quickly and accurately obtain the information users require. It has also been difficult to effectively utilize users' location information and provide information tailored to individual user needs. Furthermore, there has been a lack of effective means of providing information to users with visual and hearing impairments.

[0531] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0532] In this invention, the server includes a terminal means for receiving natural language input from a user, a natural language processing means for analyzing the user's request, a generation AI means for generating relevant information based on the analyzed request, and a display means for dynamically updating a user interface to display the generated information. This allows users to quickly and accurately obtain the information they desire, and makes it possible to effectively utilize the user's location information and provide personalized information based on their location. Furthermore, by supporting voice input and text input, information can be provided effectively to users with visual and hearing impairments.

[0533] "Terminal Means" refers to a device or interface for receiving natural language input from a user.

[0534] "Natural language processing" refers to an algorithm or system that analyzes a user's input and determines their intent.

[0535] "Generative AI means" refers to an artificial intelligence model that generates relevant information based on requests analyzed by natural language processing means.

[0536] "Display Means" means a system for visually and audibly presenting the information generated by the Generating AI Means to the user and dynamically updating the user interface.

[0537] "Data management means" refers to a system that has the function of acquiring user location information and providing it to the generating AI means.

[0538] "Database" refers to a collection of information used to search for relevant information based on a user's location.

[0539] "User interface" refers to the screen and input devices that allow a user to interact with a system.

[0540] "Voice and text input" refers to the methods by which a user inputs information into a system in the form of voice or text.

[0541] "Visually and hearing impaired users" refers to users who have difficulty obtaining information through normal means due to visual or hearing limitations.

[0542] This invention is a system that receives natural language input from users, analyzes their requests, generates relevant information, and dynamically updates the user interface. The system features the ability to obtain the user's location information and provide personalized information based on their usage history. It also supports both voice and text input to accommodate users with visual and hearing impairments.

[0543] The main components of the system are:

[0544] 1. Terminal means:

[0545] Accepts natural language input from the user. In the case of voice input, converts speech to text using speech recognition technology (e.g., Google Speech-to-Text API). Also accepts text input through the user interface.

[0546] 2. Natural Language Processing Tools:

[0547] Analyze user input and determine its intent, specifically by using a natural language processing engine (e.g., spaCy) to parse the text and extract meaningful keywords and phrases.

[0548] 3. Generation AI means:

[0549] Relevant information is generated based on the analysis results. The generation AI generates appropriate information by referencing the user's location information and usage history. OpenAI GPT-4 and other models can be used as generation AI models.

[0550] 4. Display means:

[0551] The generated information is displayed in a user interface that is dynamically updated based on the data received from the generating AI, providing the user with visual and auditory information.

[0552] As a concrete example, the following shows what happens when a user types "What restaurants are nearby?" The process obtains the user's location information and generates and displays a list of nearby restaurants. The device uses the Google Speech-to-Text API to convert speech to text, and sends this text to a natural language processing engine (spaCy) to extract keywords. The generative AI (OpenAI GPT-4) generates a list based on the user's location information and displays it on the screen.

[0553] Example prompt sentence:

[0554] Example 1: Finding a restaurant

[0555] "User is looking for nearby restaurants. GPS information is ____. Please provide a list of the nearest restaurants."

[0556] Example 2: Restaurant details

[0557] "A user is looking for more information about Restaurant A. Please provide details such as Restaurant A's address, opening hours, and menu."

[0558] This process allows users to quickly and accurately obtain the information they need, and the system can provide personalized information based on location information and is also compatible with users with visual and hearing impairments.

[0559] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0560] Step 1: Getting User Input

[0561] The device receives natural language input from the user. The user can enter voice or text through the device's user interface. For voice input, the device uses speech recognition software (e.g., a speech recognition API) to convert the speech to text. The converted text is stored in memory and passed on to the next processing step.

[0562] Specific behavior:

[0563] The user speaks into the device's microphone, "What restaurants are nearby?"

[0564] The device uses a speech recognition API to convert the speech into text, generating the text "What restaurants are nearby?"

[0565] Input and Output:

[0566] Input: User voice or text input

[0567] Output: User request in text format

[0568] Step 2: Natural Language Processing

[0569] The device sends the text input received from the user to a natural language processing engine (e.g., NLP), which analyzes the text and extracts meaningful keywords and phrases. The extracted intents and keywords are passed on to the next step.

[0570] Specific behavior:

[0571] The device sends the text "What restaurants are nearby?" to a natural language processing engine.

[0572] A natural language processing engine analyzes the text and extracts keywords such as "nearby" and "restaurant."

[0573] Input and Output:

[0574] Input: User request in text format

[0575] Output: Extracted keywords and intent

[0576] Step 3: Information generation

[0577] The server sends the user's intent, analyzed by the natural language processing engine, to the generative AI (e.g., generative AI model), which generates relevant information. It also obtains the user's location information from the data management module and provides this location information to the generative AI. Based on this information, the generative AI generates information appropriate to the user's request in real time.

[0578] Specific behavior:

[0579] The server obtains the user's location information (e.g., latitude 35.6895, longitude 139.6917) from the data management module.

[0580] The intent and location information obtained from the natural language processing engine is sent to the generation AI.

[0581] The generation AI generates a list of restaurants such as "Restaurant A (500m)" and "Restaurant B (700m)."

[0582] Input and Output:

[0583] Input: Extracted keywords and intent, user location

[0584] Output: Generated related information (e.g., a list of restaurants)

[0585] Step 4: Dynamic UI Updates

[0586] The terminal passes the output data of the generated AI received from the server to the display module, which dynamically updates the user interface based on this data and provides the user with visual and auditory information.

[0587] Specific behavior:

[0588] The terminal passes the list of restaurants received from the server to the display module.

[0589] The display module updates the user interface to show information such as "Restaurant A (Distance: 500m)" and "Restaurant B (Distance: 700m)."

[0590] Users can view a list of nearby restaurants on their device screen.

[0591] Input and Output:

[0592] Input: Generated related information (e.g., a list of restaurants)

[0593] Output: Dynamically updated user interface and visual / auditory information

[0594] (Application example 1)

[0595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0596] Conventional food delivery systems face challenges in efficiently obtaining the information users require and making it difficult to check order and delivery status in real time. They also lack an interface that is easy for visually and hearing impaired users to use. This reduces the quality of the user experience and limits the number of users.

[0597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0598] In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, means for providing related information based on the user's location information, and means for receiving natural language input in voice or text format. This makes it possible to provide a system that allows users to efficiently obtain information using natural language and check order and delivery status in real time. It also makes it possible to provide an easy-to-use interface that is compatible with users with visual and hearing impairments.

[0599] A "means for receiving natural language input from a user" is any device or software for obtaining natural language input from a user in the form of speech or text.

[0600] "Natural language processing means for analyzing user requests" refers to technologies and algorithms that analyze the natural language entered by the user and understand their intentions and requests.

[0601] The "means for generating related information based on the analyzed request" refers to a method or technology for generating necessary information based on the results of analyzing the user's request.

[0602] "Means for dynamically updating the user interface to display generated information" refers to technologies and systems that change and update the displayed content in real time to present generated information to the user on screen or via audio.

[0603] "Means for providing relevant information based on a user's location information" refers to technologies or methods for utilizing a user's current location information to provide information related to that location.

[0604] A "means for receiving natural language input in voice or text form" is a device or software that recognizes and captures natural language input by a user through voice or text.

[0605] The "means for acquiring user location information and providing it to the generating means" refers to a technology or system for acquiring user location information and passing it to the means for generating related information.

[0606] "Means having a database for searching related information" refers to a database for searching related information based on the user's location information and a method for using the database.

[0607] "Means for tracking and displaying order and delivery status in real time" refers to technologies and systems that allow users to check the fulfillment status of their orders and the progress of deliveries in real time and display them to users.

[0608] "Means having an interface to assist in ordering procedures" refers to devices or software that have an interface that provides guidance and input assistance to enable users to place orders smoothly.

[0609] As an embodiment of this invention, we present a system that allows users to smoothly use food delivery services using natural language. The system is composed of components such as a user interface, a natural language processing engine, a generative AI, a data management module, and a display module.

[0610] Hardware and software used

[0611] Hardware:

[0612] Head-mounted displays and smartphones

[0613] microphone

[0614] software:

[0615] Python

[0616] geopy library: Getting location information

[0617] requests library: API calls

[0618] transformers library: natural language processing

[0619] speech_recognition library: speech recognition

[0620] Program processing explanation

[0621] User Interface

[0622] The user interface supports both voice and text input: users can say things like "I'd like to order a pizza," and the input is captured by a microphone and converted into text.

[0623] Natural Language Processing

[0624] User requests, entered via voice or text, are analyzed by a natural language processing engine. The transformers library is used to identify the user's intent. For example, a request such as "I want to order a pizza" can be interpreted as "Find nearby restaurants that serve pizza."

[0625] Information Generation and Display

[0626] The parsed request is sent to the generation AI, which generates relevant information based on the user's location information obtained from the data management module. For example, it generates a list of pizza restaurants nearest to the user's current location. The generated information is then passed to the display module, which dynamically updates the user interface and presents it to the user.

[0627] Order and delivery status tracking

[0628] The interface supports the user with menu information and ordering procedures for the restaurant they select, and also allows them to track and display the delivery status in real time after ordering, providing users with information such as the order progress and estimated arrival time.

[0629] Examples and prompts

[0630] For example, if a user says "I'd like to order pizza" to the head-mounted display, the system will analyze the request and display a list of nearby restaurants that serve pizza. The menu of the restaurant selected from the list will also be displayed, allowing the user to immediately proceed with the ordering process. After placing an order, the user can also check the current delivery status by asking "What's the status of my order?"

[0631] Example prompt sentence:

[0632] User: I want to order a pizza.

[0633] System: Finding pizza restaurants near you...

[0634] System: Choose from Restaurant A (distance: 500m), Restaurant B (distance: 700m), or Restaurant C (distance: 900m).

[0635] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0636] Step 1:

[0637] The user provides input via voice or text. The input is in the form of natural language, such as "I'd like to order a pizza." The device takes this input and, if it's voice, converts it to text using speech recognition software. The input data is passed on to the next processing step in natural language.

[0638] Step 2:

[0639] The device passes the captured input data to a natural language processing engine, which uses the transformers library to parse the input text and determine the user's intent, which is then converted into a clear instruction: "Find nearby pizza restaurants."

[0640] Step 3:

[0641] The server sends the parsed request to the generation AI, which generates a list of pizza restaurants based on the user's location information obtained from the data management module. The generation AI uses the user's location information as coordinates to extract data on pizza restaurants in the vicinity from a database. The output data is generated as a list of pizza restaurants and is passed to the next processing step.

[0642] Step 4:

[0643] The server sends the generated list of pizza restaurants to the display module. The display module dynamically updates the user interface and presents the list of pizza restaurants to the user. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed.

[0644] Step 5:

[0645] The user selects a restaurant from the list. The user's selection information is acquired by the terminal and sent to the server. The server then retrieves the detailed menu of the selected restaurant from the database and uses a generation AI to generate interface information to assist with the ordering process.

[0646] Step 6:

[0647] The server sends the generated interface information to the display module, which displays the menu of the selected restaurant to the user and allows the user to complete the ordering process. Once the user confirms the order, the information is sent back to the server.

[0648] Step 7:

[0649] Once an order is confirmed, the server stores the order information in a database and generates delivery information. The delivery status is updated in real time.

[0650] Step 8:

[0651] When the user asks, "What's the status of my order?", the device again sends this new input to the natural language processing engine, which analyzes it and determines that it's a request to check the delivery status.

[0652] Step 9:

[0653] The server retrieves delivery status information from the database and updates the current delivery status. The information is passed to the display module, which displays the delivery status to the user. For example, statuses such as "Currently cooking," "Delivering," and "Estimated arrival time: 20 minutes later" are displayed in real time.

[0654] In this way, a system is provided that allows users to efficiently use food delivery services using natural language input.

[0655] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0656] overview

[0657] The present invention provides a system that receives natural language input from a user, analyzes the user's request, and displays and updates the generated information in real time. It also combines an emotion engine that recognizes the user's emotions. The system also acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0658] System configuration

[0659] The system consists of the following main components:

[0660] 1. User Interface (UI)

[0661] An interface that allows users to enter questions or requests in natural language.

[0662] 2. Natural Language Processing (NLP) Engine

[0663] Algorithms for parsing user input and extracting its intent.

[0664] 3. Generation AI

[0665] Generates appropriate information based on the request analyzed by the NLP engine.

[0666] 4. Data Management Module

[0667] The user's location information and past usage history are stored and provided to the generating AI.

[0668] 5. Display module

[0669] Dynamically update the UI to display the generated information.

[0670] 6. Emotion Engine

[0671] An engine that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[0672] Program processing

[0673] Getting User Input

[0674] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0675] Natural Language Processing

[0676] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0677] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0678] Emotion recognition

[0679] The device's emotion engine recognizes the user's emotions from natural language input, voice, and facial expressions. For example, it can identify the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[0680] information generation

[0681] The server sends the request content, emotion information, and location information analyzed by the NLP engine and emotion engine to the generation AI. The location information is obtained from the device's location information service.

[0682] The generative AI uses this information to generate appropriate restaurant information in real time. For example, if the user is feeling stressed, it will suggest quiet and relaxing restaurants.

[0683] Dynamic UI Updates

[0684] The device analyzes the information it receives from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants) and repeat the same process to provide real-time responses for further inquiries about details or other options.

[0685] Specific examples

[0686] Example 1: Finding a restaurant

[0687] 1. User types "What restaurants are near me?"

[0688] 2. The device sends the input to a natural language processing engine for analysis.

[0689] 3. The NLP engine identifies the intent: "Looking for nearby restaurants."

[0690] 4. The emotion engine identifies stress from the user's voice.

[0691] 5. Generative AI will generate a list of quiet and relaxing restaurants based on the user's location and emotional information.

[0692] 6. The device displays the listed information on the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0693] Example 2: Restaurant details

[0694] 1. The user types, "Tell me the details about Restaurant A."

[0695] 2. The device sends the input to a natural language processing engine for analysis.

[0696] 3. The NLP engine identifies the intent as "I want to know more about Restaurant A."

[0697] 4. The emotion engine identifies that the user is relaxed in their voice.

[0698] 5. The generation AI generates detailed information about Restaurant A, such as its address, opening hours, and menu.

[0699] 6. The device displays the detailed information in the user interface.

[0700] In this way, the system of the present invention provides the information the user wants in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[0701] The processing flow will be explained below.

[0702] Step 1:

[0703] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0704] Step 2:

[0705] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0706] Step 3:

[0707] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0708] Step 4:

[0709] The device's emotion engine recognizes the user's emotions from natural language input and voice. For example, it identifies the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[0710] Step 5:

[0711] The device sends the request content analyzed by the NLP engine, the emotional information recognized by the emotion engine, and the user's location information to the generation AI. The location information is obtained from the device's location information service.

[0712] Step 6:

[0713] The server's AI generates optimal restaurant information based on the request, emotional information, and location information received. If the user is feeling stressed, it will prioritize quiet and relaxing restaurants.

[0714] Step 7:

[0715] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[0716] Step 8:

[0717] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0718] Step 9:

[0719] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants).

[0720] Step 10:

[0721] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[0722] Step 11:

[0723] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[0724] Through these steps, the system provides users with the information they want in real time, realizing an intuitive and efficient user experience, allowing users to obtain the most appropriate information according to their emotions and circumstances.

[0725] Example 2

[0726] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0727] Conventional natural language processing systems provide only simple information without considering the user's emotions or location, making it difficult to provide appropriate information that meets the user's needs.In addition, the information provided to users with visual and hearing impairments is insufficient.

[0728] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0729] In this invention, the server includes means for recognizing a user's emotion from the user's natural language input, voice, and facial expression, means for analyzing the emotion information, and means for acquiring the user's location information and providing it to the generation means. This makes it possible to provide appropriate information based on the user's emotion and location information, and realizes a system that can also accommodate users with visual and hearing impairments.

[0730] "Natural language input" is a method by which users enter questions or requests in everyday language.

[0731] "Natural language processing" is a technology that analyzes user input and extracts their intent.

[0732] The "generator" is a function that generates related information based on the analyzed request.

[0733] "User interface" refers to a screen or operating interface that is dynamically updated to display generated information.

[0734] An "emotion engine" is a system that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[0735] "Emotional information" is data about the user's emotional state as recognized using the emotion engine.

[0736] "Location information" is data that indicates a user's current location.

[0737] "Dynamic update" refers to changing screens and data in real time as needed.

[0738] A "database" is a system for storing and searching information.

[0739] "Visually and hearing impaired users" refers to people who have limited vision or hearing, and includes special assistive devices to accommodate these users.

[0740] The present invention combines a system that receives natural language input from users, analyzes their requests, and displays and updates the generated information in real time with an emotion engine that recognizes the user's emotions. The system also has the ability to acquire the user's location information and provide relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0741] System configuration

[0742] The system consists of the following main components:

[0743] 1. User Interface (UI)

[0744] An interface that allows users to input questions or requests in natural language, supports both speech and text input, and is designed to be accessible to users with visual or hearing impairments.

[0745] 2. Natural Language Processing (NLP) Engine

[0746] It is an algorithm that analyzes user input and extracts their intent. For example, if a user types "What restaurants are nearby?", an NLP engine analyzes this input and identifies the intent as "looking for a nearby restaurant."

[0747] 3. Generation AI

[0748] It generates appropriate information based on requests analyzed by the NLP engine, and also references the emotion engine and location information to provide information that is best suited to the user's situation.

[0749] 4. Data Management Module

[0750] The user's location information and past usage history are stored and provided to the AI ​​generator. In particular, location information is obtained from the device's location information service.

[0751] 5. Display module

[0752] The UI for displaying the generated information is dynamically updated, allowing users to instantly see the information they are looking for.

[0753] 6. Emotion Engine

[0754] This engine recognizes the user's emotions from natural language input, voice, and facial expressions. The emotion engine identifies emotions such as stress or relaxation from the tone of voice and the content of the input.

[0755] Specific operation example

[0756] Example 1: Finding a restaurant

[0757] When a user types "What restaurants are nearby?", the device's voice recognition system converts the speech into text, which is then sent to a natural language processing engine. The NLP engine analyzes the input and identifies the user's intent: "I'm looking for a nearby restaurant." At the same time, the emotion engine identifies stress from the user's voice, and this information is sent to the server's generation AI. The generation AI then creates a list of "quiet, relaxing restaurants nearby" and sends the information to the device. The device dynamically updates the user interface, providing information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[0758] Example 2: Restaurant details

[0759] When a user types "Tell me more about Restaurant A," the device sends the input to a natural language processing engine for analysis. The NLP engine identifies the intent as "I want to know more about Restaurant A," and the emotion engine identifies a relaxed tone from the user's voice. The server then requests detailed information about Restaurant A, such as its address, opening hours, and menu, from the generation AI, and the generated details are sent to the device and displayed in the user interface.

[0760] Prompt Sentence Examples

[0761] "Can you recommend a quiet, relaxing restaurant nearby?"

[0762] "Please tell me more information about Restaurant A."

[0763] In this way, the system of the present invention provides the information the user desires in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[0764] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0765] Step 1: Getting User Input

[0766] A user inputs a question or request through the application's user interface. For example, the user may input "What restaurants are nearby?" by voice or text. The input may be the user's voice or text data. In the case of voice, the device's speech recognition system converts the voice to text. The converted text is "What restaurants are nearby?"

[0767] Specific behavior:

[0768] The user presses the microphone button and says, "What restaurants are nearby?"

[0769] The device's voice recognition system converts the speech into text.

[0770] The text data "What are the nearby restaurants?" is saved.

[0771] Step 2: Natural Language Processing

[0772] The device sends user input to a natural language processing (NLP) engine. The input is text data such as "What restaurants are nearby?" The NLP engine analyzes the user's input and extracts its intent. The output of the NLP engine is the analysis result, which identifies the user's intent as "looking for a nearby restaurant."

[0773] Specific behavior:

[0774] The text data "What restaurants are nearby?" is sent to the NLP engine.

[0775] An NLP engine analyzes text data and identifies intent.

[0776] The analysis result, "Looking for a nearby restaurant," is returned to the device.

[0777] Step 3: Recognize emotions

[0778] The device's emotion engine recognizes the user's emotions from their natural language input, voice, and facial expressions. The input is the user's voice and natural language input. The emotion engine identifies emotions from the tone of the voice and the content of the input. The output is emotional information that the user is "feeling stressed."

[0779] Specific behavior:

[0780] The user's voice is sent to the emotion engine.

[0781] The emotion engine analyzes the tone and pace of the voice.

[0782] Emotional information "user is feeling stressed" is stored on the device.

[0783] Step 4: Information Generation

[0784] The server sends the analyzed request content (user intent), emotional information, and location information to the generation AI. The inputs are the analysis result "looking for a nearby restaurant," emotional information "stress," and location information. The generation AI generates appropriate information based on this information. The output is the information "quiet, relaxing restaurants nearby."

[0785] Specific behavior:

[0786] The analysis result, "Looking for a nearby restaurant," emotional information, "stress," and location information are sent to the generating AI.

[0787] The generative AI creates a prompt and sends it to the AI ​​model.

[0788] The AI ​​model generates appropriate restaurant information and returns a list such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[0789] Step 5: Dynamic UI Updates

[0790] The device analyzes the information received from the generation AI and dynamically updates the user interface. The input is the generated restaurant list. The output is specific information displayed on the user interface (e.g., "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)").

[0791] Specific behavior:

[0792] The generated restaurant list is sent to the terminal.

[0793] The device's UI module parses the list and updates the user interface.

[0794] As a result, the interface will display information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[0795] (Application example 2)

[0796] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0797] In recent years, an increasing number of systems analyze natural language input from users and provide relevant information to improve user experience. However, conventional systems often do not fully utilize emotion recognition or location information, making it difficult to provide services that respond to individual users' situations and emotions. Therefore, there is a need to develop systems that take user emotions and location information into account and provide more personalized information in real time.

[0798] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, generation means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, emotion recognition means for recognizing emotions from the user's voice input and facial expressions, and means for customizing suggestions based on the emotions recognized by the emotion recognition means. This makes it possible to provide more personalized information based on the user's real-time emotions and location information.

[0799] - "Natural language input" refers to users asking questions or giving instructions in the language they normally use.

[0800] "Natural language processing means" refers to technology that analyzes the natural language entered by the user and extracts its meaning and intent.

[0801] "Generation means" refers to a technology that creates relevant information based on analyzed user requirements.

[0802] "User interface" refers to the display device and operating environment through which a user receives information.

[0803] "Dynamic updating" refers to instantly changing the displayed content in response to new user input or circumstances.

[0804] "Emotion recognition means" refers to technology that identifies emotions from the user's voice, facial expressions, etc.

[0805] "Customizing recommendations" refers to providing more personalized information and recommendations based on perceived emotions.

[0806] "Location Information" refers to geographic data about a user's current location.

[0807] A "database" refers to a system that systematically stores and searches information.

[0808] "Voice input" refers to a method in which a user gives instructions to a system through a voice input device such as a microphone.

[0809] "Text input" refers to the method of entering characters using a keyboard or touch panel.

[0810] "Accessible to the visually impaired" refers to a design that allows users with visual impairments to use the system.

[0811] "Hearing-impaired" refers to a design that allows users with hearing impairments to use the system.

[0812] System Overview

[0813] The system of this invention receives natural language input from users, analyzes their requests, and provides relevant information in real time. The system combines an emotion recognition unit that recognizes the user's emotions with a unit that acquires the user's location information to provide more personalized information. Furthermore, the system supports both voice and text input, and is designed to accommodate users with visual and hearing impairments.

[0814] Hardware and Software Configuration

[0815] Hardware:

[0816] Smart glasses: To display product information

[0817] Camera: To recognize product IDs (e.g., QR codes)

[0818] Microphone: To recognize your voice input

[0819] software:

[0820] OpenCV: To capture the camera stream and recognize product IDs

[0821] dlib: Face recognition and eye tracking

[0822] Transformers (Hugging Face): Natural Language Processing and Emotion Recognition

[0823] Geopy: Getting location information

[0824] SpeechRecognition: Recognizing voice input

[0825] Data processing and calculation

[0826] The server performs the following data processing and calculations.

[0827] 1. Natural Language Processing: Receiving and analyzing voice and text input from the user. This is done using the Transformers library.

[0828] 2. Product ID recognition: Analyze the input from the camera using OpenCV and recognize the product ID.

[0829] 3. Emotion Recognition: Based on the user's voice or input data, emotion recognition is used to identify emotions. This process is also performed using the Transformers library.

[0830] 4. Information generation: Generative AI generates optimal information based on analyzed requests, emotional information, and location information.

[0831] 5. Dynamic UI Updates: Update the user interface in real time to display the appropriate information.

[0832] Specific Examples

[0833] Example 1: Finding a restaurant

[0834] User: "What restaurants are nearby?" speaks through the smart glasses.

[0835] Device: The input voice data is converted to text using SpeechRecognition. The text data is then analyzed using Transformers' NLP engine. The analysis results identify that the user is searching for a restaurant.

[0836] Emotion recognition: Recognizes stress from the user's tone of voice and language.

[0837] Generative AI: Lists quiet and relaxing restaurants based on the user's location (obtained through Geopy) and emotional information.

[0838] User interface: The smart glasses display shows information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0839] Example 2: Retrieving product information

[0840] User: Browses a shelf of red wines in a brick-and-mortar store and asks aloud through smart glasses, "Tell me about this wine."

[0841] Camera: The camera in the smart glasses recognizes the product ID (e.g., QR code) using OpenCV.

[0842] Generative AI: Generates information about red wine based on the recognized product ID.

[0843] Emotion recognition: Recognizes when the user is relaxed and determines that no additional suggestions are necessary.

[0844] User Interface: Smart glasses display the information "Red wine, price: \2500, delicious red wine."

[0845] Prompt Sentence Examples

[0846] Product ID recognition

[0847] "Tell me about this product"

[0848] emotion recognition

[0849] "tired..."

[0850] As described above, the system of the present invention can provide personalized information to users in real time based on the user's natural language input, emotions, and location information.

[0851] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0852] Step 1:

[0853] The user inputs natural language through the smart glasses. In this case, the user asks questions such as "What are some restaurants nearby?" The input voice data is converted into text data by the voice input means. Specifically, voice data is acquired from the microphone and converted into text using the SpeechRecognition library. Input data: voice data, output data: text data.

[0854] Step 2:

[0855] The device uses a natural language processing engine (NLP engine) to extract intent from the user's natural language input. Specifically, it uses the Transformers library to analyze the input text data and extract intents such as "looking for a restaurant." Input data: text data, Output data: extracted intent.

[0856] Step 3:

[0857] The device sends data to a generation means for generating related information based on the analyzed intent. The generation means then sends the related information to a generative AI model (Transformers) based on the extracted intent and location information. Input data: extracted intent, location information; output data: related information.

[0858] Step 4:

[0859] The device uses the camera to recognize the IDs of surrounding products. It analyzes the camera image using the OpenCV library and obtains the product ID by reading the QR code or barcode. Input data: camera image, Output data: product ID.

[0860] Step 5:

[0861] The device uses emotion recognition to customize the information generated by the generative AI model according to the user's emotions. Specifically, it uses the Transformers library to identify emotions from the user's voice tone and vocabulary, and adjusts the recommendations accordingly. Input data: voice data and text data. Output data: recognized emotions.

[0862] Step 6:

[0863] The server customizes the generated related information based on the emotional information obtained by the emotion recognition means. The generative AI model provides optimal information based on the user's emotions, such as suggesting quiet and relaxing restaurants. Input data: emotional information, related information. Output data: customized suggestions.

[0864] Step 7:

[0865] The device dynamically updates the user interface and displays information created by the generative AI model on the smart glasses display. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed in real time and provided to the user. Input data: customized suggestions, output data: dynamically updated user interface.

[0866] This series of processing steps allows users to obtain personalized information in real time, enabling them to use services more efficiently.

[0867] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0868] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0869] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0870] [Third embodiment]

[0871] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0872] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0873] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0874] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0875] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0876] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0877] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0878] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0879] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0880] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0881] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0882] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0883] overview

[0884] The present invention provides a system that accepts natural language input from users, analyzes their requests, and displays and updates the generated information in real time. The system acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[0885] System configuration

[0886] The system consists of the following main components:

[0887] 1. User Interface (UI)

[0888] An interface that allows users to enter questions or requests in natural language.

[0889] 2. Natural Language Processing (NLP) Engine

[0890] Algorithms for parsing user input and extracting its intent.

[0891] 3. Generation AI

[0892] Generates appropriate information based on the request analyzed by the NLP engine.

[0893] 4. Data Management Module

[0894] The user's location information and past usage history are stored and provided to the generating AI.

[0895] 5. Display module

[0896] Dynamically update the UI to display the generated information.

[0897] Program processing

[0898] Getting User Input

[0899] The device receives natural language input from the user (e.g., "What are the nearest restaurants?"). Voice input or text input is possible through the user interface. The user's input is stored in text form and passed on to the next processing step.

[0900] Natural Language Processing

[0901] The device passes user input to a natural language processing engine to parse the request. The NLP engine extracts meaningful keywords and phrases from the input text to identify the user's intent. This process reveals a request such as "I want to find nearby restaurants."

[0902] information generation

[0903] The server sends these requests to the generation AI, which generates the most appropriate information in real time based on the user's location information obtained from the data management module. The generated information searches a database related to the user's location information to provide the most useful information for the user.

[0904] Dynamic UI Updates

[0905] The device passes the information received from the generation AI to the display module, which dynamically updates the user interface, allowing the user to instantly see the information they were looking for (e.g., a list of nearby restaurants). The same process can be repeated to provide real-time responses when the user asks for more information or other options.

[0906] Specific examples

[0907] Example 1: Finding a restaurant

[0908] user

[0909] A user types, "What restaurants are nearby?"

[0910] Terminal

[0911] It takes input at the user interface and sends it to a natural language processing engine.

[0912] NLP Engine

[0913] Parse the input to determine the intent to find "nearby restaurants."

[0914] Generation AI

[0915] Generate a list of nearby restaurants based on the user's location.

[0916] Terminal

[0917] The list is displayed in the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0918] Example 2: Restaurant details

[0919] user

[0920] User types, "Give me details about Restaurant A."

[0921] Terminal

[0922] It takes input and sends it to a natural language processing engine.

[0923] NLP Engine

[0924] Analyze the input and determine the intent to ask for "details about Restaurant A."

[0925] Generation AI

[0926] Generate detailed information about Restaurant A, such as its address, opening hours, and menu.

[0927] Terminal

[0928] Display detailed information in the user interface.

[0929] In this way, the system of the present invention smoothly carries out a series of processes, starting with the user's natural language input, through analysis, information generation, and display, thereby intuitively and efficiently providing the information the user desires.

[0930] The processing flow will be explained below.

[0931] Step 1:

[0932] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[0933] Step 2:

[0934] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[0935] Step 3:

[0936] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[0937] Step 4:

[0938] The device sends the request content and location information analyzed by the NLP engine to the generation AI. The location information is obtained from the device's location information service.

[0939] Step 5:

[0940] The server provides the request details and location information to the AI, which then prepares to generate appropriate restaurant information based on this information.

[0941] Step 6:

[0942] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[0943] Step 7:

[0944] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[0945] Step 8:

[0946] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to see a list of nearby restaurants and their details.

[0947] Step 9:

[0948] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[0949] Step 10:

[0950] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[0951] As described above, each step works in conjunction to provide users with the information they want in real time, creating an intuitive and efficient user experience.

[0952] Example 1

[0953] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0954] Conventional information provision systems have had the problem of making it difficult to quickly and accurately obtain the information users require. It has also been difficult to effectively utilize users' location information and provide information tailored to individual user needs. Furthermore, there has been a lack of effective means of providing information to users with visual and hearing impairments.

[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0956] In this invention, the server includes a terminal means for receiving natural language input from a user, a natural language processing means for analyzing the user's request, a generation AI means for generating relevant information based on the analyzed request, and a display means for dynamically updating a user interface to display the generated information. This allows users to quickly and accurately obtain the information they desire, and makes it possible to effectively utilize the user's location information and provide personalized information based on their location. Furthermore, by supporting voice input and text input, information can be provided effectively to users with visual and hearing impairments.

[0957] "Terminal Means" refers to a device or interface for receiving natural language input from a user.

[0958] "Natural language processing" refers to an algorithm or system that analyzes a user's input and determines their intent.

[0959] "Generative AI means" refers to an artificial intelligence model that generates relevant information based on requests analyzed by natural language processing means.

[0960] "Display Means" means a system for visually and audibly presenting the information generated by the Generating AI Means to the user and dynamically updating the user interface.

[0961] "Data management means" refers to a system that has the function of acquiring user location information and providing it to the generating AI means.

[0962] "Database" refers to a collection of information used to search for relevant information based on a user's location.

[0963] "User interface" refers to the screen and input devices that allow a user to interact with a system.

[0964] "Voice and text input" refers to the methods by which a user inputs information into a system in the form of voice or text.

[0965] "Visually and hearing impaired users" refers to users who have difficulty obtaining information through normal means due to visual or hearing limitations.

[0966] This invention is a system that receives natural language input from users, analyzes their requests, generates relevant information, and dynamically updates the user interface. The system features the ability to obtain the user's location information and provide personalized information based on their usage history. It also supports both voice and text input to accommodate users with visual and hearing impairments.

[0967] The main components of the system are:

[0968] 1. Terminal means:

[0969] Accepts natural language input from the user. In the case of voice input, converts speech to text using speech recognition technology (e.g., Google Speech-to-Text API). Also accepts text input through the user interface.

[0970] 2. Natural Language Processing Tools:

[0971] Analyze user input and determine its intent, specifically by using a natural language processing engine (e.g., spaCy) to parse the text and extract meaningful keywords and phrases.

[0972] 3. Generation AI means:

[0973] Relevant information is generated based on the analysis results. The generation AI generates appropriate information by referencing the user's location information and usage history. OpenAI GPT-4 and other models can be used as generation AI models.

[0974] 4. Display means:

[0975] The generated information is displayed in a user interface that is dynamically updated based on the data received from the generating AI, providing the user with visual and auditory information.

[0976] As a concrete example, the following shows what happens when a user types "What restaurants are nearby?" The process obtains the user's location information and generates and displays a list of nearby restaurants. The device uses the Google Speech-to-Text API to convert speech to text, and sends this text to a natural language processing engine (spaCy) to extract keywords. The generative AI (OpenAI GPT-4) generates a list based on the user's location information and displays it on the screen.

[0977] Example prompt sentence:

[0978] Example 1: Finding a restaurant

[0979] "User is looking for nearby restaurants. GPS information is ____. Please provide a list of the nearest restaurants."

[0980] Example 2: Restaurant details

[0981] "A user is looking for more information about Restaurant A. Please provide details such as Restaurant A's address, opening hours, and menu."

[0982] This process allows users to quickly and accurately obtain the information they need, and the system can provide personalized information based on location information and is also compatible with users with visual and hearing impairments.

[0983] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0984] Step 1: Getting User Input

[0985] The device receives natural language input from the user. The user can enter voice or text through the device's user interface. For voice input, the device uses speech recognition software (e.g., a speech recognition API) to convert the speech to text. The converted text is stored in memory and passed on to the next processing step.

[0986] Specific behavior:

[0987] The user speaks into the device's microphone, "What restaurants are nearby?"

[0988] The device uses a speech recognition API to convert the speech into text, generating the text "What restaurants are nearby?"

[0989] Input and Output:

[0990] Input: User voice or text input

[0991] Output: User request in text format

[0992] Step 2: Natural Language Processing

[0993] The device sends the text input received from the user to a natural language processing engine (e.g., NLP), which analyzes the text and extracts meaningful keywords and phrases. The extracted intents and keywords are passed on to the next step.

[0994] Specific behavior:

[0995] The device sends the text "What restaurants are nearby?" to a natural language processing engine.

[0996] A natural language processing engine analyzes the text and extracts keywords such as "nearby" and "restaurant."

[0997] Input and Output:

[0998] Input: User request in text format

[0999] Output: Extracted keywords and intent

[1000] Step 3: Information generation

[1001] The server sends the user's intent, analyzed by the natural language processing engine, to the generative AI (e.g., generative AI model), which generates relevant information. It also obtains the user's location information from the data management module and provides this location information to the generative AI. Based on this information, the generative AI generates information appropriate to the user's request in real time.

[1002] Specific behavior:

[1003] The server obtains the user's location information (e.g., latitude 35.6895, longitude 139.6917) from the data management module.

[1004] The intent and location information obtained from the natural language processing engine is sent to the generation AI.

[1005] The generation AI generates a list of restaurants such as "Restaurant A (500m)" and "Restaurant B (700m)."

[1006] Input and Output:

[1007] Input: Extracted keywords and intent, user location

[1008] Output: Generated related information (e.g., a list of restaurants)

[1009] Step 4: Dynamic UI Updates

[1010] The terminal passes the output data of the generated AI received from the server to the display module, which dynamically updates the user interface based on this data and provides the user with visual and auditory information.

[1011] Specific behavior:

[1012] The terminal passes the list of restaurants received from the server to the display module.

[1013] The display module updates the user interface to show information such as "Restaurant A (Distance: 500m)" and "Restaurant B (Distance: 700m)."

[1014] Users can view a list of nearby restaurants on their device screen.

[1015] Input and Output:

[1016] Input: Generated related information (e.g., a list of restaurants)

[1017] Output: Dynamically updated user interface and visual / auditory information

[1018] (Application example 1)

[1019] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1020] Conventional food delivery systems face challenges in efficiently obtaining the information users require and making it difficult to check order and delivery status in real time. They also lack an interface that is easy for visually and hearing impaired users to use. This reduces the quality of the user experience and limits the number of users.

[1021] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1022] In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, means for providing related information based on the user's location information, and means for receiving natural language input in voice or text format. This makes it possible to provide a system that allows users to efficiently obtain information using natural language and check order and delivery status in real time. It also makes it possible to provide an easy-to-use interface that is compatible with users with visual and hearing impairments.

[1023] A "means for receiving natural language input from a user" is any device or software for obtaining natural language input from a user in the form of speech or text.

[1024] "Natural language processing means for analyzing user requests" refers to technologies and algorithms that analyze the natural language entered by the user and understand their intentions and requests.

[1025] The "means for generating related information based on the analyzed request" refers to a method or technology for generating necessary information based on the results of analyzing the user's request.

[1026] "Means for dynamically updating the user interface to display generated information" refers to technologies and systems that change and update the displayed content in real time to present generated information to the user on screen or via audio.

[1027] "Means for providing relevant information based on a user's location information" refers to technologies or methods for utilizing a user's current location information to provide information related to that location.

[1028] A "means for receiving natural language input in voice or text form" is a device or software that recognizes and captures natural language input by a user through voice or text.

[1029] The "means for acquiring user location information and providing it to the generating means" refers to a technology or system for acquiring user location information and passing it to the means for generating related information.

[1030] "Means having a database for searching related information" refers to a database for searching related information based on the user's location information and a method for using the database.

[1031] "Means for tracking and displaying order and delivery status in real time" refers to technologies and systems that allow users to check the fulfillment status of their orders and the progress of deliveries in real time and display them to users.

[1032] "Means having an interface to assist in ordering procedures" refers to devices or software that have an interface that provides guidance and input assistance to enable users to place orders smoothly.

[1033] As an embodiment of this invention, we present a system that allows users to smoothly use food delivery services using natural language. The system is composed of components such as a user interface, a natural language processing engine, a generative AI, a data management module, and a display module.

[1034] Hardware and software used

[1035] Hardware:

[1036] Head-mounted displays and smartphones

[1037] microphone

[1038] software:

[1039] Python

[1040] geopy library: Getting location information

[1041] requests library: API calls

[1042] transformers library: natural language processing

[1043] speech_recognition library: speech recognition

[1044] Program processing explanation

[1045] User Interface

[1046] The user interface supports both voice and text input: users can say things like "I'd like to order a pizza," and the input is captured by a microphone and converted into text.

[1047] Natural Language Processing

[1048] User requests, entered via voice or text, are analyzed by a natural language processing engine. The transformers library is used to identify the user's intent. For example, a request such as "I want to order a pizza" can be interpreted as "Find nearby restaurants that serve pizza."

[1049] Information Generation and Display

[1050] The parsed request is sent to the generation AI, which generates relevant information based on the user's location information obtained from the data management module. For example, it generates a list of pizza restaurants nearest to the user's current location. The generated information is then passed to the display module, which dynamically updates the user interface and presents it to the user.

[1051] Order and delivery status tracking

[1052] The interface supports the user with menu information and ordering procedures for the restaurant they select, and also allows them to track and display the delivery status in real time after ordering, providing users with information such as the order progress and estimated arrival time.

[1053] Examples and prompts

[1054] For example, if a user says "I'd like to order pizza" to the head-mounted display, the system will analyze the request and display a list of nearby restaurants that serve pizza. The menu of the restaurant selected from the list will also be displayed, allowing the user to immediately proceed with the ordering process. After placing an order, the user can also check the current delivery status by asking "What's the status of my order?"

[1055] Example prompt sentence:

[1056] User: I want to order a pizza.

[1057] System: Finding pizza restaurants near you...

[1058] System: Choose from Restaurant A (distance: 500m), Restaurant B (distance: 700m), or Restaurant C (distance: 900m).

[1059] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1060] Step 1:

[1061] The user provides input via voice or text. The input is in the form of natural language, such as "I'd like to order a pizza." The device takes this input and, if it's voice, converts it to text using speech recognition software. The input data is passed on to the next processing step in natural language.

[1062] Step 2:

[1063] The device passes the captured input data to a natural language processing engine, which uses the transformers library to parse the input text and determine the user's intent, which is then converted into a clear instruction: "Find nearby pizza restaurants."

[1064] Step 3:

[1065] The server sends the parsed request to the generation AI, which generates a list of pizza restaurants based on the user's location information obtained from the data management module. The generation AI uses the user's location information as coordinates to extract data on pizza restaurants in the vicinity from a database. The output data is generated as a list of pizza restaurants and is passed to the next processing step.

[1066] Step 4:

[1067] The server sends the generated list of pizza restaurants to the display module. The display module dynamically updates the user interface and presents the list of pizza restaurants to the user. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed.

[1068] Step 5:

[1069] The user selects a restaurant from the list. The user's selection information is acquired by the terminal and sent to the server. The server then retrieves the detailed menu of the selected restaurant from the database and uses a generation AI to generate interface information to assist with the ordering process.

[1070] Step 6:

[1071] The server sends the generated interface information to the display module, which displays the menu of the selected restaurant to the user and allows the user to complete the ordering process. Once the user confirms the order, the information is sent back to the server.

[1072] Step 7:

[1073] Once an order is confirmed, the server stores the order information in a database and generates delivery information. The delivery status is updated in real time.

[1074] Step 8:

[1075] When the user asks, "What's the status of my order?", the device again sends this new input to the natural language processing engine, which analyzes it and determines that it's a request to check the delivery status.

[1076] Step 9:

[1077] The server retrieves delivery status information from the database and updates the current delivery status. The information is passed to the display module, which displays the delivery status to the user. For example, statuses such as "Currently cooking," "Delivering," and "Estimated arrival time: 20 minutes later" are displayed in real time.

[1078] In this way, a system is provided that allows users to efficiently use food delivery services using natural language input.

[1079] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1080] overview

[1081] The present invention provides a system that receives natural language input from a user, analyzes the user's request, and displays and updates the generated information in real time. It also combines an emotion engine that recognizes the user's emotions. The system also acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[1082] System configuration

[1083] The system consists of the following main components:

[1084] 1. User Interface (UI)

[1085] An interface that allows users to enter questions or requests in natural language.

[1086] 2. Natural Language Processing (NLP) Engine

[1087] Algorithms for parsing user input and extracting its intent.

[1088] 3. Generation AI

[1089] Generates appropriate information based on the request analyzed by the NLP engine.

[1090] 4. Data Management Module

[1091] The user's location information and past usage history are stored and provided to the generating AI.

[1092] 5. Display module

[1093] Dynamically update the UI to display the generated information.

[1094] 6. Emotion Engine

[1095] An engine that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[1096] Program processing

[1097] Getting User Input

[1098] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[1099] Natural Language Processing

[1100] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[1101] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[1102] Emotion recognition

[1103] The device's emotion engine recognizes the user's emotions from natural language input, voice, and facial expressions. For example, it can identify the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[1104] information generation

[1105] The server sends the request content, emotion information, and location information analyzed by the NLP engine and emotion engine to the generation AI. The location information is obtained from the device's location information service.

[1106] The generative AI uses this information to generate appropriate restaurant information in real time. For example, if the user is feeling stressed, it will suggest quiet and relaxing restaurants.

[1107] Dynamic UI Updates

[1108] The device analyzes the information it receives from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants) and repeat the same process to provide real-time responses for further inquiries about details or other options.

[1109] Specific examples

[1110] Example 1: Finding a restaurant

[1111] 1. User types "What restaurants are near me?"

[1112] 2. The device sends the input to a natural language processing engine for analysis.

[1113] 3. The NLP engine identifies the intent: "Looking for nearby restaurants."

[1114] 4. The emotion engine identifies stress from the user's voice.

[1115] 5. Generative AI will generate a list of quiet and relaxing restaurants based on the user's location and emotional information.

[1116] 6. The device displays the listed information on the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1117] Example 2: Restaurant details

[1118] 1. The user types, "Tell me the details about Restaurant A."

[1119] 2. The device sends the input to a natural language processing engine for analysis.

[1120] 3. The NLP engine identifies the intent as "I want to know more about Restaurant A."

[1121] 4. The emotion engine identifies that the user is relaxed in their voice.

[1122] 5. The generation AI generates detailed information about Restaurant A, such as its address, opening hours, and menu.

[1123] 6. The device displays the detailed information in the user interface.

[1124] In this way, the system of the present invention provides the information the user wants in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[1125] The processing flow will be explained below.

[1126] Step 1:

[1127] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[1128] Step 2:

[1129] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[1130] Step 3:

[1131] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[1132] Step 4:

[1133] The device's emotion engine recognizes the user's emotions from natural language input and voice. For example, it identifies the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[1134] Step 5:

[1135] The device sends the request content analyzed by the NLP engine, the emotional information recognized by the emotion engine, and the user's location information to the generation AI. The location information is obtained from the device's location information service.

[1136] Step 6:

[1137] The server's AI generates optimal restaurant information based on the request, emotional information, and location information received. If the user is feeling stressed, it will prioritize quiet and relaxing restaurants.

[1138] Step 7:

[1139] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[1140] Step 8:

[1141] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1142] Step 9:

[1143] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants).

[1144] Step 10:

[1145] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[1146] Step 11:

[1147] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[1148] Through these steps, the system provides users with the information they want in real time, realizing an intuitive and efficient user experience, allowing users to obtain the most appropriate information according to their emotions and circumstances.

[1149] Example 2

[1150] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1151] Conventional natural language processing systems provide only simple information without considering the user's emotions or location, making it difficult to provide appropriate information that meets the user's needs.In addition, the information provided to users with visual and hearing impairments is insufficient.

[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1153] In this invention, the server includes means for recognizing a user's emotion from the user's natural language input, voice, and facial expression, means for analyzing the emotion information, and means for acquiring the user's location information and providing it to the generation means. This makes it possible to provide appropriate information based on the user's emotion and location information, and realizes a system that can also accommodate users with visual and hearing impairments.

[1154] "Natural language input" is a method by which users enter questions or requests in everyday language.

[1155] "Natural language processing" is a technology that analyzes user input and extracts their intent.

[1156] The "generator" is a function that generates related information based on the analyzed request.

[1157] "User interface" refers to a screen or operating interface that is dynamically updated to display generated information.

[1158] An "emotion engine" is a system that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[1159] "Emotional information" is data about the user's emotional state as recognized using the emotion engine.

[1160] "Location information" is data that indicates a user's current location.

[1161] "Dynamic update" refers to changing screens and data in real time as needed.

[1162] A "database" is a system for storing and searching information.

[1163] "Visually and hearing impaired users" refers to people who have limited vision or hearing, and includes special assistive devices to accommodate these users.

[1164] The present invention combines a system that receives natural language input from users, analyzes their requests, and displays and updates the generated information in real time with an emotion engine that recognizes the user's emotions. The system also has the ability to acquire the user's location information and provide relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[1165] System configuration

[1166] The system consists of the following main components:

[1167] 1. User Interface (UI)

[1168] An interface that allows users to input questions or requests in natural language, supports both speech and text input, and is designed to be accessible to users with visual or hearing impairments.

[1169] 2. Natural Language Processing (NLP) Engine

[1170] It is an algorithm that analyzes user input and extracts their intent. For example, if a user types "What restaurants are nearby?", an NLP engine analyzes this input and identifies the intent as "looking for a nearby restaurant."

[1171] 3. Generation AI

[1172] It generates appropriate information based on requests analyzed by the NLP engine, and also references the emotion engine and location information to provide information that is best suited to the user's situation.

[1173] 4. Data Management Module

[1174] The user's location information and past usage history are stored and provided to the AI ​​generator. In particular, location information is obtained from the device's location information service.

[1175] 5. Display module

[1176] The UI for displaying the generated information is dynamically updated, allowing users to instantly see the information they are looking for.

[1177] 6. Emotion Engine

[1178] This engine recognizes the user's emotions from natural language input, voice, and facial expressions. The emotion engine identifies emotions such as stress or relaxation from the tone of voice and the content of the input.

[1179] Specific operation example

[1180] Example 1: Finding a restaurant

[1181] When a user types "What restaurants are nearby?", the device's voice recognition system converts the speech into text, which is then sent to a natural language processing engine. The NLP engine analyzes the input and identifies the user's intent: "I'm looking for a nearby restaurant." At the same time, the emotion engine identifies stress from the user's voice, and this information is sent to the server's generation AI. The generation AI then creates a list of "quiet, relaxing restaurants nearby" and sends the information to the device. The device dynamically updates the user interface, providing information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[1182] Example 2: Restaurant details

[1183] When a user types "Tell me more about Restaurant A," the device sends the input to a natural language processing engine for analysis. The NLP engine identifies the intent as "I want to know more about Restaurant A," and the emotion engine identifies a relaxed tone from the user's voice. The server then requests detailed information about Restaurant A, such as its address, opening hours, and menu, from the generation AI, and the generated details are sent to the device and displayed in the user interface.

[1184] Prompt Sentence Examples

[1185] "Can you recommend a quiet, relaxing restaurant nearby?"

[1186] "Please tell me more information about Restaurant A."

[1187] In this way, the system of the present invention provides the information the user desires in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[1188] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1189] Step 1: Getting User Input

[1190] A user inputs a question or request through the application's user interface. For example, the user may input "What restaurants are nearby?" by voice or text. The input may be the user's voice or text data. In the case of voice, the device's speech recognition system converts the voice to text. The converted text is "What restaurants are nearby?"

[1191] Specific behavior:

[1192] The user presses the microphone button and says, "What restaurants are nearby?"

[1193] The device's voice recognition system converts the speech into text.

[1194] The text data "What are the nearby restaurants?" is saved.

[1195] Step 2: Natural Language Processing

[1196] The device sends user input to a natural language processing (NLP) engine. The input is text data such as "What restaurants are nearby?" The NLP engine analyzes the user's input and extracts its intent. The output of the NLP engine is the analysis result, which identifies the user's intent as "looking for a nearby restaurant."

[1197] Specific behavior:

[1198] The text data "What restaurants are nearby?" is sent to the NLP engine.

[1199] An NLP engine analyzes text data and identifies intent.

[1200] The analysis result, "Looking for a nearby restaurant," is returned to the device.

[1201] Step 3: Recognize emotions

[1202] The device's emotion engine recognizes the user's emotions from their natural language input, voice, and facial expressions. The input is the user's voice and natural language input. The emotion engine identifies emotions from the tone of the voice and the content of the input. The output is emotional information that the user is "feeling stressed."

[1203] Specific behavior:

[1204] The user's voice is sent to the emotion engine.

[1205] The emotion engine analyzes the tone and pace of the voice.

[1206] Emotional information "user is feeling stressed" is stored on the device.

[1207] Step 4: Information Generation

[1208] The server sends the analyzed request content (user intent), emotional information, and location information to the generation AI. The inputs are the analysis result "looking for a nearby restaurant," emotional information "stress," and location information. The generation AI generates appropriate information based on this information. The output is the information "quiet, relaxing restaurants nearby."

[1209] Specific behavior:

[1210] The analysis result, "Looking for a nearby restaurant," emotional information, "stress," and location information are sent to the generating AI.

[1211] The generative AI creates a prompt and sends it to the AI ​​model.

[1212] The AI ​​model generates appropriate restaurant information and returns a list such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[1213] Step 5: Dynamic UI Updates

[1214] The device analyzes the information received from the generation AI and dynamically updates the user interface. The input is the generated restaurant list. The output is specific information displayed on the user interface (e.g., "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)").

[1215] Specific behavior:

[1216] The generated restaurant list is sent to the terminal.

[1217] The device's UI module parses the list and updates the user interface.

[1218] As a result, the interface will display information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[1219] (Application example 2)

[1220] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1221] In recent years, an increasing number of systems analyze natural language input from users and provide relevant information to improve user experience. However, conventional systems often do not fully utilize emotion recognition or location information, making it difficult to provide services that respond to individual users' situations and emotions. Therefore, there is a need to develop systems that take user emotions and location information into account and provide more personalized information in real time.

[1222] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, generation means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, emotion recognition means for recognizing emotions from the user's voice input and facial expressions, and means for customizing suggestions based on the emotions recognized by the emotion recognition means. This makes it possible to provide more personalized information based on the user's real-time emotions and location information.

[1223] - "Natural language input" refers to users asking questions or giving instructions in the language they normally use.

[1224] "Natural language processing means" refers to technology that analyzes the natural language entered by the user and extracts its meaning and intent.

[1225] "Generation means" refers to a technology that creates relevant information based on analyzed user requirements.

[1226] "User interface" refers to the display device and operating environment through which a user receives information.

[1227] "Dynamic updating" refers to instantly changing the displayed content in response to new user input or circumstances.

[1228] "Emotion recognition means" refers to technology that identifies emotions from the user's voice, facial expressions, etc.

[1229] "Customizing recommendations" refers to providing more personalized information and recommendations based on perceived emotions.

[1230] "Location Information" refers to geographic data about a user's current location.

[1231] A "database" refers to a system that systematically stores and searches information.

[1232] "Voice input" refers to a method in which a user gives instructions to a system through a voice input device such as a microphone.

[1233] "Text input" refers to the method of entering characters using a keyboard or touch panel.

[1234] "Accessible to the visually impaired" refers to a design that allows users with visual impairments to use the system.

[1235] "Hearing-impaired" refers to a design that allows users with hearing impairments to use the system.

[1236] System Overview

[1237] The system of this invention receives natural language input from users, analyzes their requests, and provides relevant information in real time. The system combines an emotion recognition unit that recognizes the user's emotions with a unit that acquires the user's location information to provide more personalized information. Furthermore, the system supports both voice and text input, and is designed to accommodate users with visual and hearing impairments.

[1238] Hardware and Software Configuration

[1239] Hardware:

[1240] Smart glasses: To display product information

[1241] Camera: To recognize product IDs (e.g., QR codes)

[1242] Microphone: To recognize your voice input

[1243] software:

[1244] OpenCV: To capture the camera stream and recognize product IDs

[1245] dlib: Face recognition and eye tracking

[1246] Transformers (Hugging Face): Natural Language Processing and Emotion Recognition

[1247] Geopy: Getting location information

[1248] SpeechRecognition: Recognizing voice input

[1249] Data processing and calculation

[1250] The server performs the following data processing and calculations.

[1251] 1. Natural Language Processing: Receiving and analyzing voice and text input from the user. This is done using the Transformers library.

[1252] 2. Product ID recognition: Analyze the input from the camera using OpenCV and recognize the product ID.

[1253] 3. Emotion Recognition: Based on the user's voice or input data, emotion recognition is used to identify emotions. This process is also performed using the Transformers library.

[1254] 4. Information generation: Generative AI generates optimal information based on analyzed requests, emotional information, and location information.

[1255] 5. Dynamic UI Updates: Update the user interface in real time to display the appropriate information.

[1256] Specific Examples

[1257] Example 1: Finding a restaurant

[1258] User: "What restaurants are nearby?" speaks through the smart glasses.

[1259] Device: The input voice data is converted to text using SpeechRecognition. The text data is then analyzed using Transformers' NLP engine. The analysis results identify that the user is searching for a restaurant.

[1260] Emotion recognition: Recognizes stress from the user's tone of voice and language.

[1261] Generative AI: Lists quiet and relaxing restaurants based on the user's location (obtained through Geopy) and emotional information.

[1262] User interface: The smart glasses display shows information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1263] Example 2: Retrieving product information

[1264] User: Browses a shelf of red wines in a brick-and-mortar store and asks aloud through smart glasses, "Tell me about this wine."

[1265] Camera: The camera in the smart glasses recognizes the product ID (e.g., QR code) using OpenCV.

[1266] Generative AI: Generates information about red wine based on the recognized product ID.

[1267] Emotion recognition: Recognizes when the user is relaxed and determines that no additional suggestions are necessary.

[1268] User Interface: Smart glasses display the information "Red wine, price: \2500, delicious red wine."

[1269] Prompt Sentence Examples

[1270] Product ID recognition

[1271] "Tell me about this product"

[1272] emotion recognition

[1273] "tired..."

[1274] As described above, the system of the present invention can provide personalized information to users in real time based on the user's natural language input, emotions, and location information.

[1275] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1276] Step 1:

[1277] The user inputs natural language through the smart glasses. In this case, the user asks questions such as "What are some restaurants nearby?" The input voice data is converted into text data by the voice input means. Specifically, voice data is acquired from the microphone and converted into text using the SpeechRecognition library. Input data: voice data, output data: text data.

[1278] Step 2:

[1279] The device uses a natural language processing engine (NLP engine) to extract intent from the user's natural language input. Specifically, it uses the Transformers library to analyze the input text data and extract intents such as "looking for a restaurant." Input data: text data, Output data: extracted intent.

[1280] Step 3:

[1281] The device sends data to a generation means for generating related information based on the analyzed intent. The generation means then sends the related information to a generative AI model (Transformers) based on the extracted intent and location information. Input data: extracted intent, location information; output data: related information.

[1282] Step 4:

[1283] The device uses the camera to recognize the IDs of surrounding products. It analyzes the camera image using the OpenCV library and obtains the product ID by reading the QR code or barcode. Input data: camera image, Output data: product ID.

[1284] Step 5:

[1285] The device uses emotion recognition to customize the information generated by the generative AI model according to the user's emotions. Specifically, it uses the Transformers library to identify emotions from the user's voice tone and vocabulary, and adjusts the recommendations accordingly. Input data: voice data and text data. Output data: recognized emotions.

[1286] Step 6:

[1287] The server customizes the generated related information based on the emotional information obtained by the emotion recognition means. The generative AI model provides optimal information based on the user's emotions, such as suggesting quiet and relaxing restaurants. Input data: emotional information, related information. Output data: customized suggestions.

[1288] Step 7:

[1289] The device dynamically updates the user interface and displays information created by the generative AI model on the smart glasses display. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed in real time and provided to the user. Input data: customized suggestions, output data: dynamically updated user interface.

[1290] This series of processing steps allows users to obtain personalized information in real time, enabling them to use services more efficiently.

[1291] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1292] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1293] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1294] [Fourth embodiment]

[1295] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1296] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1297] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1298] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1299] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1300] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1301] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1302] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1303] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1304] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1305] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1306] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1307] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1308] overview

[1309] The present invention provides a system that accepts natural language input from users, analyzes their requests, and displays and updates the generated information in real time. The system acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[1310] System configuration

[1311] The system consists of the following main components:

[1312] 1. User Interface (UI)

[1313] An interface that allows users to enter questions or requests in natural language.

[1314] 2. Natural Language Processing (NLP) Engine

[1315] Algorithms for parsing user input and extracting its intent.

[1316] 3. Generation AI

[1317] Generates appropriate information based on the request analyzed by the NLP engine.

[1318] 4. Data Management Module

[1319] The user's location information and past usage history are stored and provided to the generating AI.

[1320] 5. Display module

[1321] Dynamically update the UI to display the generated information.

[1322] Program processing

[1323] Getting User Input

[1324] The device receives natural language input from the user (e.g., "What are the nearest restaurants?"). Voice input or text input is possible through the user interface. The user's input is stored in text form and passed on to the next processing step.

[1325] Natural Language Processing

[1326] The device passes user input to a natural language processing engine to parse the request. The NLP engine extracts meaningful keywords and phrases from the input text to identify the user's intent. This process reveals a request such as "I want to find nearby restaurants."

[1327] information generation

[1328] The server sends these requests to the generation AI, which generates the most appropriate information in real time based on the user's location information obtained from the data management module. The generated information searches a database related to the user's location information to provide the most useful information for the user.

[1329] Dynamic UI Updates

[1330] The device passes the information received from the generation AI to the display module, which dynamically updates the user interface, allowing the user to instantly see the information they were looking for (e.g., a list of nearby restaurants). The same process can be repeated to provide real-time responses when the user asks for more information or other options.

[1331] Specific examples

[1332] Example 1: Finding a restaurant

[1333] user

[1334] A user types, "What restaurants are nearby?"

[1335] Terminal

[1336] It takes input at the user interface and sends it to a natural language processing engine.

[1337] NLP Engine

[1338] Parse the input to determine the intent to find "nearby restaurants."

[1339] Generation AI

[1340] Generate a list of nearby restaurants based on the user's location.

[1341] Terminal

[1342] The list is displayed in the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1343] Example 2: Restaurant details

[1344] user

[1345] User types, "Give me details about Restaurant A."

[1346] Terminal

[1347] It takes input and sends it to a natural language processing engine.

[1348] NLP Engine

[1349] Analyze the input and determine the intent to ask for "details about Restaurant A."

[1350] Generation AI

[1351] Generate detailed information about Restaurant A, such as its address, opening hours, and menu.

[1352] Terminal

[1353] Display detailed information in the user interface.

[1354] In this way, the system of the present invention smoothly carries out a series of processes, starting with the user's natural language input, through analysis, information generation, and display, thereby intuitively and efficiently providing the information the user desires.

[1355] The processing flow will be explained below.

[1356] Step 1:

[1357] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[1358] Step 2:

[1359] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[1360] Step 3:

[1361] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[1362] Step 4:

[1363] The device sends the request content and location information analyzed by the NLP engine to the generation AI. The location information is obtained from the device's location information service.

[1364] Step 5:

[1365] The server provides the request details and location information to the AI, which then prepares to generate appropriate restaurant information based on this information.

[1366] Step 6:

[1367] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[1368] Step 7:

[1369] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1370] Step 8:

[1371] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to see a list of nearby restaurants and their details.

[1372] Step 9:

[1373] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[1374] Step 10:

[1375] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[1376] As described above, each step works in conjunction to provide users with the information they want in real time, creating an intuitive and efficient user experience.

[1377] Example 1

[1378] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1379] Conventional information provision systems have had the problem of making it difficult to quickly and accurately obtain the information users require. It has also been difficult to effectively utilize users' location information and provide information tailored to individual user needs. Furthermore, there has been a lack of effective means of providing information to users with visual and hearing impairments.

[1380] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1381] In this invention, the server includes a terminal means for receiving natural language input from a user, a natural language processing means for analyzing the user's request, a generation AI means for generating relevant information based on the analyzed request, and a display means for dynamically updating a user interface to display the generated information. This allows users to quickly and accurately obtain the information they desire, and makes it possible to effectively utilize the user's location information and provide personalized information based on their location. Furthermore, by supporting voice input and text input, information can be provided effectively to users with visual and hearing impairments.

[1382] "Terminal Means" refers to a device or interface for receiving natural language input from a user.

[1383] "Natural language processing" refers to an algorithm or system that analyzes a user's input and determines their intent.

[1384] "Generative AI means" refers to an artificial intelligence model that generates relevant information based on requests analyzed by natural language processing means.

[1385] "Display Means" means a system for visually and audibly presenting the information generated by the Generating AI Means to the user and dynamically updating the user interface.

[1386] "Data management means" refers to a system that has the function of acquiring user location information and providing it to the generating AI means.

[1387] "Database" refers to a collection of information used to search for relevant information based on a user's location.

[1388] "User interface" refers to the screen and input devices that allow a user to interact with a system.

[1389] "Voice and text input" refers to the methods by which a user inputs information into a system in the form of voice or text.

[1390] "Visually and hearing impaired users" refers to users who have difficulty obtaining information through normal means due to visual or hearing limitations.

[1391] This invention is a system that receives natural language input from users, analyzes their requests, generates relevant information, and dynamically updates the user interface. The system features the ability to obtain the user's location information and provide personalized information based on their usage history. It also supports both voice and text input to accommodate users with visual and hearing impairments.

[1392] The main components of the system are:

[1393] 1. Terminal means:

[1394] Accepts natural language input from the user. In the case of voice input, converts speech to text using speech recognition technology (e.g., Google Speech-to-Text API). Also accepts text input through the user interface.

[1395] 2. Natural Language Processing Tools:

[1396] Analyze user input and determine its intent, specifically by using a natural language processing engine (e.g., spaCy) to parse the text and extract meaningful keywords and phrases.

[1397] 3. Generation AI means:

[1398] Relevant information is generated based on the analysis results. The generation AI generates appropriate information by referencing the user's location information and usage history. OpenAI GPT-4 and other models can be used as generation AI models.

[1399] 4. Display means:

[1400] The generated information is displayed in a user interface that is dynamically updated based on the data received from the generating AI, providing the user with visual and auditory information.

[1401] As a concrete example, the following shows what happens when a user types "What restaurants are nearby?" The process obtains the user's location information and generates and displays a list of nearby restaurants. The device uses the Google Speech-to-Text API to convert speech to text, and sends this text to a natural language processing engine (spaCy) to extract keywords. The generative AI (OpenAI GPT-4) generates a list based on the user's location information and displays it on the screen.

[1402] Example prompt sentence:

[1403] Example 1: Finding a restaurant

[1404] "User is looking for nearby restaurants. GPS information is ____. Please provide a list of the nearest restaurants."

[1405] Example 2: Restaurant details

[1406] "A user is looking for more information about Restaurant A. Please provide details such as Restaurant A's address, opening hours, and menu."

[1407] This process allows users to quickly and accurately obtain the information they need, and the system can provide personalized information based on location information and is also compatible with users with visual and hearing impairments.

[1408] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1409] Step 1: Getting User Input

[1410] The device receives natural language input from the user. The user can enter voice or text through the device's user interface. For voice input, the device uses speech recognition software (e.g., a speech recognition API) to convert the speech to text. The converted text is stored in memory and passed on to the next processing step.

[1411] Specific behavior:

[1412] The user speaks into the device's microphone, "What restaurants are nearby?"

[1413] The device uses a speech recognition API to convert the speech into text, generating the text "What restaurants are nearby?"

[1414] Input and Output:

[1415] Input: User voice or text input

[1416] Output: User request in text format

[1417] Step 2: Natural Language Processing

[1418] The device sends the text input received from the user to a natural language processing engine (e.g., NLP), which analyzes the text and extracts meaningful keywords and phrases. The extracted intents and keywords are passed on to the next step.

[1419] Specific behavior:

[1420] The device sends the text "What restaurants are nearby?" to a natural language processing engine.

[1421] A natural language processing engine analyzes the text and extracts keywords such as "nearby" and "restaurant."

[1422] Input and Output:

[1423] Input: User request in text format

[1424] Output: Extracted keywords and intent

[1425] Step 3: Information generation

[1426] The server sends the user's intent, analyzed by the natural language processing engine, to the generative AI (e.g., generative AI model), which generates relevant information. It also obtains the user's location information from the data management module and provides this location information to the generative AI. Based on this information, the generative AI generates information appropriate to the user's request in real time.

[1427] Specific behavior:

[1428] The server obtains the user's location information (e.g., latitude 35.6895, longitude 139.6917) from the data management module.

[1429] The intent and location information obtained from the natural language processing engine is sent to the generation AI.

[1430] The generation AI generates a list of restaurants such as "Restaurant A (500m)" and "Restaurant B (700m)."

[1431] Input and Output:

[1432] Input: Extracted keywords and intent, user location

[1433] Output: Generated related information (e.g., a list of restaurants)

[1434] Step 4: Dynamic UI Updates

[1435] The terminal passes the output data of the generated AI received from the server to the display module, which dynamically updates the user interface based on this data and provides the user with visual and auditory information.

[1436] Specific behavior:

[1437] The terminal passes the list of restaurants received from the server to the display module.

[1438] The display module updates the user interface to show information such as "Restaurant A (Distance: 500m)" and "Restaurant B (Distance: 700m)."

[1439] Users can view a list of nearby restaurants on their device screen.

[1440] Input and Output:

[1441] Input: Generated related information (e.g., a list of restaurants)

[1442] Output: Dynamically updated user interface and visual / auditory information

[1443] (Application example 1)

[1444] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1445] Conventional food delivery systems face challenges in efficiently obtaining the information users require and making it difficult to check order and delivery status in real time. They also lack an interface that is easy for visually and hearing impaired users to use. This reduces the quality of the user experience and limits the number of users.

[1446] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1447] In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, means for providing related information based on the user's location information, and means for receiving natural language input in voice or text format. This makes it possible to provide a system that allows users to efficiently obtain information using natural language and check order and delivery status in real time. It also makes it possible to provide an easy-to-use interface that is compatible with users with visual and hearing impairments.

[1448] A "means for receiving natural language input from a user" is any device or software for obtaining natural language input from a user in the form of speech or text.

[1449] "Natural language processing means for analyzing user requests" refers to technologies and algorithms that analyze the natural language entered by the user and understand their intentions and requests.

[1450] The "means for generating related information based on the analyzed request" refers to a method or technology for generating necessary information based on the results of analyzing the user's request.

[1451] "Means for dynamically updating the user interface to display generated information" refers to technologies and systems that change and update the displayed content in real time to present generated information to the user on screen or via audio.

[1452] "Means for providing relevant information based on a user's location information" refers to technologies or methods for utilizing a user's current location information to provide information related to that location.

[1453] A "means for receiving natural language input in voice or text form" is a device or software that recognizes and captures natural language input by a user through voice or text.

[1454] The "means for acquiring user location information and providing it to the generating means" refers to a technology or system for acquiring user location information and passing it to the means for generating related information.

[1455] "Means having a database for searching related information" refers to a database for searching related information based on the user's location information and a method for using the database.

[1456] "Means for tracking and displaying order and delivery status in real time" refers to technologies and systems that allow users to check the fulfillment status of their orders and the progress of deliveries in real time and display them to users.

[1457] "Means having an interface to assist in ordering procedures" refers to devices or software that have an interface that provides guidance and input assistance to enable users to place orders smoothly.

[1458] As an embodiment of this invention, we present a system that allows users to smoothly use food delivery services using natural language. The system is composed of components such as a user interface, a natural language processing engine, a generative AI, a data management module, and a display module.

[1459] Hardware and software used

[1460] Hardware:

[1461] Head-mounted displays and smartphones

[1462] microphone

[1463] software:

[1464] Python

[1465] geopy library: Getting location information

[1466] requests library: API calls

[1467] transformers library: natural language processing

[1468] speech_recognition library: speech recognition

[1469] Program processing explanation

[1470] User Interface

[1471] The user interface supports both voice and text input: users can say things like "I'd like to order a pizza," and the input is captured by a microphone and converted into text.

[1472] Natural Language Processing

[1473] User requests, entered via voice or text, are analyzed by a natural language processing engine. The transformers library is used to identify the user's intent. For example, a request such as "I want to order a pizza" can be interpreted as "Find nearby restaurants that serve pizza."

[1474] Information Generation and Display

[1475] The parsed request is sent to the generation AI, which generates relevant information based on the user's location information obtained from the data management module. For example, it generates a list of pizza restaurants nearest to the user's current location. The generated information is then passed to the display module, which dynamically updates the user interface and presents it to the user.

[1476] Order and delivery status tracking

[1477] The interface supports the user with menu information and ordering procedures for the restaurant they select, and also allows them to track and display the delivery status in real time after ordering, providing users with information such as the order progress and estimated arrival time.

[1478] Examples and prompts

[1479] For example, if a user says "I'd like to order pizza" to the head-mounted display, the system will analyze the request and display a list of nearby restaurants that serve pizza. The menu of the restaurant selected from the list will also be displayed, allowing the user to immediately proceed with the ordering process. After placing an order, the user can also check the current delivery status by asking "What's the status of my order?"

[1480] Example prompt sentence:

[1481] User: I want to order a pizza.

[1482] System: Finding pizza restaurants near you...

[1483] System: Choose from Restaurant A (distance: 500m), Restaurant B (distance: 700m), or Restaurant C (distance: 900m).

[1484] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1485] Step 1:

[1486] The user provides input via voice or text. The input is in the form of natural language, such as "I'd like to order a pizza." The device takes this input and, if it's voice, converts it to text using speech recognition software. The input data is passed on to the next processing step in natural language.

[1487] Step 2:

[1488] The device passes the captured input data to a natural language processing engine, which uses the transformers library to parse the input text and determine the user's intent, which is then converted into a clear instruction: "Find nearby pizza restaurants."

[1489] Step 3:

[1490] The server sends the parsed request to the generation AI, which generates a list of pizza restaurants based on the user's location information obtained from the data management module. The generation AI uses the user's location information as coordinates to extract data on pizza restaurants in the vicinity from a database. The output data is generated as a list of pizza restaurants and is passed to the next processing step.

[1491] Step 4:

[1492] The server sends the generated list of pizza restaurants to the display module. The display module dynamically updates the user interface and presents the list of pizza restaurants to the user. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed.

[1493] Step 5:

[1494] The user selects a restaurant from the list. The user's selection information is acquired by the terminal and sent to the server. The server then retrieves the detailed menu of the selected restaurant from the database and uses a generation AI to generate interface information to assist with the ordering process.

[1495] Step 6:

[1496] The server sends the generated interface information to the display module, which displays the menu of the selected restaurant to the user and allows the user to complete the ordering process. Once the user confirms the order, the information is sent back to the server.

[1497] Step 7:

[1498] Once an order is confirmed, the server stores the order information in a database and generates delivery information. The delivery status is updated in real time.

[1499] Step 8:

[1500] When the user asks, "What's the status of my order?", the device again sends this new input to the natural language processing engine, which analyzes it and determines that it's a request to check the delivery status.

[1501] Step 9:

[1502] The server retrieves delivery status information from the database and updates the current delivery status. The information is passed to the display module, which displays the delivery status to the user. For example, statuses such as "Currently cooking," "Delivering," and "Estimated arrival time: 20 minutes later" are displayed in real time.

[1503] In this way, a system is provided that allows users to efficiently use food delivery services using natural language input.

[1504] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1505] overview

[1506] The present invention provides a system that receives natural language input from a user, analyzes the user's request, and displays and updates the generated information in real time. It also combines an emotion engine that recognizes the user's emotions. The system also acquires the user's location information and provides relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[1507] System configuration

[1508] The system consists of the following main components:

[1509] 1. User Interface (UI)

[1510] An interface that allows users to enter questions or requests in natural language.

[1511] 2. Natural Language Processing (NLP) Engine

[1512] Algorithms for parsing user input and extracting its intent.

[1513] 3. Generation AI

[1514] Generates appropriate information based on the request analyzed by the NLP engine.

[1515] 4. Data Management Module

[1516] The user's location information and past usage history are stored and provided to the generating AI.

[1517] 5. Display module

[1518] Dynamically update the UI to display the generated information.

[1519] 6. Emotion Engine

[1520] An engine that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[1521] Program processing

[1522] Getting User Input

[1523] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[1524] Natural Language Processing

[1525] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[1526] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[1527] Emotion recognition

[1528] The device's emotion engine recognizes the user's emotions from natural language input, voice, and facial expressions. For example, it can identify the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[1529] information generation

[1530] The server sends the request content, emotion information, and location information analyzed by the NLP engine and emotion engine to the generation AI. The location information is obtained from the device's location information service.

[1531] The generative AI uses this information to generate appropriate restaurant information in real time. For example, if the user is feeling stressed, it will suggest quiet and relaxing restaurants.

[1532] Dynamic UI Updates

[1533] The device analyzes the information it receives from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants) and repeat the same process to provide real-time responses for further inquiries about details or other options.

[1534] Specific examples

[1535] Example 1: Finding a restaurant

[1536] 1. User types "What restaurants are near me?"

[1537] 2. The device sends the input to a natural language processing engine for analysis.

[1538] 3. The NLP engine identifies the intent: "Looking for nearby restaurants."

[1539] 4. The emotion engine identifies stress from the user's voice.

[1540] 5. Generative AI will generate a list of quiet and relaxing restaurants based on the user's location and emotional information.

[1541] 6. The device displays the listed information on the user interface, providing information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1542] Example 2: Restaurant details

[1543] 1. The user types, "Tell me the details about Restaurant A."

[1544] 2. The device sends the input to a natural language processing engine for analysis.

[1545] 3. The NLP engine identifies the intent as "I want to know more about Restaurant A."

[1546] 4. The emotion engine identifies that the user is relaxed in their voice.

[1547] 5. The generation AI generates detailed information about Restaurant A, such as its address, opening hours, and menu.

[1548] 6. The device displays the detailed information in the user interface.

[1549] In this way, the system of the present invention provides the information the user wants in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[1550] The processing flow will be explained below.

[1551] Step 1:

[1552] The user provides natural language input, such as "What restaurants are nearby?", through the application's user interface. The user can use voice input or text input.

[1553] Step 2:

[1554] The device receives input from the user and stores the text data, which is then sent to a natural language processing (NLP) engine for analysis.

[1555] Step 3:

[1556] The device's NLP engine analyzes the user's input and extracts their intent. For example, it may extract "nearby restaurants" as a keyword and determine that the user is searching for a restaurant.

[1557] Step 4:

[1558] The device's emotion engine recognizes the user's emotions from natural language input and voice. For example, it identifies the emotion "the user is feeling stressed" from the tone of voice and the content of the input.

[1559] Step 5:

[1560] The device sends the request content analyzed by the NLP engine, the emotional information recognized by the emotion engine, and the user's location information to the generation AI. The location information is obtained from the device's location information service.

[1561] Step 6:

[1562] The server's AI generates optimal restaurant information based on the request, emotional information, and location information received. If the user is feeling stressed, it will prioritize quiet and relaxing restaurants.

[1563] Step 7:

[1564] The AI ​​searches the server's restaurant database and lists the restaurants closest to the user's current location, along with detailed information such as the distance, ratings, and opening hours of each restaurant.

[1565] Step 8:

[1566] The AI ​​then sends the generated restaurant information to the device, including specific information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1567] Step 9:

[1568] The device analyzes the information received from the generative AI and dynamically updates the user interface, allowing users to instantly see the information they are looking for (e.g., a list of nearby restaurants).

[1569] Step 10:

[1570] The user can then ask further questions or make requests based on the displayed information, for example, by typing in an additional question such as "Tell me more about Restaurant A."

[1571] Step 11:

[1572] The terminal again sends the new user input to the natural language processing engine and repeats the process described above.

[1573] Through these steps, the system provides users with the information they want in real time, realizing an intuitive and efficient user experience, allowing users to obtain the most appropriate information according to their emotions and circumstances.

[1574] Example 2

[1575] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1576] Conventional natural language processing systems provide only simple information without considering the user's emotions or location, making it difficult to provide appropriate information that meets the user's needs.In addition, the information provided to users with visual and hearing impairments is insufficient.

[1577] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1578] In this invention, the server includes means for recognizing a user's emotion from the user's natural language input, voice, and facial expression, means for analyzing the emotion information, and means for acquiring the user's location information and providing it to the generation means. This makes it possible to provide appropriate information based on the user's emotion and location information, and realizes a system that can also accommodate users with visual and hearing impairments.

[1579] "Natural language input" is a method by which users enter questions or requests in everyday language.

[1580] "Natural language processing" is a technology that analyzes user input and extracts their intent.

[1581] The "generator" is a function that generates related information based on the analyzed request.

[1582] "User interface" refers to a screen or operating interface that is dynamically updated to display generated information.

[1583] An "emotion engine" is a system that recognizes a user's emotions from their natural language input, voice, and facial expressions.

[1584] "Emotional information" is data about the user's emotional state as recognized using the emotion engine.

[1585] "Location information" is data that indicates a user's current location.

[1586] "Dynamic update" refers to changing screens and data in real time as needed.

[1587] A "database" is a system for storing and searching information.

[1588] "Visually and hearing impaired users" refers to people who have limited vision or hearing, and includes special assistive devices to accommodate these users.

[1589] The present invention combines a system that receives natural language input from users, analyzes their requests, and displays and updates the generated information in real time with an emotion engine that recognizes the user's emotions. The system also has the ability to acquire the user's location information and provide relevant information based on that location. The system supports both voice and text input, and is suitable for users with visual and hearing impairments.

[1590] System configuration

[1591] The system consists of the following main components:

[1592] 1. User Interface (UI)

[1593] An interface that allows users to input questions or requests in natural language, supports both speech and text input, and is designed to be accessible to users with visual or hearing impairments.

[1594] 2. Natural Language Processing (NLP) Engine

[1595] It is an algorithm that analyzes user input and extracts their intent. For example, if a user types "What restaurants are nearby?", an NLP engine analyzes this input and identifies the intent as "looking for a nearby restaurant."

[1596] 3. Generation AI

[1597] It generates appropriate information based on requests analyzed by the NLP engine, and also references the emotion engine and location information to provide information that is best suited to the user's situation.

[1598] 4. Data Management Module

[1599] The user's location information and past usage history are stored and provided to the AI ​​generator. In particular, location information is obtained from the device's location information service.

[1600] 5. Display module

[1601] The UI for displaying the generated information is dynamically updated, allowing users to instantly see the information they are looking for.

[1602] 6. Emotion Engine

[1603] This engine recognizes the user's emotions from natural language input, voice, and facial expressions. The emotion engine identifies emotions such as stress or relaxation from the tone of voice and the content of the input.

[1604] Specific operation example

[1605] Example 1: Finding a restaurant

[1606] When a user types "What restaurants are nearby?", the device's voice recognition system converts the speech into text, which is then sent to a natural language processing engine. The NLP engine analyzes the input and identifies the user's intent: "I'm looking for a nearby restaurant." At the same time, the emotion engine identifies stress from the user's voice, and this information is sent to the server's generation AI. The generation AI then creates a list of "quiet, relaxing restaurants nearby" and sends the information to the device. The device dynamically updates the user interface, providing information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[1607] Example 2: Restaurant details

[1608] When a user types "Tell me more about Restaurant A," the device sends the input to a natural language processing engine for analysis. The NLP engine identifies the intent as "I want to know more about Restaurant A," and the emotion engine identifies a relaxed tone from the user's voice. The server then requests detailed information about Restaurant A, such as its address, opening hours, and menu, from the generation AI, and the generated details are sent to the device and displayed in the user interface.

[1609] Prompt Sentence Examples

[1610] "Can you recommend a quiet, relaxing restaurant nearby?"

[1611] "Please tell me more information about Restaurant A."

[1612] In this way, the system of the present invention provides the information the user desires in real time based on the user's natural language input and emotional information, realizing an intuitive and efficient user experience.

[1613] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1614] Step 1: Getting User Input

[1615] A user inputs a question or request through the application's user interface. For example, the user may input "What restaurants are nearby?" by voice or text. The input may be the user's voice or text data. In the case of voice, the device's speech recognition system converts the voice to text. The converted text is "What restaurants are nearby?"

[1616] Specific behavior:

[1617] The user presses the microphone button and says, "What restaurants are nearby?"

[1618] The device's voice recognition system converts the speech into text.

[1619] The text data "What are the nearby restaurants?" is saved.

[1620] Step 2: Natural Language Processing

[1621] The device sends user input to a natural language processing (NLP) engine. The input is text data such as "What restaurants are nearby?" The NLP engine analyzes the user's input and extracts its intent. The output of the NLP engine is the analysis result, which identifies the user's intent as "looking for a nearby restaurant."

[1622] Specific behavior:

[1623] The text data "What restaurants are nearby?" is sent to the NLP engine.

[1624] An NLP engine analyzes text data and identifies intent.

[1625] The analysis result, "Looking for a nearby restaurant," is returned to the device.

[1626] Step 3: Recognize emotions

[1627] The device's emotion engine recognizes the user's emotions from their natural language input, voice, and facial expressions. The input is the user's voice and natural language input. The emotion engine identifies emotions from the tone of the voice and the content of the input. The output is emotional information that the user is "feeling stressed."

[1628] Specific behavior:

[1629] The user's voice is sent to the emotion engine.

[1630] The emotion engine analyzes the tone and pace of the voice.

[1631] Emotional information "user is feeling stressed" is stored on the device.

[1632] Step 4: Information Generation

[1633] The server sends the analyzed request content (user intent), emotional information, and location information to the generation AI. The inputs are the analysis result "looking for a nearby restaurant," emotional information "stress," and location information. The generation AI generates appropriate information based on this information. The output is the information "quiet, relaxing restaurants nearby."

[1634] Specific behavior:

[1635] The analysis result, "Looking for a nearby restaurant," emotional information, "stress," and location information are sent to the generating AI.

[1636] The generative AI creates a prompt and sends it to the AI ​​model.

[1637] The AI ​​model generates appropriate restaurant information and returns a list such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[1638] Step 5: Dynamic UI Updates

[1639] The device analyzes the information received from the generation AI and dynamically updates the user interface. The input is the generated restaurant list. The output is specific information displayed on the user interface (e.g., "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)").

[1640] Specific behavior:

[1641] The generated restaurant list is sent to the terminal.

[1642] The device's UI module parses the list and updates the user interface.

[1643] As a result, the interface will display information such as "Restaurant A (distance: 500m, quiet)" and "Restaurant B (distance: 700m)."

[1644] (Application example 2)

[1645] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1646] In recent years, an increasing number of systems analyze natural language input from users and provide relevant information to improve user experience. However, conventional systems often do not fully utilize emotion recognition or location information, making it difficult to provide services that respond to individual users' situations and emotions. Therefore, there is a need to develop systems that take user emotions and location information into account and provide more personalized information in real time.

[1647] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving natural language input from a user, natural language processing means for analyzing the user's request, generation means for generating related information based on the analyzed request, means for dynamically updating a user interface to display the generated information, emotion recognition means for recognizing emotions from the user's voice input and facial expressions, and means for customizing suggestions based on the emotions recognized by the emotion recognition means. This makes it possible to provide more personalized information based on the user's real-time emotions and location information.

[1648] - "Natural language input" refers to users asking questions or giving instructions in the language they normally use.

[1649] "Natural language processing means" refers to technology that analyzes the natural language entered by the user and extracts its meaning and intent.

[1650] "Generation means" refers to a technology that creates relevant information based on analyzed user requirements.

[1651] "User interface" refers to the display device and operating environment through which a user receives information.

[1652] "Dynamic updating" refers to instantly changing the displayed content in response to new user input or circumstances.

[1653] "Emotion recognition means" refers to technology that identifies emotions from the user's voice, facial expressions, etc.

[1654] "Customizing recommendations" refers to providing more personalized information and recommendations based on perceived emotions.

[1655] "Location Information" refers to geographic data about a user's current location.

[1656] A "database" refers to a system that systematically stores and searches information.

[1657] "Voice input" refers to a method in which a user gives instructions to a system through a voice input device such as a microphone.

[1658] "Text input" refers to the method of entering characters using a keyboard or touch panel.

[1659] "Accessible to the visually impaired" refers to a design that allows users with visual impairments to use the system.

[1660] "Hearing-impaired" refers to a design that allows users with hearing impairments to use the system.

[1661] System Overview

[1662] The system of this invention receives natural language input from users, analyzes their requests, and provides relevant information in real time. The system combines an emotion recognition unit that recognizes the user's emotions with a unit that acquires the user's location information to provide more personalized information. Furthermore, the system supports both voice and text input, and is designed to accommodate users with visual and hearing impairments.

[1663] Hardware and Software Configuration

[1664] Hardware:

[1665] Smart glasses: To display product information

[1666] Camera: To recognize product IDs (e.g., QR codes)

[1667] Microphone: To recognize your voice input

[1668] software:

[1669] OpenCV: To capture the camera stream and recognize product IDs

[1670] dlib: Face recognition and eye tracking

[1671] Transformers (Hugging Face): Natural Language Processing and Emotion Recognition

[1672] Geopy: Getting location information

[1673] SpeechRecognition: Recognizing voice input

[1674] Data processing and calculation

[1675] The server performs the following data processing and calculations.

[1676] 1. Natural Language Processing: Receiving and analyzing voice and text input from the user. This is done using the Transformers library.

[1677] 2. Product ID recognition: Analyze the input from the camera using OpenCV and recognize the product ID.

[1678] 3. Emotion Recognition: Based on the user's voice or input data, emotion recognition is used to identify emotions. This process is also performed using the Transformers library.

[1679] 4. Information generation: Generative AI generates optimal information based on analyzed requests, emotional information, and location information.

[1680] 5. Dynamic UI Updates: Update the user interface in real time to display the appropriate information.

[1681] Specific Examples

[1682] Example 1: Finding a restaurant

[1683] User: "What restaurants are nearby?" speaks through the smart glasses.

[1684] Device: The input voice data is converted to text using SpeechRecognition. The text data is then analyzed using Transformers' NLP engine. The analysis results identify that the user is searching for a restaurant.

[1685] Emotion recognition: Recognizes stress from the user's tone of voice and language.

[1686] Generative AI: Lists quiet and relaxing restaurants based on the user's location (obtained through Geopy) and emotional information.

[1687] User interface: The smart glasses display shows information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)."

[1688] Example 2: Retrieving product information

[1689] User: Browses a shelf of red wines in a brick-and-mortar store and asks aloud through smart glasses, "Tell me about this wine."

[1690] Camera: The camera in the smart glasses recognizes the product ID (e.g., QR code) using OpenCV.

[1691] Generative AI: Generates information about red wine based on the recognized product ID.

[1692] Emotion recognition: Recognizes when the user is relaxed and determines that no additional suggestions are necessary.

[1693] User Interface: Smart glasses display the information "Red wine, price: \2500, delicious red wine."

[1694] Prompt Sentence Examples

[1695] Product ID recognition

[1696] "Tell me about this product"

[1697] emotion recognition

[1698] "tired..."

[1699] As described above, the system of the present invention can provide personalized information to users in real time based on the user's natural language input, emotions, and location information.

[1700] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1701] Step 1:

[1702] The user inputs natural language through the smart glasses. In this case, the user asks questions such as "What are some restaurants nearby?" The input voice data is converted into text data by the voice input means. Specifically, voice data is acquired from the microphone and converted into text using the SpeechRecognition library. Input data: voice data, output data: text data.

[1703] Step 2:

[1704] The device uses a natural language processing engine (NLP engine) to extract intent from the user's natural language input. Specifically, it uses the Transformers library to analyze the input text data and extract intents such as "looking for a restaurant." Input data: text data, Output data: extracted intent.

[1705] Step 3:

[1706] The device sends data to a generation means for generating related information based on the analyzed intent. The generation means then sends the related information to a generative AI model (Transformers) based on the extracted intent and location information. Input data: extracted intent, location information; output data: related information.

[1707] Step 4:

[1708] The device uses the camera to recognize the IDs of surrounding products. It analyzes the camera image using the OpenCV library and obtains the product ID by reading the QR code or barcode. Input data: camera image, Output data: product ID.

[1709] Step 5:

[1710] The device uses emotion recognition to customize the information generated by the generative AI model according to the user's emotions. Specifically, it uses the Transformers library to identify emotions from the user's voice tone and vocabulary, and adjusts the recommendations accordingly. Input data: voice data and text data. Output data: recognized emotions.

[1711] Step 6:

[1712] The server customizes the generated related information based on the emotional information obtained by the emotion recognition means. The generative AI model provides optimal information based on the user's emotions, such as suggesting quiet and relaxing restaurants. Input data: emotional information, related information. Output data: customized suggestions.

[1713] Step 7:

[1714] The device dynamically updates the user interface and displays information created by the generative AI model on the smart glasses display. For example, information such as "Restaurant A (distance: 500m)" and "Restaurant B (distance: 700m)" is displayed in real time and provided to the user. Input data: customized suggestions, output data: dynamically updated user interface.

[1715] This series of processing steps allows users to obtain personalized information in real time, enabling them to use services more efficiently.

[1716] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1717] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1718] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1719] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1720] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1721] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1722] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1723] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1724] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1725] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1726] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1727] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1728] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1729] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1730] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1731] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1732] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1733] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1734] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1735] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1736] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1737] The following is further disclosed regarding the above embodiment.

[1738] (Claim 1)

[1739] means for receiving natural language input from a user;

[1740] natural language processing means for analyzing user requests;

[1741] generating means for generating related information based on the analyzed request;

[1742] means for dynamically updating a user interface to display the generated information;

[1743] A system including:

[1744] (Claim 2)

[1745] A means for acquiring user location information and providing it to a generating means;

[1746] a means for searching a database for relevant information based on the user's location information;

[1747] The system of claim 1 further comprising:

[1748] (Claim 3)

[1749] 10. The system of claim 1, wherein the user interface that displays the generated information has means for supporting voice input and text input, and further comprises means for accommodating visually and hearing impaired users.

[1750] "Example 1"

[1751] (Claim 1)

[1752] terminal means for receiving natural language input from a user;

[1753] natural language processing means for analyzing user requests;

[1754] A generating AI means for generating relevant information based on the analyzed request;

[1755] display means for dynamically updating a user interface to display the generated information;

[1756] A system including:

[1757] (Claim 2)

[1758] A data management means for acquiring user location information and providing it to the generating AI means;

[1759] a data management means having a database for searching related information based on user location information;

[1760] The system of claim 1 further comprising:

[1761] (Claim 3)

[1762] 2. The system of claim 1, wherein the user interface that displays the generated information has a terminal means that supports voice input and text input, and further includes a display means that is accommodating users with visual and hearing impairments.

[1763] "Application Example 1"

[1764] (Claim 1)

[1765] means for receiving natural language input from a user;

[1766] natural language processing means for analyzing user requests;

[1767] generating means for generating related information based on the analyzed request;

[1768] means for dynamically updating a user interface to display the generated information;

[1769] means for providing relevant information based on a user's location;

[1770] means for receiving natural language input in speech or text form;

[1771] A system including:

[1772] (Claim 2)

[1773] A means for acquiring user location information and providing it to a generating means;

[1774] a means having a database for searching for relevant information based on the user's location information;

[1775] A way to track and view order and delivery status in real time, and

[1776] 10. The system of claim 1, further comprising:

[1777] (Claim 3)

[1778] a user interface for displaying the generated information, the user interface having means for supporting voice and text input;

[1779] Provision for accessibility to visually and hearing impaired users, and

[1780] means for providing an interface for assisting in the ordering process;

[1781] 10. The system of claim 1, further comprising:

[1782] "Example 2: Combining Emotion Engines"

[1783] (Claim 1)

[1784] means for receiving natural language input from a user;

[1785] natural language processing means for analyzing user requests;

[1786] generating means for generating related information based on the analyzed request;

[1787] means for dynamically updating a user interface to display the generated information;

[1788] A means for recognizing a user's emotions from the user's natural language input, voice, and facial expressions;

[1789] A means for analyzing emotional information;

[1790] A system including:

[1791] (Claim 2)

[1792] A means for acquiring user location information and providing it to a generating means;

[1793] a means for searching a database for relevant information based on the user's location information;

[1794] 2. The system according to claim 1, wherein the generating means generates the information using the user's emotional information and location information.

[1795] (Claim 3)

[1796] 10. The system of claim 1, wherein the user interface that displays the generated information has means for supporting voice input and text input, and further comprises means for accommodating visually and hearing impaired users.

[1797] "Application example 2 when combining emotion engines"

[1798] (Claim 1)

[1799] means for receiving natural language input from a user;

[1800] natural language processing means for analyzing user requests;

[1801] generating means for generating related information based on the analyzed request;

[1802] means for dynamically updating a user interface to display the generated information;

[1803] an emotion recognition means for recognizing emotions from a user's voice input and facial expressions;

[1804] means for customizing suggestions based on emotions recognized by the emotion recognition means;

[1805] A system including:

[1806] (Claim 2)

[1807] A means for acquiring user location information and providing it to a generating means;

[1808] a means for searching a database for relevant information based on the user's location information;

[1809] A means for providing optimal relevant information to a user based on the user's location information and emotion information;

[1810] The system of claim 1 further comprising:

[1811] (Claim 3)

[1812] a user interface for displaying the generated information, the user interface having means for supporting voice and text input;

[1813] 10. The system of claim 1, further comprising means for accommodating visually and hearing impaired users. [Explanation of symbols]

[1814] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving natural language input from a user; natural language processing means for analyzing user requests; generating means for generating related information based on the analyzed request; means for dynamically updating a user interface to display the generated information; A system including:

2. A means for acquiring user location information and providing it to a generating means; a means for searching a database for relevant information based on the user's location information; The system of claim 1 further comprising:

3. 10. The system of claim 1, wherein the user interface for displaying the generated information has means for supporting voice input and text input, and further comprises means for accommodating visually and hearing impaired users.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A