System

The system addresses the limitations of conventional humanoid robots by integrating image recognition, natural language processing, and multilingual translation to facilitate natural interaction and efficient service provision, including autonomous navigation.

JP2026028049APending Publication Date: 2026-02-19SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024130347
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Conventional humanoid robots face challenges in accurately recognizing user emotions and facial expressions, have limited multilingual capabilities, and require manual processes, hindering natural conversations and efficient service provision.

Method used

A system equipped with image recognition, natural language processing, multilingual translation, and database connection means to analyze user faces and gestures, understand user intent, and provide personalized responses while navigating autonomously.

Benefits of technology

Enables natural interaction and efficient service provision by accurately recognizing users, supporting multiple languages, and autonomously guiding users to destinations, reducing workload and improving service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028049000001_ABST
    Figure 2026028049000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes an image recognition means, a natural language processing means, a multilingual translation means, a database connection means, and a means for recognizing the face and expression of a user by the image recognition means, analyzing the question of the user by the natural language processing means, generating information corresponding to the language of the user by the multilingual translation means, and acquiring a user profile and a conversation history by the database connection means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional humanoid robots have difficulty accurately recognizing users' emotions and facial expressions, making it difficult to have natural conversations or respond to advanced questions. Furthermore, development of robots with multilingual capabilities, autonomous navigation, and learning capabilities has been insufficient. Furthermore, many processes must be performed manually, limiting improvements to operational efficiency. The present invention aims to solve these issues and provide an advanced robot with emotion recognition, natural language support, multilingual capabilities, autonomous mobility, information provision, and learning capabilities. [Means for solving the problem]

[0005] The present invention is a system equipped with image recognition means, natural language processing means, multilingual translation means, and database connection means. First, the image recognition means recognizes the user's face and facial expressions, and then the natural language processing means analyzes the user's question. Based on this information, the multilingual translation means generates information corresponding to the user's language. Finally, the database connection means is used to acquire the user profile and dialogue history. This enables accurate recognition of the user's emotions and gestures, enabling natural and sophisticated response to questions. Furthermore, the various means within the system work together to efficiently provide information and perform business processing.

[0006] "Image recognition means" refers to technology that uses cameras and sensors to capture a user's face and facial expressions, and then analyzes and identifies them.

[0007] "Natural language processing means" refers to technology that analyzes natural language text entered by a user, understands its meaning and intent, and generates a response.

[0008] "Multilingual translation tool" refers to technology that automatically translates text from one language into multiple other languages.

[0009] "Database connection means" refers to the technology that enables a system to connect to a database and retrieve and store information such as user profiles and interaction history.

[0010] "User" refers to the person who interacts with the system and inputs questions and commands.

[0011] A "user profile" is a data set that aggregates information about a user, including their individual characteristics and interaction history.

[0012] "Interaction history" refers to a record of past interactions between a user and a system.

[0013] "Emotion" refers to the psychological state inferred from the user's facial expressions and voice.

[0014] A "gesture" refers to an action in which a user expresses their intentions or emotions through physical movements.

[0015] "Autonomous movement" refers to the ability of a robot to move to a destination on its own without external instructions.

[0016] "Providing information" refers to the act of providing appropriate data or advice in response to a user's questions or requests. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention relates to a multimodal AI system using a humanoid robot. This system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. Specific embodiments of the system are described below.

[0039] System initialization

[0040] The server performs the initial setup of the system, including loading the image recognition model, natural language processing (NLP) model, and multilingual translation model, and setting up the database, so that the entire system is ready to run efficiently and smoothly.

[0041] User Awareness

[0042] The robot uses a camera to recognize the user. During this process, the camera captures the user's face and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[0043] Start a conversation

[0044] The robot uses the user information obtained by the recognition means to initiate a conversation using natural language processing means. First, it retrieves the user profile from the database and generates an appropriate greeting based on that information. This allows the user to begin a natural dialogue with the system.

[0045] Answering questions

[0046] When a user types a question into the robot, it analyzes it using natural language processing. During the analysis, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. For example, if a user asks about the weather, the system retrieves weather information and provides it to the user.

[0047] Autonomous Mobility and Navigation

[0048] The robot can move autonomously to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path. During this process, it is designed to reach the destination while avoiding obstacles and other people. For example, if the user says, "Please guide me to the reception desk," the robot will calculate the route to the reception desk and actually begin guiding the user.

[0049] Specific examples

[0050] Example 1: Store directions

[0051] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expression, and analyzes it using image recognition. The robot then analyzes the intent of the question using natural language processing, and the system generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and actually guides the user to the reception desk.

[0052] Example 2: Bank balance inquiry

[0053] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot responds, "My current account balance is 100,000 yen." This allows the user to obtain information quickly and accurately.

[0054] Through these functions, the system of the present invention aims to achieve natural interaction with users and highly efficient service provision, thereby reducing workload and improving service quality.

[0055] The processing flow will be explained below.

[0056] Step 1:

[0057] The server initializes the entire system. First, it loads the image recognition, natural language processing, and multilingual translation models. Next, it sets up the database used by the system and prepares it to store user profiles and interaction history.

[0058] Step 2:

[0059] The robot uses a camera to capture images to recognize the user's face and facial expressions, and the captured images are passed to an image recognition means to analyze the user's face, emotions, and gestures.

[0060] Step 3:

[0061] Based on the analysis results, the robot retrieves the user's profile from the database, allowing it to refer to each user's past interaction history and specific information.

[0062] Step 4:

[0063] The robot uses natural language processing tools to initiate a conversation with the user, generating appropriate greetings and introductory phrases based on the user's profile, emotions, and gestures, and then delivering them to the user via voice or text.

[0064] Step 5:

[0065] When a user types a question into the robot, the question is analyzed using natural language processing tools, which extracts the intent of the question and related entities, thereby providing a clear understanding of the question.

[0066] Step 6:

[0067] Based on the analysis results, the robot retrieves the appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. The response is then provided to the user.

[0068] Step 7:

[0069] When a user requests guidance to a specific location, the robot begins navigation to the destination. First, it calculates the optimal path to the destination and then autonomously moves along that path, avoiding obstacles and other people along the way to safely reach the destination.

[0070] Step 8:

[0071] The robot records the results of its interactions with the user and their actions in a database, which can be used for future interactions. This record includes the interaction history, the user's reactions, and any new information acquired. This allows the system to continuously learn and improve the quality of its services.

[0072] In this way, the system of the present invention realizes natural interaction with the user and provides a variety of services.

[0073] Example 1

[0074] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0075] Conventional humanoid robot systems have had problems with natural interaction with users and efficient service provision. Specifically, they suffer from low recognition accuracy, delayed responses, and limited ability to support different languages. Furthermore, there is still room for improvement in terms of recognizing users' faces and gestures, identifying emotions, providing appropriate information, and navigation.

[0076] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0077] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and initial setting means. This allows the server to recognize the user's face and facial expressions, identify their ID and emotions, and provide appropriate information in multiple languages, thereby realizing natural interaction with the user and enabling efficient service provision.

[0078] "Image recognition means" refers to means that has the function of analyzing image data acquired using an imaging device such as a camera and recognizing a specific object (for example, a user's face or gesture).

[0079] "Natural language processing means" refers to means that include algorithms and models for analyzing text data entered by a user and understanding their intent and emotions.

[0080] A "multilingual translation means" is a means that has the function of translating text between different languages, and that understands input from a user in different languages ​​and generates corresponding information.

[0081] "Database connection means" refers to a means that has the function of communicating with a database server and acquiring and saving the necessary information.

[0082] The "means for performing initial settings" refers to a means for performing basic settings for system operation at the start, including loading various models and setting up a database.

[0083] "Means for recognizing a user's face and facial expressions" refers to means that has the function of extracting specific features from a user's facial image captured by a camera, identifying the user based on the results, and determining their emotional state.

[0084] "Means for identifying a user's ID and emotions" refers to a means for analyzing data obtained through image recognition to determine who the user is (ID) and their emotional state at the time.

[0085] "Means for obtaining a user profile" means means capable of obtaining information related to a user from a database, including past conversation history and preferences.

[0086] The "means for generating a greeting message" is a means having a function for automatically generating a greeting message suitable for a user based on the acquired user profile.

[0087] "Means for analyzing the intent of a question and related entities" refers to a means for analyzing the question entered by the user and identifying the intent of the question and related keywords (entities).

[0088] The "means for generating an appropriate answer to a question" is a means for obtaining relevant information based on the intent and entities of the analyzed question, and automatically generating an appropriate answer based on that information.

[0089] The "means for continuously capturing image data" refers to a means having a function for activating a camera and continuously acquiring image data at regular intervals.

[0090] The "means for preprocessing image data" refers to a means that has the function of performing preprocessing such as resizing, noise removal, and color conversion on the acquired image data.

[0091] "Means for processing sensor information in real time and moving autonomously" refers to means that analyze information from various sensors in real time, calculate a route according to the environment, and move autonomously.

[0092] The present invention relates to a multimodal AI system using a humanoid robot. The system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. It also includes means for recognizing a user's face and facial expressions, identifying the user's ID and emotion, acquiring a user profile, generating a greeting, analyzing the intent of a question and related entities, generating appropriate answers to the question, continuously capturing image data, preprocessing the image data, and processing sensor information in real time to move autonomously.

[0093] 1. System initialization

[0094] The server performs the initial system configuration. First, for image recognition, it loads a ResNet50 model using TensorFlow using an NVIDIA GPU. Next, for natural language processing, it loads the BERT model using Hugging Face's Transformers library. It also configures the Google Translate API for multilingual translation, and starts a MySQL database to create the necessary tables and indexes.

[0095] 2. User Awareness

[0096] The robot captures the user's face and gestures using a built-in camera (e.g., a standard HD camera). The image data is preprocessed with OpenCV and then analyzed using a ResNet50 model running on NVIDIA GPUs, which identifies the user's identity, face, and facial expressions.

[0097] 3. Start a conversation

[0098] The robot starts a conversation using the data mentioned above. First, it uses the user's ID to retrieve the user profile from a MySQL database. Then, it uses that information to generate an appropriate greeting using Hugging Face's BERT model. For example, it might generate a greeting like, "Hello, Tanaka-san. How's your day?"

[0099] 4. Answering questions

[0100] Users input questions into the robot, which then uses voice input (e.g., a standard microphone) to receive the question and convert it into text using the Google Speech-to-Text API. It then uses Hugging Face's BERT model to analyze the intent of the question and relevant entities, and based on the analysis results, queries the appropriate API (e.g., the OpenWeatherMap API) or database to generate an answer to the question.

[0101] 5. Autonomous Movement and Navigation

[0102] The robot autonomously moves to a location specified by the user. First, the user specifies the destination via voice input or touchscreen. Next, the robot uses SLAM technology to generate a map of the environment and calculates the optimal route using algorithms such as the A algorithm. The robot autonomously moves along the calculated route using a LiDAR sensor and motor control unit, avoiding obstacles and reaching the destination.

[0103] Specific examples

[0104] In-store directions

[0105] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and first identifies the user using image recognition. It then analyzes the intent of the question using natural language processing and generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and autonomously guides the user to the reception desk using SLAM technology and a motor control unit.

[0106] Bank balance inquiry

[0107] When a user asks, "What is my account balance?", the robot first authenticates the user and retrieves account information from the database. After successful authentication, the robot responds, "Your current account balance is 100,000 yen."

[0108] Prompt Sentence Examples

[0109] "When a user asks for weather information, how do you get the information from the API and respond?"

[0110] "Please explain in detail the algorithm that recognizes user emotions."

[0111] As explained above, the system of the present invention realizes natural interaction with users and aims to reduce the workload and improve service quality through highly efficient service provision.

[0112] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0113] Step 1: Initialize the system

[0114] The server performs the initial system setup, inputting the image recognition model, natural language processing model, multilingual translation model, and database configuration information.

[0115] 1. As a deep learning model for image recognition, we load a ResNet50 model in TensorFlow using an NVIDIA GPU. The input is the TensorFlow model file, and the output is the model loaded in memory.

[0116] 2. Load the BERT model from Hugging Face's Transformers library as the natural language processing model. The input is the model file, and the output is the BERT model loaded in memory.

[0117] 3. Set up the Google Translate API as a multilingual translation model. The input is the API key and configuration information, and the output is the API availability status.

[0118] 4. Starts a MySQL database and creates the necessary tables and indexes. The input is database configuration information, and the output is a configured database ready to use.

[0119] Step 2: User Awareness

[0120] The robot uses a camera to recognize the user, and the input is image data obtained from the camera.

[0121] 1. Start a camera (e.g., HD camera) and continuously capture image data. The input is the camera sensor data, and the output is the captured image data.

[0122] 2. The acquired image data is preprocessed using the OpenCV library, which includes resizing, noise removal, color conversion, etc. The input is raw image data, and the output is preprocessed image data.

[0123] 3. The preprocessed image data is input to a ResNet50 model running on an NVIDIA GPU to identify the user's ID and emotion. The input is the preprocessed image data, and the output is the identified user's ID and emotion information.

[0124] Step 3: Start a conversation

[0125] The robot starts a conversation based on the recognized user information, including the user's ID and emotional information.

[0126] 1. The robot uses the user's ID to retrieve the user profile from the MySQL database. The input is the user's ID and the output is the retrieved user profile.

[0127] 2. Based on the retrieved user profile, the robot uses Hugging Face's BERT model to generate an appropriate greeting, such as "Hello, Tanaka-san. How are you today?" The input is the user profile, and the output is the generated greeting.

[0128] Step 4: Answer questions

[0129] The user inputs a question to the robot, and the input is voice data.

[0130] 1. Collect user questions using a voice input function (e.g., microphone). The input is voice data, and the output is the question data captured as voice.

[0131] 2. The robot converts voice data into text data using the Google Speech-to-Text API. The input is voice data and the output is text data.

[0132] 3. The text data is analyzed using Hugging Face's BERT model to extract the question intent and related entities. The input is the text data, and the output is the analysis results.

[0133] 4. Based on the analysis results, the robot queries the appropriate API (e.g., OpenWeatherMap API) or database to generate an answer to the question. The input is the analysis results and the necessary API information, and the output is the generated answer.

[0134] Step 5: Autonomous Movement and Navigation

[0135] The robot moves autonomously to a specified location, and the input is a destination specified by the user.

[0136] 1. The user specifies a destination using voice input or a touch screen. The input is the destination information, and the output is the specified destination.

[0137] 2. The robot uses SLAM technology to generate a map of the environment and calculate the optimal path. The input is sensor information and destination information, and the output is the calculated path information.

[0138] 3. The robot uses a LiDAR sensor and a motor control unit to move autonomously based on a calculated path. The input is path information and real-time sensor information, and the output is autonomous movement behavior.

[0139] (Application example 1)

[0140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0141] Modern stores and public facilities require a means for visitors to quickly and accurately find their destination and obtain information. However, conventional guidance systems and information centers face problems such as labor shortages and language barriers. Furthermore, visitors often have to search for the information they need themselves, which is inconvenient. A system that solves these issues and allows visitors to easily reach their destination is needed.

[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0143] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, and database connection means. This enables the server to recognize a user's face and facial expressions, analyze the user's questions, generate information corresponding to the user's language, and acquire a user profile and dialogue history. The server also includes connection means with a smart device, means for activating and controlling the robot via the smart device, autonomous robot movement means, and means for guiding the robot to a user-specified destination using the autonomous movement means. This allows visitors to easily call the robot from their smartphones, and the robot automatically guides the visitor to their destination, significantly improving visitor convenience.

[0144] "Image recognition means" refers to a device or system that uses a camera or sensor to recognize a user's face, facial expressions, and gestures.

[0145] "Natural language processing means" is a technology that analyzes questions and instructions uttered by users and understands their intent and content.

[0146] "Multilingual translation means" refers to technology that translates between different languages ​​and generates information corresponding to the language used by the user.

[0147] "Database connection means" refers to technology for accessing a database that stores user profiles, interaction history, and other necessary data, and acquiring that information.

[0148] "Means for connecting with smart devices" refers to technology that connects devices such as smartphones and tablets to the system, enabling two-way communication.

[0149] "Means for activating and controlling robots" refers to technology for activating robots via smart devices and controlling their movements and functions.

[0150] "Autonomous robot mobility means" refers to technology that enables a robot to move autonomously while recognizing its surrounding environment.

[0151] "Generating information" is the process of generating appropriate answers or guidance information based on the user's questions or instructions.

[0152] A "user profile" is data that compiles information about a user (e.g., the user's ID, past interaction history, preferences, etc.).

[0153] "Dialogue history" is data that records past conversations and interactions with a user.

[0154] "Autonomous movement" refers to a robot's ability to move while adapting to its surrounding environment based on its own judgment.

[0155] The present invention relates to a system that controls a robot in cooperation with a smart device to provide guidance and information to users in stores and public facilities. This system is realized using the following various means.

[0156] 1. Image Recognition Methods:

[0157] The server uses cameras and sensors to recognize the user's face, facial expressions, and gestures, using software libraries such as OpenCV and TensorFlow, which are known for their image processing technology.

[0158] 2. Natural Language Processing Tools:

[0159] The server uses natural language processing techniques to analyze the user's questions and instructions. This process uses various NLP (Natural Language Processing) libraries and tools (e.g., Google's Dialogflow, OpenAI's GPT-3).

[0160] 3. Multilingual translation tools:

[0161] The server uses a multilingual translation engine to translate between different languages ​​and generate information corresponding to the user's language, for example, the Google Translate API.

[0162] 4. Database connection method:

[0163] The server retrieves information such as the user's profile and interaction history from a database, using Firebase or AWS RDS to dynamically manage user information.

[0164] 5. Connectivity with smart devices:

[0165] The smartphone application communicates with the robot via Bluetooth or Wi-Fi to activate and control it, allowing users to summon the robot using their smartphone.

[0166] 6. Autonomous robotic mobility:

[0167] The robot autonomously navigates to its destination using SLAM (Simultaneous Localization and Mapping) technology with Lidar sensors and cameras, using ROS Lidar and Google Cartographer.

[0168] 7. Information generation:

[0169] The server generates appropriate answers and guidance information in response to user questions and instructions, using a generative AI model in the process.

[0170] 8. User Awareness and Guidance:

[0171] The robot recognizes the user with a camera, analyzes the intent of the question using natural language processing technology, and provides guidance. If the user asks, "Where is the dairy section?", the robot will use the store's map information to answer, "The dairy section is on the right side of this floor. I'll show you." and begin guiding the user. Using its autonomous mobility, the robot will accurately guide the user to their destination.

[0172] 9. Example prompt:

[0173] An example of a prompt sentence to input to the generative AI model is as follows:

[0174] "A user is in a store asking about the dairy section. The robot recognizes the user with a camera and analyzes the intent of the question using NLP technology. It then retrieves store map information from a database and provides directions. Please explain the specific processing flow in detail."

[0175] In this way, the system of the present invention can provide an environment in which the user can quickly and accurately obtain information regardless of the environment.

[0176] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0177] Step 1:

[0178] The user starts the smartphone app and presses the "Start Guide Robot" button, at which point the smartphone uses its camera and microphone to check the surrounding environment.

[0179] input:

[0180] User presses a button

[0181] output:

[0182] Camera and microphone ambient data

[0183] Specific behavior:

[0184] The camera captures image data of the surroundings, and the microphone captures audio data.

[0185] Step 2:

[0186] The device calls the robot via Bluetooth or Wi-Fi and sends commands for the robot to move to the user's location.

[0187] input:

[0188] Camera and microphone data

[0189] output:

[0190] Robot commands

[0191] Specific behavior:

[0192] It connects to the robot communication module via Bluetooth or Wi-Fi and sends the command "move to the user."

[0193] Step 3:

[0194] The robot recognizes the user using a camera and analyzes the user's face and gestures using image recognition means.

[0195] input:

[0196] Image data acquired by the camera

[0197] output:

[0198] User ID, facial expression, and gesture information

[0199] Specific behavior:

[0200] Image processing libraries such as OpenCV and TensorFlow are used to recognize and analyze the user's face and gestures and obtain user information.

[0201] Step 4:

[0202] The server uses natural language processing means to analyze the user's question and understand its intent.

[0203] input:

[0204] User utterances

[0205] output:

[0206] User questions and their intent

[0207] Specific behavior:

[0208] NLP tools such as Google's Dialogflow and OpenAI's GPT-3 are used to interpret the content of the speech and extract the intent of the question.

[0209] Step 5:

[0210] The server uses database connectivity to obtain the user's profile and interaction history and generates appropriate information.

[0211] input:

[0212] User ID, question content

[0213] output:

[0214] User profile and corresponding answer information

[0215] Specific behavior:

[0216] It connects to Firebase or AWS RDS, obtains the user's past interaction history and profile information, and then uses a generative AI model to generate appropriate answers and guidance information.

[0217] Step 6:

[0218] The server uses a multilingual translation means to translate the generated information into the language used by the user.

[0219] input:

[0220] Generated answer information

[0221] output:

[0222] Information translated into the user's language

[0223] Specific behavior:

[0224] Use the Google Translate API to translate the generated information into the user's language.

[0225] Step 7:

[0226] The robot presents appropriate information to the user via voice or text and begins providing guidance in response to the user's questions.

[0227] input:

[0228] Translated information

[0229] output:

[0230] User guidance information

[0231] Specific behavior:

[0232] The text information is converted into speech using a speech synthesis engine (e.g., Google TTS) and provided to the user. The autonomous mobility system is also used to calculate the route to the user's specified destination and begin providing guidance.

[0233] Step 8:

[0234] The robot uses autonomous mobility to guide the user to their destination while avoiding obstacles.

[0235] input:

[0236] User-specified destination information and surrounding environment information

[0237] output:

[0238] Guidance to the destination, reaching the final point

[0239] Specific behavior:

[0240] It uses Lidar sensors and SLAM technology (e.g., ROS Lidar, Google Cartographer) to map the surrounding environment, calculates routes and avoids obstacles in real time, and guides the user to a specified destination.

[0241] As described above, by acquiring and analyzing the necessary input data at each step and generating appropriate output, the system of the present invention can smoothly answer the user's questions and provide guidance. The system of the present invention provides visitors with fast and accurate guidance.

[0242] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0243] The present invention relates to a multimodal AI system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine to achieve natural interaction with users. Specific embodiments of the system are described below.

[0244] System initialization

[0245] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[0246] User Awareness

[0247] The robot uses a camera to recognize the user. During this process, the camera captures the user's face, facial expressions, and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[0248] Emotion recognition

[0249] The robot uses an emotion engine to analyze the user's emotions from the captured facial expressions and voice. The emotion engine analyzes the data from each facial expression to determine whether the user is happy, confused, tired, etc.

[0250] Start a conversation

[0251] The robot uses the user information obtained by the recognition means and emotion engine to initiate a conversation using natural language processing means. First, it retrieves the user profile from a database and generates an appropriate greeting based on that information. This greeting is adjusted taking into account the user's emotional state. This allows the user to begin a natural dialogue with the system.

[0252] Answering questions

[0253] When a user inputs a question into the robot, the question is analyzed using natural language processing. During the analysis process, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. Furthermore, by incorporating the analysis results of the emotion engine into the response, the robot responds with an appropriate tone and content according to the user's emotional state.

[0254] Autonomous Mobility and Navigation

[0255] The robot can autonomously navigate to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path, safely reaching the destination while avoiding obstacles and other people along the way. It can also adjust the tone and content of its guidance depending on the user's emotional state during the journey.

[0256] Specific examples

[0257] Example 1: Store directions

[0258] If a user asks, "Where is the reception desk?", the robot will capture the user's face and facial expressions with a camera and analyze them using image recognition and an emotion engine. It will then analyze the intent of the question using natural language processing, and the system will generate an appropriate response. The robot will respond, "The reception desk is on the left side of this floor. I'll show you," and actually guide the user to the reception desk. If the robot determines that the user appears confused, it can add a particularly detailed explanation.

[0259] Example 2: Bank balance inquiry

[0260] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and responds, "Your current account balance is 100,000 yen." If the user is nervous, the robot can also add additional comments to help them relax.

[0261] Through these functions, the system of the present invention realizes natural interaction with users and provides a variety of services. Utilizing the emotion engine enables advanced responses that take user emotions into consideration, significantly improving the quality of services.

[0262] The processing flow will be explained below.

[0263] Step 1:

[0264] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[0265] Step 2:

[0266] The robot utilizes a camera to capture the user's face, facial expressions, and gestures, which are then analyzed using image recognition to obtain the user's identity, emotions, and gestures.

[0267] Step 3:

[0268] The robot uses an emotion engine to analyze the user's emotional state, extracting emotional data from facial expressions, voice tone, and gestures to determine whether the user is happy, confused, tired, etc.

[0269] Step 4:

[0270] The robot uses natural language processing to initiate a conversation based on the acquired user information. First, it retrieves the user profile from a database and generates a greeting based on that information. The greeting is adjusted to reflect the user's emotional state.

[0271] Step 5:

[0272] When a user types a question into the robot, the question is analyzed using natural language processing, which extracts the intent of the question and related entities to clearly understand the question.

[0273] Step 6:

[0274] Based on the analysis results, the robot retrieves appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. At this time, the robot will take into account the user's emotional state obtained from the emotion engine and adjust the tone and content of the response.

[0275] Step 7:

[0276] When a user requests guidance to a specified location, the robot begins navigation. First, it calculates the optimal path to the destination and autonomously moves along that path. During the journey, it navigates safely while avoiding obstacles and other people. It also monitors the user's emotional state in real time and provides guidance accordingly.

[0277] Step 8:

[0278] The robot records the results of its interactions with the user in a database. Dialogue history and emotional data are accumulated and used for future interactions. This allows the system to continuously learn and improve the quality of its services.

[0279] Specific examples

[0280] Example 1: Store directions

[0281] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expressions, and analyzes them using image recognition and an emotion engine. Using the user's emotional state and profile information obtained based on the analysis, the robot generates a response using natural language processing. The robot responds in a friendly tone, saying, "The reception desk is on the left side of this floor. I'll show you there," and begins navigation.

[0282] Example 2: Bank balance inquiry

[0283] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from a database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and generate a response accordingly. It responds in a calm tone, saying, "Your current account balance is 100,000 yen," and can also add additional comments to ease the tension during the question.

[0284] This embodiment allows users to receive flexible responses that take their emotions into consideration, improving the quality of service.

[0285] Example 2

[0286] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0287] Conventional dialogue systems have limited interaction with users and lack the ability to respond with emotion or to handle diverse languages. Furthermore, their user recognition and autonomous movement capabilities are limited, making it difficult to achieve natural and effective dialogue. This results in a poor user experience and reduced system utilization efficiency.

[0288] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an image recognition means, a natural language processing means, a multilingual translation means, a database connection means, an emotion analysis means, a user recognition means, and an autonomous movement means. This makes it possible to recognize the user's face and facial expression, analyze questions, generate information in multiple languages ​​while taking into account the emotional state, acquire a user profile and dialogue history, and autonomously move to a specified destination.

[0289] "Image recognition means" refers to a means of acquiring image data such as a user's face, facial expressions, and gestures using a camera or sensor, and analyzing that data.

[0290] A "natural language processing means" is a means for analyzing text or voice questions entered by a user, understanding their intent, and generating an appropriate response.

[0291] The "multilingual translation means" is a means for translating text input in different languages ​​and generating information corresponding to multiple languages.

[0292] "Database connection means" refers to a means for accessing a database to obtain and store related data such as user profiles and interaction history.

[0293] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data to identify their emotional state.

[0294] "User recognition means" refers to a means for identifying a user's ID and distinguishing between individual users.

[0295] An "autonomous vehicle" is a vehicle that moves autonomously to a specified destination while avoiding obstacles.

[0296] The present invention relates to a system that integrates image recognition means, natural language processing means, multilingual translation means, database connection means, emotion analysis means, user recognition means, and autonomous mobility means to realize natural interactions with users. Detailed embodiments of the system are described below.

[0297] System configuration

[0298] Hardware

[0299] This system uses the following hardware components:

[0300] Camera (e.g. high-resolution webcam)

[0301] microphone

[0302] speaker

[0303] Robot body (including motors and sensors for autonomous movement)

[0304] server

[0305] Database (e.g. MySQL, PostgreSQL)

[0306] software

[0307] The system consists of the following software components:

[0308] Image recognition models (e.g., TensorFlow, PyTorch)

[0309] Natural language processing models (e.g., GPT-4)

[0310] Multilingual translation models (e.g., Google Translate API)

[0311] Sentiment analysis engine (e.g. Affectiva SDK)

[0312] Database management systems (e.g., MySQL, PostgreSQL)

[0313] Autonomous movement algorithms (e.g., SLAM technology)

[0314] System Operation

[0315] System initialization

[0316] The server first initializes the entire system. This initialization includes loading image recognition models, natural language processing models, multilingual translation models, and a sentiment analysis engine, as well as setting up a database. Specifically, the server loads trained models into memory using TensorFlow or PyTorch, and initializes the sentiment analysis engine using the Affectiva SDK. The database uses MySQL or PostgreSQL, and is prepared to store user profiles and interaction history.

[0317] User Awareness

[0318] The robot recognizes the user through a camera. This involves capturing the user's face, facial expressions, and gestures, and then analyzing the data using libraries such as OpenCV. An emotion analysis engine is also integrated into this process to analyze the user's emotional state. The user's ID and emotional state are then used as the basis for the system to provide personalized responses.

[0319] Emotion recognition

[0320] The robot analyzes the user's emotions from the captured facial expressions and voice. It uses an emotion analysis engine to identify specific emotions (e.g., happiness, sadness, surprise), allowing the system to understand the user's current emotional state and prepare to generate a response accordingly.

[0321] Start a conversation

[0322] The robot initiates a conversation using an NLP model based on the user's profile and emotional state. The server queries a database to retrieve the user's profile. Using an NLP model (e.g., GPT-4), it generates an appropriate greeting for the user and speaks it through the speaker. The greeting is adjusted based on the user's emotional state.

[0323] Answering questions

[0324] When a user types a question into the robot, it is analyzed using a natural language processing model. Analysis includes extracting the intent of the question and relevant entities. The server then references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer, which the robot then speaks in a tone that best suits its emotional state.

[0325] Autonomous Mobility and Navigation

[0326] The robot autonomously navigates to a destination specified by the user. It uses an algorithm based on SLAM technology to calculate the optimal path, and moves along that path while avoiding obstacles. Along the way, it adjusts the tone and content of its guidance according to the user's emotional state.

[0327] Examples of specific examples and prompts

[0328] Example 1: Store directions

[0329] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and analyzes them using image recognition and an emotion engine. It then analyzes the intent of the question using natural language processing, generates and speaks an appropriate response, and guides the user to the reception desk, providing detailed explanations if the user appears confused.

[0330] Example prompt for a generative AI model: "The user is asking where the reception desk is. The robot will politely provide directions. Please respond especially politely if the user seems confused."

[0331] Example 2: Bank balance inquiry

[0332] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves the account information from the database. If authentication is successful, the emotion engine analyzes the user's emotional state and responds with the current account balance. If the user is nervous, it adds a comment to relax them.

[0333] Example prompt for a generative AI model: "The user asks for their account balance. The robot responds by adding a relaxing comment based on the user's emotional state."

[0334] The above is a specific embodiment of the present invention, and the system allows users to experience natural conversations that take emotion into consideration.

[0335] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0336] Step 1: Initialize the system

[0337] The server initializes the entire system. First, it loads an image recognition model into memory using TensorFlow or PyTorch. Next, it loads a natural language processing (NLP) model such as OpenAI's GPT-4, and then initializes a multilingual translation model such as the Google Translate API and a sentiment analysis engine using the Affectiva SDK. It also sets up a database using MySQL or PostgreSQL to store user profiles and interaction history. This completes the initialization process, making the various models ready for processing. It receives model configuration files and database configuration information as input, and the model loading is complete as output.

[0338] Step 2: Getting started with user awareness

[0339] The robot recognizes the user through a camera. Specifically, it uses the camera to capture the user's face, facial expressions, and gestures, and analyzes the data using the OpenCV library. As a result of the analysis, the user's facial feature points are extracted, and the emotion analysis engine analyzes the user's emotional state. The input is the captured image data, and the output is the user's ID and emotional state.

[0340] Step 3: Detailed analysis of emotional state

[0341] The robot performs a detailed analysis of the user's emotions from the captured facial expressions and voice. Using the Affectiva SDK, it analyzes facial expression data to identify specific emotions (such as joy, sadness, or surprise). It also collects the user's voice data and analyzes its emotional nuances. It receives facial expression and voice data as input and outputs the identified emotional state.

[0342] Step 4: Start a conversation

[0343] The robot starts a conversation using an NLP model based on the user's profile and emotional state. The server executes an SQL query to retrieve the user profile from the database. Using the NLP model (GPT-4), it generates an appropriate greeting based on the user's profile and emotional state. It then speaks this to the user through the speaker. The input is the user profile and emotional state retrieved from the database, and the output is the generated greeting.

[0344] Step 5: Answer questions

[0345] The user inputs a question to the robot. The user uses a voice recognition system to input the question in text format. The server receives the text data and uses an NLP model to analyze the intent of the question and extract relevant entities. Based on this, it references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer. It then takes into account the results of the sentiment analysis engine to adjust the tone and content of the response and speaks the answer to the user through the speaker. The input is the user's question text, and the output is the generated answer.

[0346] Step 6: Autonomous Movement and Navigation

[0347] The robot moves autonomously to a destination specified by the user. First, the user inputs instructions for the destination. The robot uses SLAM technology to calculate the optimal path to the destination and moves along that path. During movement, it uses LiDAR and ultrasonic sensors to detect obstacles and avoid them to proceed safely. It also adjusts the tone and content of its guidance according to the user's emotional state. It receives destination instructions and sensor data as input, and outputs the movement results along the optimal path.

[0348] (Application example 2)

[0349] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0350] Today's elderly often have difficulty communicating with store staff and navigating the store when shopping or using services in physical stores. Furthermore, language barriers and a lack of emotional support reduce the elderly's satisfaction. This reduces the opportunities for elderly people to enjoy shopping in physical stores and causes stress, which is an issue.

[0351] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0352] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. This enables elderly people to shop or use services in brick-and-mortar stores by recognizing faces and facial expressions, analyzing questions, generating language-compatible information, acquiring profiles and dialogue histories, and analyzing emotional states, thereby providing appropriate responses and in-store navigation based on their emotions.

[0353] "Image recognition means" refers to technology that uses a device such as a camera to recognize a user's face, facial expressions, and gestures.

[0354] "Natural language processing means" is a technology that analyzes questions and requests entered by users and understands their intentions.

[0355] "Multilingual translation means" refers to technology for translating and generating information in accordance with the language used by the user.

[0356] "Database connection means" refers to technology for connecting to a database that stores user profiles and interaction history, and obtaining the necessary information.

[0357] The "emotion engine" is a technology that analyzes the user's emotional state from their facial expressions and tone of voice to determine what emotions the user is feeling.

[0358] A "system" is a collection of components that integrate multiple technical means to provide specific functions or services.

[0359] "In-store navigation" is a function that guides users to find desired locations and products within a store.

[0360] "Support for the elderly" means providing support and services to help the elderly live their daily lives more conveniently and safely.

[0361] "Multilingual support" means providing services and information in different languages ​​to users who speak different languages.

[0362] "Appropriate response" means providing answers or guidance appropriate to the situation based on the user's question or condition.

[0363] A "profile" is data that records basic information and personal characteristics about a user.

[0364] System configuration

[0365] This invention is realized by a system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. The components of this system are described in detail below.

[0366] Hardware and Software

[0367] Image Recognition: The system recognizes the user's face and facial expressions using a camera-equipped smartphone or head-mounted display. It uses the OpenCV (cv2) library and a model trained in TensorFlow.

[0368] Natural Language Processing: Natural language processing uses the Transformers library pipeline to analyze question answers. Generative AI models such as BERT and RoBERTa are used.

[0369] Multilingual translation method: For multilingual translation, the DeepTranslator library is used, and natural translation is provided in conjunction with the Google Translate API.

[0370] Database connection method: User profiles and interaction history are managed using SQLite, and the necessary information is obtained.

[0371] Emotion engine: Emotion analysis uses the Transformers library pipeline to analyze emotions from voice and facial expressions.

[0372] System initialization

[0373] Initialization takes place on the server, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing user profiles and interaction history.

[0374] User recognition and sentiment analysis

[0375] The device's camera captures the user's face, facial expressions, and gestures, and analyzes them using image recognition.The emotion engine analyzes the user's emotional state and determines how they are feeling.

[0376] Initiating conversations and answering questions

[0377] The system analyzes questions using natural language processing and generates an appropriate greeting based on the user's profile. It also adjusts the tone and content of the response based on the user's emotional state, using an emotion engine. The system offers conversations that are especially designed to ensure the elderly can use the service with peace of mind.

[0378] In-store navigation

[0379] Based on the user's gestures, the system guides the user to the desired location or product within the store. The system also reflects the user's emotional state during navigation, adjusting the navigation to ensure a safe and secure journey.

[0380] Examples and prompts

[0381] For example, if an elderly person asks, "Do you have the shirt I'm looking for in stock?", the smartphone camera will recognize the user's face and perform image recognition and emotion analysis. Natural language processing will then be used to obtain the appropriate stock information, which will be translated into multiple languages ​​and displayed to the user.

[0382] Prompt Sentence Examples

[0383] "Do you have the shirt I'm looking for in stock?"

[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0385] Step 1:

[0386] The server initializes the entire system, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing the user profile and interaction history.

[0387] Input: Image recognition model, natural language processing model, multilingual translation model, emotion engine, database

[0388] Output: Initialized models and database connections

[0389] Step 2:

[0390] The device uses a camera to capture the user's face, facial expressions, and gestures, and the captured image data is analyzed using image recognition tools.

[0391] Input: Video data of the user's face, expressions, and gestures

[0392] Output: Parsed user ID and emotional state

[0393] Step 3:

[0394] Based on the analysis results, the device further analyzes the user's emotional state using an emotion engine, thereby specifically grasping the user's emotional state.

[0395] Input: User's video data and analysis results

[0396] Output: Detailed emotional state data

[0397] Step 4:

[0398] The server retrieves the user profile and interaction history from the database and generates an appropriate greeting for the user. It uses natural language processing means to provide a greeting based on the user's emotional state.

[0399] Input: User ID, emotional state data, user profile and interaction history from database

[0400] Output: Emotion-based greeting

[0401] Step 5:

[0402] The user inputs a question into the terminal. The server analyzes the question using natural language processing to extract the intent of the question and related information. It then generates an appropriate answer based on the analysis results.

[0403] Input: User question

[0404] Output: Analyzed question intent and answer suggestions

[0405] Step 6:

[0406] The server uses a multilingual translation means to translate the generated answer into the user's language, and the translated result is displayed to the user.

[0407] Input: Analyzed question intent and answer suggestions

[0408] Output: The answer translated into the user's language

[0409] Step 7:

[0410] The device monitors the user's gestures and provides in-store navigation as needed, adjusting the tone and content of the guidance based on the user's emotional state.

[0411] Input: User gesture data and emotional state

[0412] Output: Navigation prompts and tuned tones

[0413] Through this series of steps, the system can provide an environment where seniors can shop and use services in physical stores in a natural and safe manner.

[0414] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0415] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0416] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0417] [Second embodiment]

[0418] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0419] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0420] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0421] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0422] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0423] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0424] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0425] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0426] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0427] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0428] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0429] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0430] The present invention relates to a multimodal AI system using a humanoid robot. This system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. Specific embodiments of the system are described below.

[0431] System initialization

[0432] The server performs the initial setup of the system, including loading the image recognition model, natural language processing (NLP) model, and multilingual translation model, and setting up the database, so that the entire system is ready to run efficiently and smoothly.

[0433] User Awareness

[0434] The robot uses a camera to recognize the user. During this process, the camera captures the user's face and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[0435] Start a conversation

[0436] The robot uses the user information obtained by the recognition means to initiate a conversation using natural language processing means. First, it retrieves the user profile from the database and generates an appropriate greeting based on that information. This allows the user to begin a natural dialogue with the system.

[0437] Answering questions

[0438] When a user types a question into the robot, it analyzes it using natural language processing. During the analysis, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. For example, if a user asks about the weather, the system retrieves weather information and provides it to the user.

[0439] Autonomous Mobility and Navigation

[0440] The robot can move autonomously to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path. During this process, it is designed to reach the destination while avoiding obstacles and other people. For example, if the user says, "Please guide me to the reception desk," the robot will calculate the route to the reception desk and actually begin guiding the user.

[0441] Specific examples

[0442] Example 1: Store directions

[0443] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expression, and analyzes it using image recognition. The robot then analyzes the intent of the question using natural language processing, and the system generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and actually guides the user to the reception desk.

[0444] Example 2: Bank balance inquiry

[0445] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot responds, "My current account balance is 100,000 yen." This allows the user to obtain information quickly and accurately.

[0446] Through these functions, the system of the present invention aims to achieve natural interaction with users and highly efficient service provision, thereby reducing workload and improving service quality.

[0447] The processing flow will be explained below.

[0448] Step 1:

[0449] The server initializes the entire system. First, it loads the image recognition, natural language processing, and multilingual translation models. Next, it sets up the database used by the system and prepares it to store user profiles and interaction history.

[0450] Step 2:

[0451] The robot uses a camera to capture images to recognize the user's face and facial expressions, and the captured images are passed to an image recognition means to analyze the user's face, emotions, and gestures.

[0452] Step 3:

[0453] Based on the analysis results, the robot retrieves the user's profile from the database, allowing it to refer to each user's past interaction history and specific information.

[0454] Step 4:

[0455] The robot uses natural language processing tools to initiate a conversation with the user, generating appropriate greetings and introductory phrases based on the user's profile, emotions, and gestures, and then delivering them to the user via voice or text.

[0456] Step 5:

[0457] When a user types a question into the robot, the question is analyzed using natural language processing tools, which extracts the intent of the question and related entities, thereby providing a clear understanding of the question.

[0458] Step 6:

[0459] Based on the analysis results, the robot retrieves the appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. The response is then provided to the user.

[0460] Step 7:

[0461] When a user requests guidance to a specific location, the robot begins navigation to the destination. First, it calculates the optimal path to the destination and then autonomously moves along that path, avoiding obstacles and other people along the way to safely reach the destination.

[0462] Step 8:

[0463] The robot records the results of its interactions with the user and their actions in a database, which can be used for future interactions. This record includes the interaction history, the user's reactions, and any new information acquired. This allows the system to continuously learn and improve the quality of its services.

[0464] In this way, the system of the present invention realizes natural interaction with the user and provides a variety of services.

[0465] Example 1

[0466] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0467] Conventional humanoid robot systems have had problems with natural interaction with users and efficient service provision. Specifically, they suffer from low recognition accuracy, delayed responses, and limited ability to support different languages. Furthermore, there is still room for improvement in terms of recognizing users' faces and gestures, identifying emotions, providing appropriate information, and navigation.

[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0469] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and initial setting means. This allows the server to recognize the user's face and facial expressions, identify their ID and emotions, and provide appropriate information in multiple languages, thereby realizing natural interaction with the user and enabling efficient service provision.

[0470] "Image recognition means" refers to means that has the function of analyzing image data acquired using an imaging device such as a camera and recognizing a specific object (for example, a user's face or gesture).

[0471] "Natural language processing means" refers to means that include algorithms and models for analyzing text data entered by a user and understanding their intent and emotions.

[0472] A "multilingual translation means" is a means that has the function of translating text between different languages, and that understands input from a user in different languages ​​and generates corresponding information.

[0473] "Database connection means" refers to a means that has the function of communicating with a database server and acquiring and saving the necessary information.

[0474] The "means for performing initial settings" refers to a means for performing basic settings for system operation at the start, including loading various models and setting up a database.

[0475] "Means for recognizing a user's face and facial expressions" refers to means that has the function of extracting specific features from a user's facial image captured by a camera, identifying the user based on the results, and determining their emotional state.

[0476] "Means for identifying a user's ID and emotions" refers to a means for analyzing data obtained through image recognition to determine who the user is (ID) and their emotional state at the time.

[0477] "Means for obtaining a user profile" means means capable of obtaining information related to a user from a database, including past conversation history and preferences.

[0478] The "means for generating a greeting message" is a means having a function for automatically generating a greeting message suitable for a user based on the acquired user profile.

[0479] "Means for analyzing the intent of a question and related entities" refers to a means for analyzing the question entered by the user and identifying the intent of the question and related keywords (entities).

[0480] The "means for generating an appropriate answer to a question" is a means for obtaining relevant information based on the intent and entities of the analyzed question, and automatically generating an appropriate answer based on that information.

[0481] The "means for continuously capturing image data" refers to a means having a function for activating a camera and continuously acquiring image data at regular intervals.

[0482] The "means for preprocessing image data" refers to a means that has the function of performing preprocessing such as resizing, noise removal, and color conversion on the acquired image data.

[0483] "Means for processing sensor information in real time and moving autonomously" refers to means that analyze information from various sensors in real time, calculate a route according to the environment, and move autonomously.

[0484] The present invention relates to a multimodal AI system using a humanoid robot. The system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. It also includes means for recognizing a user's face and facial expressions, identifying the user's ID and emotion, acquiring a user profile, generating a greeting, analyzing the intent of a question and related entities, generating appropriate answers to the question, continuously capturing image data, preprocessing the image data, and processing sensor information in real time to move autonomously.

[0485] 1. System initialization

[0486] The server performs the initial system configuration. First, for image recognition, it loads a ResNet50 model using TensorFlow using an NVIDIA GPU. Next, for natural language processing, it loads the BERT model using Hugging Face's Transformers library. It also configures the Google Translate API for multilingual translation, and starts a MySQL database to create the necessary tables and indexes.

[0487] 2. User Awareness

[0488] The robot captures the user's face and gestures using a built-in camera (e.g., a standard HD camera). The image data is preprocessed with OpenCV and then analyzed using a ResNet50 model running on NVIDIA GPUs, which identifies the user's identity, face, and facial expressions.

[0489] 3. Start a conversation

[0490] The robot starts a conversation using the data mentioned above. First, it uses the user's ID to retrieve the user profile from a MySQL database. Then, it uses that information to generate an appropriate greeting using Hugging Face's BERT model. For example, it might generate a greeting like, "Hello, Tanaka-san. How's your day?"

[0491] 4. Answering questions

[0492] Users input questions into the robot, which then uses voice input (e.g., a standard microphone) to receive the question and convert it into text using the Google Speech-to-Text API. It then uses Hugging Face's BERT model to analyze the intent of the question and relevant entities, and based on the analysis results, queries the appropriate API (e.g., the OpenWeatherMap API) or database to generate an answer to the question.

[0493] 5. Autonomous Movement and Navigation

[0494] The robot autonomously moves to a location specified by the user. First, the user specifies the destination via voice input or touchscreen. Next, the robot uses SLAM technology to generate a map of the environment and calculates the optimal route using algorithms such as the A algorithm. The robot autonomously moves along the calculated route using a LiDAR sensor and motor control unit, avoiding obstacles and reaching the destination.

[0495] Specific examples

[0496] In-store directions

[0497] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and first identifies the user using image recognition. It then analyzes the intent of the question using natural language processing and generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and autonomously guides the user to the reception desk using SLAM technology and a motor control unit.

[0498] Bank balance inquiry

[0499] When a user asks, "What is my account balance?", the robot first authenticates the user and retrieves account information from the database. After successful authentication, the robot responds, "Your current account balance is 100,000 yen."

[0500] Prompt Sentence Examples

[0501] "When a user asks for weather information, how do you get the information from the API and respond?"

[0502] "Please explain in detail the algorithm that recognizes user emotions."

[0503] As explained above, the system of the present invention realizes natural interaction with users and aims to reduce the workload and improve service quality through highly efficient service provision.

[0504] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0505] Step 1: Initialize the system

[0506] The server performs the initial system setup, inputting the image recognition model, natural language processing model, multilingual translation model, and database configuration information.

[0507] 1. As a deep learning model for image recognition, we load a ResNet50 model in TensorFlow using an NVIDIA GPU. The input is the TensorFlow model file, and the output is the model loaded in memory.

[0508] 2. Load the BERT model from Hugging Face's Transformers library as the natural language processing model. The input is the model file, and the output is the BERT model loaded in memory.

[0509] 3. Set up the Google Translate API as a multilingual translation model. The input is the API key and configuration information, and the output is the API availability status.

[0510] 4. Starts a MySQL database and creates the necessary tables and indexes. The input is database configuration information, and the output is a configured database ready to use.

[0511] Step 2: User Awareness

[0512] The robot uses a camera to recognize the user, and the input is image data obtained from the camera.

[0513] 1. Start a camera (e.g., HD camera) and continuously capture image data. The input is the camera sensor data, and the output is the captured image data.

[0514] 2. The acquired image data is preprocessed using the OpenCV library, which includes resizing, noise removal, color conversion, etc. The input is raw image data, and the output is preprocessed image data.

[0515] 3. The preprocessed image data is input to a ResNet50 model running on an NVIDIA GPU to identify the user's ID and emotion. The input is the preprocessed image data, and the output is the identified user's ID and emotion information.

[0516] Step 3: Start a conversation

[0517] The robot starts a conversation based on the recognized user information, including the user's ID and emotional information.

[0518] 1. The robot uses the user's ID to retrieve the user profile from the MySQL database. The input is the user's ID and the output is the retrieved user profile.

[0519] 2. Based on the retrieved user profile, the robot uses Hugging Face's BERT model to generate an appropriate greeting, such as "Hello, Tanaka-san. How are you today?" The input is the user profile, and the output is the generated greeting.

[0520] Step 4: Answer questions

[0521] The user inputs a question to the robot, and the input is voice data.

[0522] 1. Collect user questions using a voice input function (e.g., microphone). The input is voice data, and the output is the question data captured as voice.

[0523] 2. The robot converts voice data into text data using the Google Speech-to-Text API. The input is voice data and the output is text data.

[0524] 3. The text data is analyzed using Hugging Face's BERT model to extract the question intent and related entities. The input is the text data, and the output is the analysis results.

[0525] 4. Based on the analysis results, the robot queries the appropriate API (e.g., OpenWeatherMap API) or database to generate an answer to the question. The input is the analysis results and the necessary API information, and the output is the generated answer.

[0526] Step 5: Autonomous Movement and Navigation

[0527] The robot moves autonomously to a specified location, and the input is a destination specified by the user.

[0528] 1. The user specifies a destination using voice input or a touch screen. The input is the destination information, and the output is the specified destination.

[0529] 2. The robot uses SLAM technology to generate a map of the environment and calculate the optimal path. The input is sensor information and destination information, and the output is the calculated path information.

[0530] 3. The robot uses a LiDAR sensor and a motor control unit to move autonomously based on a calculated path. The input is path information and real-time sensor information, and the output is autonomous movement behavior.

[0531] (Application example 1)

[0532] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0533] Modern stores and public facilities require a means for visitors to quickly and accurately find their destination and obtain information. However, conventional guidance systems and information centers face problems such as labor shortages and language barriers. Furthermore, visitors often have to search for the information they need themselves, which is inconvenient. A system that solves these issues and allows visitors to easily reach their destination is needed.

[0534] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0535] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, and database connection means. This enables the server to recognize a user's face and facial expressions, analyze the user's questions, generate information corresponding to the user's language, and acquire a user profile and dialogue history. The server also includes connection means with a smart device, means for activating and controlling the robot via the smart device, autonomous robot movement means, and means for guiding the robot to a user-specified destination using the autonomous movement means. This allows visitors to easily call the robot from their smartphones, and the robot automatically guides the visitor to their destination, significantly improving visitor convenience.

[0536] "Image recognition means" refers to a device or system that uses a camera or sensor to recognize a user's face, facial expressions, and gestures.

[0537] "Natural language processing means" is a technology that analyzes questions and instructions uttered by users and understands their intent and content.

[0538] "Multilingual translation means" refers to technology that translates between different languages ​​and generates information corresponding to the language used by the user.

[0539] "Database connection means" refers to technology for accessing a database that stores user profiles, interaction history, and other necessary data, and acquiring that information.

[0540] "Means for connecting with smart devices" refers to technology that connects devices such as smartphones and tablets to the system, enabling two-way communication.

[0541] "Means for activating and controlling robots" refers to technology for activating robots via smart devices and controlling their movements and functions.

[0542] "Autonomous robot mobility means" refers to technology that enables a robot to move autonomously while recognizing its surrounding environment.

[0543] "Generating information" is the process of generating appropriate answers or guidance information based on the user's questions or instructions.

[0544] A "user profile" is data that compiles information about a user (e.g., the user's ID, past interaction history, preferences, etc.).

[0545] "Dialogue history" is data that records past conversations and interactions with a user.

[0546] "Autonomous movement" refers to a robot's ability to move while adapting to its surrounding environment based on its own judgment.

[0547] The present invention relates to a system that controls a robot in cooperation with a smart device to provide guidance and information to users in stores and public facilities. This system is realized using the following various means.

[0548] 1. Image Recognition Methods:

[0549] The server uses cameras and sensors to recognize the user's face, facial expressions, and gestures, using software libraries such as OpenCV and TensorFlow, which are known for their image processing technology.

[0550] 2. Natural Language Processing Tools:

[0551] The server uses natural language processing techniques to analyze the user's questions and instructions. This process uses various NLP (Natural Language Processing) libraries and tools (e.g., Google's Dialogflow, OpenAI's GPT-3).

[0552] 3. Multilingual translation tools:

[0553] The server uses a multilingual translation engine to translate between different languages ​​and generate information corresponding to the user's language, for example, the Google Translate API.

[0554] 4. Database connection method:

[0555] The server retrieves information such as the user's profile and interaction history from a database, using Firebase or AWS RDS to dynamically manage user information.

[0556] 5. Connectivity with smart devices:

[0557] The smartphone application communicates with the robot via Bluetooth or Wi-Fi to activate and control it, allowing users to summon the robot using their smartphone.

[0558] 6. Autonomous robotic mobility:

[0559] The robot autonomously navigates to its destination using SLAM (Simultaneous Localization and Mapping) technology with Lidar sensors and cameras, using ROS Lidar and Google Cartographer.

[0560] 7. Information generation:

[0561] The server generates appropriate answers and guidance information in response to user questions and instructions, using a generative AI model in the process.

[0562] 8. User Awareness and Guidance:

[0563] The robot recognizes the user with a camera, analyzes the intent of the question using natural language processing technology, and provides guidance. If the user asks, "Where is the dairy section?", the robot will use the store's map information to answer, "The dairy section is on the right side of this floor. I'll show you." and begin guiding the user. Using its autonomous mobility, the robot will accurately guide the user to their destination.

[0564] 9. Example prompt:

[0565] An example of a prompt sentence to input to the generative AI model is as follows:

[0566] "A user is in a store asking about the dairy section. The robot recognizes the user with a camera and analyzes the intent of the question using NLP technology. It then retrieves store map information from a database and provides directions. Please explain the specific processing flow in detail."

[0567] In this way, the system of the present invention can provide an environment in which the user can quickly and accurately obtain information regardless of the environment.

[0568] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0569] Step 1:

[0570] The user starts the smartphone app and presses the "Start Guide Robot" button, at which point the smartphone uses its camera and microphone to check the surrounding environment.

[0571] input:

[0572] User presses a button

[0573] output:

[0574] Camera and microphone ambient data

[0575] Specific behavior:

[0576] The camera captures image data of the surroundings, and the microphone captures audio data.

[0577] Step 2:

[0578] The device calls the robot via Bluetooth or Wi-Fi and sends commands for the robot to move to the user's location.

[0579] input:

[0580] Camera and microphone data

[0581] output:

[0582] Robot commands

[0583] Specific behavior:

[0584] It connects to the robot communication module via Bluetooth or Wi-Fi and sends the command "move to the user."

[0585] Step 3:

[0586] The robot recognizes the user using a camera and analyzes the user's face and gestures using image recognition means.

[0587] input:

[0588] Image data acquired by the camera

[0589] output:

[0590] User ID, facial expression, and gesture information

[0591] Specific behavior:

[0592] Image processing libraries such as OpenCV and TensorFlow are used to recognize and analyze the user's face and gestures and obtain user information.

[0593] Step 4:

[0594] The server uses natural language processing means to analyze the user's question and understand its intent.

[0595] input:

[0596] User utterances

[0597] output:

[0598] User questions and their intent

[0599] Specific behavior:

[0600] NLP tools such as Google's Dialogflow and OpenAI's GPT-3 are used to interpret the content of the speech and extract the intent of the question.

[0601] Step 5:

[0602] The server uses database connectivity to obtain the user's profile and interaction history and generates appropriate information.

[0603] input:

[0604] User ID, question content

[0605] output:

[0606] User profile and corresponding answer information

[0607] Specific behavior:

[0608] It connects to Firebase or AWS RDS, obtains the user's past interaction history and profile information, and then uses a generative AI model to generate appropriate answers and guidance information.

[0609] Step 6:

[0610] The server uses a multilingual translation means to translate the generated information into the language used by the user.

[0611] input:

[0612] Generated answer information

[0613] output:

[0614] Information translated into the user's language

[0615] Specific behavior:

[0616] Use the Google Translate API to translate the generated information into the user's language.

[0617] Step 7:

[0618] The robot presents appropriate information to the user via voice or text and begins providing guidance in response to the user's questions.

[0619] input:

[0620] Translated information

[0621] output:

[0622] User guidance information

[0623] Specific behavior:

[0624] The text information is converted into speech using a speech synthesis engine (e.g., Google TTS) and provided to the user. The autonomous mobility system is also used to calculate the route to the user's specified destination and begin providing guidance.

[0625] Step 8:

[0626] The robot uses autonomous mobility to guide the user to their destination while avoiding obstacles.

[0627] input:

[0628] User-specified destination information and surrounding environment information

[0629] output:

[0630] Guidance to the destination, reaching the final point

[0631] Specific behavior:

[0632] It uses Lidar sensors and SLAM technology (e.g., ROS Lidar, Google Cartographer) to map the surrounding environment, calculates routes and avoids obstacles in real time, and guides the user to a specified destination.

[0633] As described above, by acquiring and analyzing the necessary input data at each step and generating appropriate output, the system of the present invention can smoothly answer the user's questions and provide guidance. The system of the present invention provides visitors with fast and accurate guidance.

[0634] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0635] The present invention relates to a multimodal AI system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine to achieve natural interaction with users. Specific embodiments of the system are described below.

[0636] System initialization

[0637] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[0638] User Awareness

[0639] The robot uses a camera to recognize the user. During this process, the camera captures the user's face, facial expressions, and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[0640] Emotion recognition

[0641] The robot uses an emotion engine to analyze the user's emotions from the captured facial expressions and voice. The emotion engine analyzes the data from each facial expression to determine whether the user is happy, confused, tired, etc.

[0642] Start a conversation

[0643] The robot uses the user information obtained by the recognition means and emotion engine to initiate a conversation using natural language processing means. First, it retrieves the user profile from a database and generates an appropriate greeting based on that information. This greeting is adjusted taking into account the user's emotional state. This allows the user to begin a natural dialogue with the system.

[0644] Answering questions

[0645] When a user inputs a question into the robot, the question is analyzed using natural language processing. During the analysis process, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. Furthermore, by incorporating the analysis results of the emotion engine into the response, the robot responds with an appropriate tone and content according to the user's emotional state.

[0646] Autonomous Mobility and Navigation

[0647] The robot can autonomously navigate to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path, safely reaching the destination while avoiding obstacles and other people along the way. It can also adjust the tone and content of its guidance depending on the user's emotional state during the journey.

[0648] Specific examples

[0649] Example 1: Store directions

[0650] If a user asks, "Where is the reception desk?", the robot will capture the user's face and facial expressions with a camera and analyze them using image recognition and an emotion engine. It will then analyze the intent of the question using natural language processing, and the system will generate an appropriate response. The robot will respond, "The reception desk is on the left side of this floor. I'll show you," and actually guide the user to the reception desk. If the robot determines that the user appears confused, it can add a particularly detailed explanation.

[0651] Example 2: Bank balance inquiry

[0652] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and responds, "Your current account balance is 100,000 yen." If the user is nervous, the robot can also add additional comments to help them relax.

[0653] Through these functions, the system of the present invention realizes natural interaction with users and provides a variety of services. Utilizing the emotion engine enables advanced responses that take user emotions into consideration, significantly improving the quality of services.

[0654] The processing flow will be explained below.

[0655] Step 1:

[0656] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[0657] Step 2:

[0658] The robot utilizes a camera to capture the user's face, facial expressions, and gestures, which are then analyzed using image recognition to obtain the user's identity, emotions, and gestures.

[0659] Step 3:

[0660] The robot uses an emotion engine to analyze the user's emotional state, extracting emotional data from facial expressions, voice tone, and gestures to determine whether the user is happy, confused, tired, etc.

[0661] Step 4:

[0662] The robot uses natural language processing to initiate a conversation based on the acquired user information. First, it retrieves the user profile from a database and generates a greeting based on that information. The greeting is adjusted to reflect the user's emotional state.

[0663] Step 5:

[0664] When a user types a question into the robot, the question is analyzed using natural language processing, which extracts the intent of the question and related entities to clearly understand the question.

[0665] Step 6:

[0666] Based on the analysis results, the robot retrieves appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. At this time, the robot will take into account the user's emotional state obtained from the emotion engine and adjust the tone and content of the response.

[0667] Step 7:

[0668] When a user requests guidance to a specified location, the robot begins navigation. First, it calculates the optimal path to the destination and autonomously moves along that path. During the journey, it navigates safely while avoiding obstacles and other people. It also monitors the user's emotional state in real time and provides guidance accordingly.

[0669] Step 8:

[0670] The robot records the results of its interactions with the user in a database. Dialogue history and emotional data are accumulated and used for future interactions. This allows the system to continuously learn and improve the quality of its services.

[0671] Specific examples

[0672] Example 1: Store directions

[0673] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expressions, and analyzes them using image recognition and an emotion engine. Using the user's emotional state and profile information obtained based on the analysis, the robot generates a response using natural language processing. The robot responds in a friendly tone, saying, "The reception desk is on the left side of this floor. I'll show you there," and begins navigation.

[0674] Example 2: Bank balance inquiry

[0675] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from a database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and generate a response accordingly. It responds in a calm tone, saying, "Your current account balance is 100,000 yen," and can also add additional comments to ease the tension during the question.

[0676] This embodiment allows users to receive flexible responses that take their emotions into consideration, improving the quality of service.

[0677] Example 2

[0678] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0679] Conventional dialogue systems have limited interaction with users and lack the ability to respond with emotion or to handle diverse languages. Furthermore, their user recognition and autonomous movement capabilities are limited, making it difficult to achieve natural and effective dialogue. This results in a poor user experience and reduced system utilization efficiency.

[0680] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an image recognition means, a natural language processing means, a multilingual translation means, a database connection means, an emotion analysis means, a user recognition means, and an autonomous movement means. This makes it possible to recognize the user's face and facial expression, analyze questions, generate information in multiple languages ​​while taking into account the emotional state, acquire a user profile and dialogue history, and autonomously move to a specified destination.

[0681] "Image recognition means" refers to a means of acquiring image data such as a user's face, facial expressions, and gestures using a camera or sensor, and analyzing that data.

[0682] A "natural language processing means" is a means for analyzing text or voice questions entered by a user, understanding their intent, and generating an appropriate response.

[0683] The "multilingual translation means" is a means for translating text input in different languages ​​and generating information corresponding to multiple languages.

[0684] "Database connection means" refers to a means for accessing a database to obtain and store related data such as user profiles and interaction history.

[0685] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data to identify their emotional state.

[0686] "User recognition means" refers to a means for identifying a user's ID and distinguishing between individual users.

[0687] An "autonomous vehicle" is a vehicle that moves autonomously to a specified destination while avoiding obstacles.

[0688] The present invention relates to a system that integrates image recognition means, natural language processing means, multilingual translation means, database connection means, emotion analysis means, user recognition means, and autonomous mobility means to realize natural interactions with users. Detailed embodiments of the system are described below.

[0689] System configuration

[0690] Hardware

[0691] This system uses the following hardware components:

[0692] Camera (e.g. high-resolution webcam)

[0693] microphone

[0694] speaker

[0695] Robot body (including motors and sensors for autonomous movement)

[0696] server

[0697] Database (e.g. MySQL, PostgreSQL)

[0698] software

[0699] The system consists of the following software components:

[0700] Image recognition models (e.g., TensorFlow, PyTorch)

[0701] Natural language processing models (e.g., GPT-4)

[0702] Multilingual translation models (e.g., Google Translate API)

[0703] Sentiment analysis engine (e.g. Affectiva SDK)

[0704] Database management systems (e.g., MySQL, PostgreSQL)

[0705] Autonomous movement algorithms (e.g., SLAM technology)

[0706] System Operation

[0707] System initialization

[0708] The server first initializes the entire system. This initialization includes loading image recognition models, natural language processing models, multilingual translation models, and a sentiment analysis engine, as well as setting up a database. Specifically, the server loads trained models into memory using TensorFlow or PyTorch, and initializes the sentiment analysis engine using the Affectiva SDK. The database uses MySQL or PostgreSQL, and is prepared to store user profiles and interaction history.

[0709] User Awareness

[0710] The robot recognizes the user through a camera. This involves capturing the user's face, facial expressions, and gestures, and then analyzing the data using libraries such as OpenCV. An emotion analysis engine is also integrated into this process to analyze the user's emotional state. The user's ID and emotional state are then used as the basis for the system to provide personalized responses.

[0711] Emotion recognition

[0712] The robot analyzes the user's emotions from the captured facial expressions and voice. It uses an emotion analysis engine to identify specific emotions (e.g., happiness, sadness, surprise), allowing the system to understand the user's current emotional state and prepare to generate a response accordingly.

[0713] Start a conversation

[0714] The robot initiates a conversation using an NLP model based on the user's profile and emotional state. The server queries a database to retrieve the user's profile. Using an NLP model (e.g., GPT-4), it generates an appropriate greeting for the user and speaks it through the speaker. The greeting is adjusted based on the user's emotional state.

[0715] Answering questions

[0716] When a user types a question into the robot, it is analyzed using a natural language processing model. Analysis includes extracting the intent of the question and relevant entities. The server then references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer, which the robot then speaks in a tone that best suits its emotional state.

[0717] Autonomous Mobility and Navigation

[0718] The robot autonomously navigates to a destination specified by the user. It uses an algorithm based on SLAM technology to calculate the optimal path, and moves along that path while avoiding obstacles. Along the way, it adjusts the tone and content of its guidance according to the user's emotional state.

[0719] Examples of concrete examples and prompts

[0720] Example 1: Store directions

[0721] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and analyzes them using image recognition and an emotion engine. It then analyzes the intent of the question using natural language processing, generates and speaks an appropriate response, and guides the user to the reception desk, providing detailed explanations if the user appears confused.

[0722] Example prompt for a generative AI model: "The user is asking where the reception desk is. The robot will politely provide directions. Please respond especially politely if the user seems confused."

[0723] Example 2: Bank balance inquiry

[0724] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves the account information from the database. If authentication is successful, the emotion engine analyzes the user's emotional state and responds with the current account balance. If the user is nervous, it adds a comment to relax them.

[0725] Example prompt for a generative AI model: "The user asks for their account balance. The robot responds by adding a relaxing comment based on the user's emotional state."

[0726] The above is a specific embodiment of the present invention, and the system allows users to experience natural conversations that take emotion into consideration.

[0727] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0728] Step 1: Initialize the system

[0729] The server initializes the entire system. First, it loads an image recognition model into memory using TensorFlow or PyTorch. Next, it loads a natural language processing (NLP) model such as OpenAI's GPT-4, and then initializes a multilingual translation model such as the Google Translate API and a sentiment analysis engine using the Affectiva SDK. It also sets up a database using MySQL or PostgreSQL to store user profiles and interaction history. This completes the initialization process, making the various models ready for processing. It receives model configuration files and database configuration information as input, and the model loading is complete as output.

[0730] Step 2: Getting started with user awareness

[0731] The robot recognizes the user through a camera. Specifically, it uses the camera to capture the user's face, facial expressions, and gestures, and analyzes the data using the OpenCV library. As a result of the analysis, the user's facial feature points are extracted, and the emotion analysis engine analyzes the user's emotional state. The input is the captured image data, and the output is the user's ID and emotional state.

[0732] Step 3: Detailed analysis of emotional state

[0733] The robot performs a detailed analysis of the user's emotions from the captured facial expressions and voice. Using the Affectiva SDK, it analyzes facial expression data to identify specific emotions (such as joy, sadness, or surprise). It also collects the user's voice data and analyzes its emotional nuances. It receives facial expression and voice data as input and outputs the identified emotional state.

[0734] Step 4: Start a conversation

[0735] The robot starts a conversation using an NLP model based on the user's profile and emotional state. The server executes an SQL query to retrieve the user profile from the database. Using the NLP model (GPT-4), it generates an appropriate greeting based on the user's profile and emotional state. It then speaks this to the user through the speaker. The input is the user profile and emotional state retrieved from the database, and the output is the generated greeting.

[0736] Step 5: Answer questions

[0737] The user inputs a question to the robot. The user uses a voice recognition system to input the question in text format. The server receives the text data and uses an NLP model to analyze the intent of the question and extract relevant entities. Based on this, it references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer. It then takes into account the results of the sentiment analysis engine to adjust the tone and content of the response and speaks the answer to the user through the speaker. The input is the user's question text, and the output is the generated answer.

[0738] Step 6: Autonomous Movement and Navigation

[0739] The robot moves autonomously to a destination specified by the user. First, the user inputs instructions for the destination. The robot uses SLAM technology to calculate the optimal path to the destination and moves along that path. During movement, it uses LiDAR and ultrasonic sensors to detect obstacles and avoid them to proceed safely. It also adjusts the tone and content of its guidance according to the user's emotional state. It receives destination instructions and sensor data as input, and outputs the movement results along the optimal path.

[0740] (Application example 2)

[0741] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0742] Today's elderly often have difficulty communicating with store staff and navigating the store when shopping or using services in physical stores. Furthermore, language barriers and a lack of emotional support reduce the elderly's satisfaction. This reduces the opportunities for elderly people to enjoy shopping in physical stores and causes stress, which is an issue.

[0743] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0744] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. This enables elderly people to shop or use services in brick-and-mortar stores by recognizing faces and facial expressions, analyzing questions, generating language-compatible information, acquiring profiles and dialogue histories, and analyzing emotional states, thereby providing appropriate responses and in-store navigation based on their emotions.

[0745] "Image recognition means" refers to technology that uses a device such as a camera to recognize a user's face, facial expressions, and gestures.

[0746] "Natural language processing means" is a technology that analyzes questions and requests entered by users and understands their intentions.

[0747] "Multilingual translation means" refers to technology for translating and generating information in accordance with the language used by the user.

[0748] "Database connection means" refers to technology for connecting to a database that stores user profiles and interaction history, and obtaining the necessary information.

[0749] The "emotion engine" is a technology that analyzes the user's emotional state from their facial expressions and tone of voice to determine what emotions the user is feeling.

[0750] A "system" is a collection of components that integrate multiple technical means to provide specific functions or services.

[0751] "In-store navigation" is a function that guides users to find desired locations and products within a store.

[0752] "Support for the elderly" means providing support and services to help the elderly live their daily lives more conveniently and safely.

[0753] "Multilingual support" means providing services and information in different languages ​​to users who speak different languages.

[0754] "Appropriate response" means providing answers or guidance appropriate to the situation based on the user's question or condition.

[0755] A "profile" is data that records basic information and personal characteristics about a user.

[0756] System configuration

[0757] This invention is realized by a system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. The components of this system are described in detail below.

[0758] Hardware and Software

[0759] Image Recognition: The system recognizes the user's face and facial expressions using a smartphone with a camera or a head-mounted display. It uses the OpenCV (cv2) library and a model trained in TensorFlow.

[0760] Natural Language Processing: Natural language processing uses the Transformers library pipeline to analyze question answers, using generative AI models such as BERT and RoBERTa.

[0761] Multilingual translation method: For multilingual translation, the DeepTranslator library is used, and natural translation is provided in conjunction with the Google Translate API.

[0762] Database connection method: User profiles and interaction history are managed using SQLite, and the necessary information is obtained.

[0763] Emotion engine: Emotion analysis uses the Transformers library pipeline to analyze emotions from voice and facial expressions.

[0764] System initialization

[0765] Initialization takes place on the server, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing user profiles and interaction history.

[0766] User recognition and sentiment analysis

[0767] The device's camera captures the user's face, facial expressions, and gestures, and analyzes them using image recognition.The emotion engine analyzes the user's emotional state and determines how they are feeling.

[0768] Initiating conversations and answering questions

[0769] The system analyzes questions using natural language processing and generates an appropriate greeting based on the user's profile. It also adjusts the tone and content of the response based on the user's emotional state, using an emotion engine. The system offers conversations that are especially designed for seniors, ensuring their peace of mind.

[0770] In-store navigation

[0771] Based on the user's gestures, the system guides the user to the desired location or product within the store. The system also reflects the user's emotional state during navigation, adjusting the navigation to ensure a safe and secure journey.

[0772] Examples and prompts

[0773] For example, if an elderly person asks, "Do you have the shirt I'm looking for in stock?", the smartphone camera will recognize the user's face and perform image recognition and emotion analysis. Natural language processing will then be used to obtain the appropriate stock information, which will be translated into multiple languages ​​and displayed to the user.

[0774] Prompt Sentence Examples

[0775] "Do you have the shirt I'm looking for in stock?"

[0776] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0777] Step 1:

[0778] The server initializes the entire system, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing the user profile and interaction history.

[0779] Input: Image recognition model, natural language processing model, multilingual translation model, emotion engine, database

[0780] Output: Initialized models and database connections

[0781] Step 2:

[0782] The device uses a camera to capture the user's face, facial expressions, and gestures, and the captured image data is analyzed using image recognition tools.

[0783] Input: Video data of the user's face, expressions, and gestures

[0784] Output: Parsed user ID and emotional state

[0785] Step 3:

[0786] Based on the analysis results, the device further analyzes the user's emotional state using an emotion engine, thereby specifically grasping the user's emotional state.

[0787] Input: User's video data and analysis results

[0788] Output: Detailed emotional state data

[0789] Step 4:

[0790] The server retrieves the user profile and interaction history from the database and generates an appropriate greeting for the user. It uses natural language processing means to provide a greeting based on the user's emotional state.

[0791] Input: User ID, emotional state data, user profile and interaction history from database

[0792] Output: Emotion-based greeting

[0793] Step 5:

[0794] The user inputs a question into the terminal. The server analyzes the question using natural language processing to extract the intent of the question and related information. It then generates an appropriate answer based on the analysis results.

[0795] Input: User question

[0796] Output: Analyzed question intent and answer suggestions

[0797] Step 6:

[0798] The server uses a multilingual translation means to translate the generated answer into the user's language, and the translated result is displayed to the user.

[0799] Input: Analyzed question intent and answer suggestions

[0800] Output: The answer translated into the user's language

[0801] Step 7:

[0802] The device monitors the user's gestures and provides in-store navigation as needed, adjusting the tone and content of the guidance based on the user's emotional state.

[0803] Input: User gesture data and emotional state

[0804] Output: Navigation prompts and tuned tones

[0805] Through this series of steps, the system can provide an environment where seniors can shop and use services in physical stores in a natural and safe manner.

[0806] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0807] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0809] [Third embodiment]

[0810] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0811] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0812] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0813] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0814] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0815] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0816] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0817] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0818] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0819] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0820] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0821] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0822] The present invention relates to a multimodal AI system using a humanoid robot. This system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. Specific embodiments of the system are described below.

[0823] System initialization

[0824] The server performs the initial setup of the system, including loading the image recognition model, natural language processing (NLP) model, and multilingual translation model, and setting up the database, so that the entire system is ready to run efficiently and smoothly.

[0825] User Awareness

[0826] The robot uses a camera to recognize the user. During this process, the camera captures the user's face and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[0827] Start a conversation

[0828] The robot uses the user information obtained by the recognition means to initiate a conversation using natural language processing means. First, it retrieves the user profile from the database and generates an appropriate greeting based on that information. This allows the user to begin a natural dialogue with the system.

[0829] Answering questions

[0830] When a user types a question into the robot, it analyzes it using natural language processing. During the analysis, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. For example, if a user asks about the weather, the system retrieves weather information and provides it to the user.

[0831] Autonomous Mobility and Navigation

[0832] The robot can move autonomously to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path. During this process, it is designed to reach the destination while avoiding obstacles and other people. For example, if the user says, "Please guide me to the reception desk," the robot will calculate the route to the reception desk and actually begin guiding the user.

[0833] Specific examples

[0834] Example 1: Store directions

[0835] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expression, and analyzes it using image recognition. The robot then analyzes the intent of the question using natural language processing, and the system generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and actually guides the user to the reception desk.

[0836] Example 2: Bank balance inquiry

[0837] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot responds, "My current account balance is 100,000 yen." This allows the user to obtain information quickly and accurately.

[0838] Through these functions, the system of the present invention aims to achieve natural interaction with users and highly efficient service provision, thereby reducing workload and improving service quality.

[0839] The processing flow will be explained below.

[0840] Step 1:

[0841] The server initializes the entire system. First, it loads the image recognition, natural language processing, and multilingual translation models. Next, it sets up the database used by the system and prepares it to store user profiles and interaction history.

[0842] Step 2:

[0843] The robot uses a camera to capture images to recognize the user's face and facial expressions, and the captured images are passed to an image recognition means to analyze the user's face, emotions, and gestures.

[0844] Step 3:

[0845] Based on the analysis results, the robot retrieves the user's profile from the database, allowing it to refer to each user's past interaction history and specific information.

[0846] Step 4:

[0847] The robot uses natural language processing tools to initiate a conversation with the user, generating appropriate greetings and introductory phrases based on the user's profile, emotions, and gestures, and then delivering them to the user via voice or text.

[0848] Step 5:

[0849] When a user types a question into the robot, the question is analyzed using natural language processing tools, which extracts the intent of the question and related entities, thereby providing a clear understanding of the question.

[0850] Step 6:

[0851] Based on the analysis results, the robot retrieves the appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. The response is then provided to the user.

[0852] Step 7:

[0853] When a user requests guidance to a specific location, the robot begins navigation to the destination. First, it calculates the optimal path to the destination and then autonomously moves along that path, avoiding obstacles and other people along the way to safely reach the destination.

[0854] Step 8:

[0855] The robot records the results of its interactions with the user and their actions in a database, which can be used for future interactions. This record includes the interaction history, the user's reactions, and any new information acquired. This allows the system to continuously learn and improve the quality of its services.

[0856] In this way, the system of the present invention realizes natural interaction with the user and provides a variety of services.

[0857] Example 1

[0858] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0859] Conventional humanoid robot systems have had problems with natural interaction with users and efficient service provision. Specifically, they suffer from low recognition accuracy, delayed responses, and limited ability to support different languages. Furthermore, there is still room for improvement in terms of recognizing users' faces and gestures, identifying emotions, providing appropriate information, and navigation.

[0860] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0861] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and initial setting means. This allows the server to recognize the user's face and facial expressions, identify their ID and emotions, and provide appropriate information in multiple languages, thereby realizing natural interaction with the user and enabling efficient service provision.

[0862] "Image recognition means" refers to means that has the function of analyzing image data acquired using an imaging device such as a camera and recognizing a specific object (for example, a user's face or gesture).

[0863] "Natural language processing means" refers to means that include algorithms and models for analyzing text data entered by a user and understanding their intent and emotions.

[0864] A "multilingual translation means" is a means that has the function of translating text between different languages, and that understands input from a user in different languages ​​and generates corresponding information.

[0865] "Database connection means" refers to a means that has the function of communicating with a database server and acquiring and saving the necessary information.

[0866] The "means for performing initial settings" refers to a means for performing basic settings for system operation at the start, including loading various models and setting up a database.

[0867] "Means for recognizing a user's face and facial expressions" refers to means that has the function of extracting specific features from a user's facial image captured by a camera, identifying the user based on the results, and determining their emotional state.

[0868] "Means for identifying a user's ID and emotions" refers to a means for analyzing data obtained through image recognition to determine who the user is (ID) and their emotional state at the time.

[0869] "Means for obtaining a user profile" means means capable of obtaining information related to a user from a database, including past conversation history and preferences.

[0870] The "means for generating a greeting message" is a means having a function for automatically generating a greeting message suitable for a user based on the acquired user profile.

[0871] "Means for analyzing the intent of a question and related entities" refers to a means for analyzing the question entered by the user and identifying the intent of the question and related keywords (entities).

[0872] The "means for generating an appropriate answer to a question" is a means for obtaining relevant information based on the intent and entities of the analyzed question, and automatically generating an appropriate answer based on that information.

[0873] The "means for continuously capturing image data" refers to a means having a function for activating a camera and continuously acquiring image data at regular intervals.

[0874] The "means for preprocessing image data" refers to a means having a function for performing preprocessing such as resizing, noise removal, and color conversion on the acquired image data.

[0875] "Means for processing sensor information in real time and moving autonomously" refers to means that have the ability to analyze information from various sensors in real time, calculate a route according to the environment, and move autonomously.

[0876] The present invention relates to a multimodal AI system using a humanoid robot. The system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. It also includes means for recognizing a user's face and facial expressions, identifying the user's ID and emotion, acquiring a user profile, generating a greeting, analyzing the intent of a question and related entities, generating appropriate answers to the question, continuously capturing image data, preprocessing the image data, and processing sensor information in real time to move autonomously.

[0877] 1. System initialization

[0878] The server performs the initial system configuration. First, for image recognition, it loads a ResNet50 model using TensorFlow using an NVIDIA GPU. Next, for natural language processing, it loads the BERT model using Hugging Face's Transformers library. It also configures the Google Translate API for multilingual translation, and starts a MySQL database to create the necessary tables and indexes.

[0879] 2. User Awareness

[0880] The robot captures the user's face and gestures using a built-in camera (e.g., a standard HD camera). The image data is preprocessed with OpenCV and then analyzed using a ResNet50 model running on NVIDIA GPUs, which identifies the user's identity, face, and facial expressions.

[0881] 3. Start a conversation

[0882] The robot starts a conversation using the data mentioned above. First, it uses the user's ID to retrieve the user profile from a MySQL database. Then, it uses that information to generate an appropriate greeting using Hugging Face's BERT model. For example, it might generate a greeting like, "Hello, Tanaka-san. How's your day?"

[0883] 4. Answering questions

[0884] Users input questions into the robot, which then uses voice input (e.g., a standard microphone) to receive the question and convert it into text using the Google Speech-to-Text API. It then uses Hugging Face's BERT model to analyze the intent of the question and relevant entities, and based on the analysis results, queries the appropriate API (e.g., the OpenWeatherMap API) or database to generate an answer to the question.

[0885] 5. Autonomous Movement and Navigation

[0886] The robot autonomously moves to a location specified by the user. First, the user specifies the destination via voice input or touchscreen. Next, the robot uses SLAM technology to generate a map of the environment and calculates the optimal route using algorithms such as the A algorithm. The robot autonomously moves along the calculated route using a LiDAR sensor and motor control unit, avoiding obstacles and reaching the destination.

[0887] Specific examples

[0888] In-store directions

[0889] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and first identifies the user using image recognition. It then analyzes the intent of the question using natural language processing and generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and autonomously guides the user to the reception desk using SLAM technology and a motor control unit.

[0890] Bank balance inquiry

[0891] When a user asks, "What is my account balance?", the robot first authenticates the user and retrieves account information from the database. After successful authentication, the robot responds, "Your current account balance is 100,000 yen."

[0892] Prompt Sentence Examples

[0893] "When a user asks for weather information, how do you get the information from the API and respond?"

[0894] "Please explain in detail the algorithm that recognizes user emotions."

[0895] As explained above, the system of the present invention realizes natural interaction with users and aims to reduce the workload and improve service quality through highly efficient service provision.

[0896] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0897] Step 1: Initialize the system

[0898] The server performs the initial system setup, inputting the image recognition model, natural language processing model, multilingual translation model, and database configuration information.

[0899] 1. As a deep learning model for image recognition, we load a ResNet50 model in TensorFlow using an NVIDIA GPU. The input is the TensorFlow model file, and the output is the model loaded in memory.

[0900] 2. Load the BERT model from Hugging Face's Transformers library as the natural language processing model. The input is the model file, and the output is the BERT model loaded in memory.

[0901] 3. Set up the Google Translate API as a multilingual translation model. The input is the API key and configuration information, and the output is the API availability status.

[0902] 4. Starts a MySQL database and creates the necessary tables and indexes. The input is database configuration information, and the output is a configured database ready to use.

[0903] Step 2: User Awareness

[0904] The robot uses a camera to recognize the user, and the input is image data obtained from the camera.

[0905] 1. Start a camera (e.g., HD camera) and continuously capture image data. The input is the camera sensor data, and the output is the captured image data.

[0906] 2. The acquired image data is preprocessed using the OpenCV library, which includes resizing, noise removal, color conversion, etc. The input is raw image data, and the output is preprocessed image data.

[0907] 3. The preprocessed image data is input to a ResNet50 model running on an NVIDIA GPU to identify the user's ID and emotion. The input is the preprocessed image data, and the output is the identified user's ID and emotion information.

[0908] Step 3: Start a conversation

[0909] The robot starts a conversation based on the recognized user information, including the user's ID and emotional information.

[0910] 1. The robot uses the user's ID to retrieve the user profile from the MySQL database. The input is the user's ID and the output is the retrieved user profile.

[0911] 2. Based on the retrieved user profile, the robot uses Hugging Face's BERT model to generate an appropriate greeting, such as "Hello, Tanaka-san. How are you today?" The input is the user profile, and the output is the generated greeting.

[0912] Step 4: Answer questions

[0913] The user inputs a question to the robot, and the input is voice data.

[0914] 1. Collect user questions using a voice input function (e.g., microphone). The input is voice data, and the output is the question data captured as voice.

[0915] 2. The robot converts voice data into text data using the Google Speech-to-Text API. The input is voice data and the output is text data.

[0916] 3. The text data is analyzed using Hugging Face's BERT model to extract the question intent and related entities. The input is the text data, and the output is the analysis results.

[0917] 4. Based on the analysis results, the robot queries the appropriate API (e.g., OpenWeatherMap API) or database to generate an answer to the question. The input is the analysis results and the necessary API information, and the output is the generated answer.

[0918] Step 5: Autonomous Movement and Navigation

[0919] The robot moves autonomously to a specified location, and the input is a destination specified by the user.

[0920] 1. The user specifies a destination using voice input or a touch screen. The input is the destination information, and the output is the specified destination.

[0921] 2. The robot uses SLAM technology to generate a map of the environment and calculate the optimal path. The input is sensor information and destination information, and the output is the calculated path information.

[0922] 3. The robot uses a LiDAR sensor and a motor control unit to move autonomously based on a calculated path. The input is path information and real-time sensor information, and the output is autonomous movement behavior.

[0923] (Application example 1)

[0924] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0925] Modern stores and public facilities require a means for visitors to quickly and accurately find their destination and obtain information. However, conventional guidance systems and information centers face problems such as labor shortages and language barriers. Furthermore, visitors often have to search for the information they need themselves, which is inconvenient. A system that solves these issues and allows visitors to easily reach their destination is needed.

[0926] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0927] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, and database connection means. This enables the server to recognize a user's face and facial expressions, analyze the user's questions, generate information corresponding to the user's language, and acquire a user profile and dialogue history. The server also includes connection means with a smart device, means for activating and controlling the robot via the smart device, autonomous robot movement means, and means for guiding the robot to a user-specified destination using the autonomous movement means. This allows visitors to easily call the robot from their smartphones, and the robot automatically guides the visitor to their destination, significantly improving visitor convenience.

[0928] "Image recognition means" refers to a device or system that uses a camera or sensor to recognize a user's face, facial expressions, and gestures.

[0929] "Natural language processing means" is a technology that analyzes questions and instructions uttered by users and understands their intent and content.

[0930] "Multilingual translation means" refers to technology that translates between different languages ​​and generates information corresponding to the language used by the user.

[0931] "Database connection means" refers to technology for accessing a database that stores user profiles, interaction history, and other necessary data, and acquiring that information.

[0932] "Means for connecting with smart devices" refers to technology that connects devices such as smartphones and tablets to the system, enabling two-way communication.

[0933] "Means for activating and controlling robots" refers to technology for activating robots via smart devices and controlling their movements and functions.

[0934] "Autonomous robot mobility means" refers to technology that enables a robot to move autonomously while recognizing its surrounding environment.

[0935] "Generating information" is the process of generating appropriate answers or guidance information based on the user's questions or instructions.

[0936] A "user profile" is data that compiles information about a user (e.g., the user's ID, past interaction history, preferences, etc.).

[0937] "Dialogue history" is data that records past conversations and interactions with a user.

[0938] "Autonomous movement" refers to a robot's ability to move while adapting to its surrounding environment based on its own judgment.

[0939] The present invention relates to a system that controls a robot in cooperation with a smart device to provide guidance and information to users in stores and public facilities. This system is realized using the following various means.

[0940] 1. Image Recognition Methods:

[0941] The server uses cameras and sensors to recognize the user's face, facial expressions, and gestures, using software libraries such as OpenCV and TensorFlow, which are known for their image processing technology.

[0942] 2. Natural Language Processing Tools:

[0943] The server uses natural language processing techniques to analyze the user's questions and instructions. This process uses various NLP (Natural Language Processing) libraries and tools (e.g., Google's Dialogflow, OpenAI's GPT-3).

[0944] 3. Multilingual translation tools:

[0945] The server uses a multilingual translation engine to translate between different languages ​​and generate information corresponding to the user's language, for example, the Google Translate API.

[0946] 4. Database connection method:

[0947] The server retrieves information such as the user's profile and interaction history from a database, using Firebase or AWS RDS to dynamically manage user information.

[0948] 5. Connectivity with smart devices:

[0949] The smartphone application communicates with the robot via Bluetooth or Wi-Fi to activate and control it, allowing users to summon the robot using their smartphone.

[0950] 6. Autonomous robotic mobility:

[0951] The robot autonomously navigates to its destination using SLAM (Simultaneous Localization and Mapping) technology with Lidar sensors and cameras, using ROS Lidar and Google Cartographer.

[0952] 7. Information generation:

[0953] The server generates appropriate answers and guidance information in response to user questions and instructions, using a generative AI model in the process.

[0954] 8. User Awareness and Guidance:

[0955] The robot recognizes the user with a camera, analyzes the intent of the question using natural language processing technology, and provides guidance. If the user asks, "Where is the dairy section?", the robot will use the store's map information to answer, "The dairy section is on the right side of this floor. I'll show you." and begin guiding the user. Using its autonomous mobility, the robot will accurately guide the user to their destination.

[0956] 9. Example prompt:

[0957] An example of a prompt sentence to input to the generative AI model is as follows:

[0958] "A user is in a store asking about the dairy section. The robot recognizes the user with a camera and analyzes the intent of the question using NLP technology. It then retrieves store map information from a database and provides directions. Please explain the specific processing flow in detail."

[0959] In this way, the system of the present invention can provide an environment in which the user can quickly and accurately obtain information regardless of the environment.

[0960] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0961] Step 1:

[0962] The user starts the smartphone app and presses the "Start Guide Robot" button, at which point the smartphone uses its camera and microphone to check the surrounding environment.

[0963] input:

[0964] User presses a button

[0965] output:

[0966] Camera and microphone ambient data

[0967] Specific behavior:

[0968] The camera captures image data of the surroundings, and the microphone captures audio data.

[0969] Step 2:

[0970] The device calls the robot via Bluetooth or Wi-Fi and sends commands for the robot to move to the user's location.

[0971] input:

[0972] Camera and microphone data

[0973] output:

[0974] Robot commands

[0975] Specific behavior:

[0976] It connects to the robot communication module via Bluetooth or Wi-Fi and sends the command "move to the user."

[0977] Step 3:

[0978] The robot recognizes the user using a camera and analyzes the user's face and gestures using image recognition means.

[0979] input:

[0980] Image data acquired by the camera

[0981] output:

[0982] User ID, facial expression, and gesture information

[0983] Specific behavior:

[0984] Image processing libraries such as OpenCV and TensorFlow are used to recognize and analyze the user's face and gestures and obtain user information.

[0985] Step 4:

[0986] The server uses natural language processing means to analyze the user's question and understand its intent.

[0987] input:

[0988] User utterances

[0989] output:

[0990] User questions and their intent

[0991] Specific behavior:

[0992] NLP tools such as Google's Dialogflow and OpenAI's GPT-3 are used to interpret the content of the speech and extract the intent of the question.

[0993] Step 5:

[0994] The server uses database connectivity to obtain the user's profile and interaction history and generates appropriate information.

[0995] input:

[0996] User ID, question content

[0997] output:

[0998] User profile and corresponding answer information

[0999] Specific behavior:

[1000] It connects to Firebase or AWS RDS, obtains the user's past interaction history and profile information, and then uses a generative AI model to generate appropriate answers and guidance information.

[1001] Step 6:

[1002] The server uses a multilingual translation means to translate the generated information into the language used by the user.

[1003] input:

[1004] Generated answer information

[1005] output:

[1006] Information translated into the user's language

[1007] Specific behavior:

[1008] Use the Google Translate API to translate the generated information into the user's language.

[1009] Step 7:

[1010] The robot presents appropriate information to the user via voice or text and begins providing guidance in response to the user's questions.

[1011] input:

[1012] Translated information

[1013] output:

[1014] User guidance information

[1015] Specific behavior:

[1016] The text information is converted into speech using a speech synthesis engine (e.g., Google TTS) and provided to the user. The autonomous mobility system is also used to calculate the route to the user's specified destination and begin providing guidance.

[1017] Step 8:

[1018] The robot uses autonomous mobility to guide the user to their destination while avoiding obstacles.

[1019] input:

[1020] User-specified destination information and surrounding environment information

[1021] output:

[1022] Guidance to the destination, reaching the final point

[1023] Specific behavior:

[1024] It uses Lidar sensors and SLAM technology (e.g., ROS Lidar, Google Cartographer) to map the surrounding environment, calculates routes and avoids obstacles in real time, and guides the user to a specified destination.

[1025] As described above, by acquiring and analyzing the necessary input data at each step and generating appropriate output, the system of the present invention can smoothly answer the user's questions and provide guidance. The system of the present invention provides visitors with fast and accurate guidance.

[1026] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1027] The present invention relates to a multimodal AI system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine to achieve natural interaction with users. Specific embodiments of the system are described below.

[1028] System initialization

[1029] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[1030] User Awareness

[1031] The robot uses a camera to recognize the user. During this process, the camera captures the user's face, facial expressions, and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[1032] Emotion recognition

[1033] The robot uses an emotion engine to analyze the user's emotions from the captured facial expressions and voice. The emotion engine analyzes the data from each facial expression to determine whether the user is happy, confused, tired, etc.

[1034] Start a conversation

[1035] The robot uses the user information obtained by the recognition means and emotion engine to initiate a conversation using natural language processing means. First, it retrieves the user profile from a database and generates an appropriate greeting based on that information. This greeting is adjusted taking into account the user's emotional state. This allows the user to begin a natural dialogue with the system.

[1036] Answering questions

[1037] When a user inputs a question into the robot, the question is analyzed using natural language processing. During the analysis process, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. Furthermore, by incorporating the analysis results of the emotion engine into the response, the robot responds with an appropriate tone and content according to the user's emotional state.

[1038] Autonomous Mobility and Navigation

[1039] The robot can autonomously navigate to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path, safely reaching the destination while avoiding obstacles and other people along the way. It can also adjust the tone and content of its guidance depending on the user's emotional state during the journey.

[1040] Specific examples

[1041] Example 1: Store directions

[1042] If a user asks, "Where is the reception desk?", the robot will capture the user's face and facial expressions with a camera and analyze them using image recognition and an emotion engine. It will then analyze the intent of the question using natural language processing, and the system will generate an appropriate response. The robot will respond, "The reception desk is on the left side of this floor. I'll show you," and actually guide the user to the reception desk. If the robot determines that the user appears confused, it can add a particularly detailed explanation.

[1043] Example 2: Bank balance inquiry

[1044] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and responds, "Your current account balance is 100,000 yen." If the user is nervous, the robot can also add additional comments to help them relax.

[1045] Through these functions, the system of the present invention realizes natural interaction with users and provides a variety of services. Utilizing the emotion engine enables advanced responses that take user emotions into consideration, significantly improving the quality of services.

[1046] The processing flow will be explained below.

[1047] Step 1:

[1048] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[1049] Step 2:

[1050] The robot utilizes a camera to capture the user's face, facial expressions, and gestures, which are then analyzed using image recognition to obtain the user's identity, emotions, and gestures.

[1051] Step 3:

[1052] The robot uses an emotion engine to analyze the user's emotional state, extracting emotional data from facial expressions, voice tone, and gestures to determine whether the user is happy, confused, tired, etc.

[1053] Step 4:

[1054] The robot uses natural language processing to initiate a conversation based on the acquired user information. First, it retrieves the user profile from a database and generates a greeting based on that information. The greeting is adjusted to reflect the user's emotional state.

[1055] Step 5:

[1056] When a user types a question into the robot, the question is analyzed using natural language processing, which extracts the intent of the question and related entities to clearly understand the question.

[1057] Step 6:

[1058] Based on the analysis results, the robot retrieves appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. At this time, the robot will take into account the user's emotional state obtained from the emotion engine and adjust the tone and content of the response.

[1059] Step 7:

[1060] When a user requests guidance to a specified location, the robot begins navigation. First, it calculates the optimal path to the destination and autonomously moves along that path. During the journey, it navigates safely while avoiding obstacles and other people. It also monitors the user's emotional state in real time and provides guidance accordingly.

[1061] Step 8:

[1062] The robot records the results of its interactions with the user in a database. Dialogue history and emotional data are accumulated and used for future interactions. This allows the system to continuously learn and improve the quality of its services.

[1063] Specific examples

[1064] Example 1: Store directions

[1065] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expressions, and analyzes them using image recognition and an emotion engine. Using the user's emotional state and profile information obtained based on the analysis, the robot generates a response using natural language processing. The robot responds in a friendly tone, saying, "The reception desk is on the left side of this floor. I'll show you there," and begins navigation.

[1066] Example 2: Bank balance inquiry

[1067] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from a database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and generate a response accordingly. It responds in a calm tone, saying, "Your current account balance is 100,000 yen," and can also add additional comments to ease the tension during the question.

[1068] This embodiment allows users to receive flexible responses that take their emotions into consideration, improving the quality of service.

[1069] Example 2

[1070] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1071] Conventional dialogue systems have limited interaction with users and lack the ability to respond with emotion or to handle diverse languages. Furthermore, their user recognition and autonomous movement capabilities are limited, making it difficult to achieve natural and effective dialogue. This results in a poor user experience and reduced system utilization efficiency.

[1072] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an image recognition means, a natural language processing means, a multilingual translation means, a database connection means, an emotion analysis means, a user recognition means, and an autonomous movement means. This makes it possible to recognize the user's face and facial expression, analyze questions, generate information in multiple languages ​​while taking into account the emotional state, acquire a user profile and dialogue history, and autonomously move to a specified destination.

[1073] "Image recognition means" refers to a means of acquiring image data such as a user's face, facial expressions, and gestures using a camera or sensor, and analyzing that data.

[1074] A "natural language processing means" is a means for analyzing text or voice questions entered by a user, understanding their intent, and generating an appropriate response.

[1075] The "multilingual translation means" is a means for translating text input in different languages ​​and generating information corresponding to multiple languages.

[1076] "Database connection means" refers to a means for accessing a database to obtain and store related data such as user profiles and interaction history.

[1077] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data to identify their emotional state.

[1078] "User recognition means" refers to a means for identifying a user's ID and distinguishing between individual users.

[1079] An "autonomous vehicle" is a vehicle that moves autonomously to a specified destination while avoiding obstacles.

[1080] The present invention relates to a system that integrates image recognition means, natural language processing means, multilingual translation means, database connection means, emotion analysis means, user recognition means, and autonomous mobility means to realize natural interactions with users. Detailed embodiments of the system are described below.

[1081] System configuration

[1082] Hardware

[1083] This system uses the following hardware components:

[1084] Camera (e.g. high-resolution webcam)

[1085] microphone

[1086] speaker

[1087] Robot body (including motors and sensors for autonomous movement)

[1088] server

[1089] Database (e.g. MySQL, PostgreSQL)

[1090] software

[1091] The system consists of the following software components:

[1092] Image recognition models (e.g., TensorFlow, PyTorch)

[1093] Natural language processing models (e.g., GPT-4)

[1094] Multilingual translation models (e.g., Google Translate API)

[1095] Sentiment analysis engine (e.g. Affectiva SDK)

[1096] Database management systems (e.g., MySQL, PostgreSQL)

[1097] Autonomous movement algorithms (e.g., SLAM technology)

[1098] System Operation

[1099] System initialization

[1100] The server first initializes the entire system. This initialization includes loading image recognition models, natural language processing models, multilingual translation models, and a sentiment analysis engine, as well as setting up a database. Specifically, the server loads trained models into memory using TensorFlow or PyTorch, and initializes the sentiment analysis engine using the Affectiva SDK. The database uses MySQL or PostgreSQL, and is prepared to store user profiles and interaction history.

[1101] User Awareness

[1102] The robot recognizes the user through a camera. This involves capturing the user's face, facial expressions, and gestures, and then analyzing the data using libraries such as OpenCV. An emotion analysis engine is also integrated into this process to analyze the user's emotional state. The user's ID and emotional state are then used as the basis for the system to provide personalized responses.

[1103] Emotion recognition

[1104] The robot analyzes the user's emotions from the captured facial expressions and voice. It uses an emotion analysis engine to identify specific emotions (e.g., happiness, sadness, surprise), allowing the system to understand the user's current emotional state and prepare to generate a response accordingly.

[1105] Start a conversation

[1106] The robot initiates a conversation using an NLP model based on the user's profile and emotional state. The server queries a database to retrieve the user's profile. Using an NLP model (e.g., GPT-4), it generates an appropriate greeting for the user and speaks it through the speaker. The greeting is adjusted based on the user's emotional state.

[1107] Answering questions

[1108] When a user types a question into the robot, it is analyzed using a natural language processing model. Analysis includes extracting the intent of the question and relevant entities. The server then references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer, which the robot then speaks in a tone that best suits its emotional state.

[1109] Autonomous Mobility and Navigation

[1110] The robot autonomously navigates to a destination specified by the user. It uses an algorithm based on SLAM technology to calculate the optimal path, and moves along that path while avoiding obstacles. Along the way, it adjusts the tone and content of its guidance according to the user's emotional state.

[1111] Examples of concrete examples and prompts

[1112] Example 1: Store directions

[1113] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and analyzes them using image recognition and an emotion engine. It then analyzes the intent of the question using natural language processing, generates and speaks an appropriate response, and guides the user to the reception desk, providing detailed explanations if the user appears confused.

[1114] Example prompt for a generative AI model: "The user is asking where the reception desk is. The robot will politely provide directions. Please respond especially politely if the user seems confused."

[1115] Example 2: Bank balance inquiry

[1116] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves the account information from the database. If authentication is successful, the emotion engine analyzes the user's emotional state and responds with the current account balance. If the user is nervous, it adds a comment to relax them.

[1117] Example prompt for a generative AI model: "The user asks for their account balance. The robot responds by adding a relaxing comment based on the user's emotional state."

[1118] The above is a specific embodiment of the present invention, and the system allows users to experience natural conversations that take emotion into consideration.

[1119] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1120] Step 1: Initialize the system

[1121] The server initializes the entire system. First, it loads an image recognition model into memory using TensorFlow or PyTorch. Next, it loads a natural language processing (NLP) model such as OpenAI's GPT-4, and then initializes a multilingual translation model such as the Google Translate API and a sentiment analysis engine using the Affectiva SDK. It also sets up a database using MySQL or PostgreSQL to store user profiles and interaction history. This completes the initialization process, making the various models ready for processing. It receives model configuration files and database configuration information as input, and the model loading is complete as output.

[1122] Step 2: Getting started with user awareness

[1123] The robot recognizes the user through a camera. Specifically, it uses the camera to capture the user's face, facial expressions, and gestures, and analyzes the data using the OpenCV library. As a result of the analysis, the user's facial feature points are extracted, and the emotion analysis engine analyzes the user's emotional state. The input is the captured image data, and the output is the user's ID and emotional state.

[1124] Step 3: Detailed analysis of emotional state

[1125] The robot performs a detailed analysis of the user's emotions from the captured facial expressions and voice. Using the Affectiva SDK, it analyzes facial expression data to identify specific emotions (such as joy, sadness, or surprise). It also collects the user's voice data and analyzes its emotional nuances. It receives facial expression and voice data as input and outputs the identified emotional state.

[1126] Step 4: Start a conversation

[1127] The robot starts a conversation using an NLP model based on the user's profile and emotional state. The server executes an SQL query to retrieve the user profile from the database. Using the NLP model (GPT-4), it generates an appropriate greeting based on the user's profile and emotional state. It then speaks this to the user through the speaker. The input is the user profile and emotional state retrieved from the database, and the output is the generated greeting.

[1128] Step 5: Answer questions

[1129] The user inputs a question to the robot. The user uses a voice recognition system to input the question in text format. The server receives the text data and uses an NLP model to analyze the intent of the question and extract relevant entities. Based on this, it references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer. It then takes into account the results of the sentiment analysis engine to adjust the tone and content of the response and speaks the answer to the user through the speaker. The input is the user's question text, and the output is the generated answer.

[1130] Step 6: Autonomous Movement and Navigation

[1131] The robot moves autonomously to a destination specified by the user. First, the user inputs instructions for the destination. The robot uses SLAM technology to calculate the optimal path to the destination and moves along that path. During movement, it uses LiDAR and ultrasonic sensors to detect obstacles and avoid them to proceed safely. It also adjusts the tone and content of its guidance according to the user's emotional state. It receives destination instructions and sensor data as input, and outputs the movement results along the optimal path.

[1132] (Application example 2)

[1133] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1134] Today's elderly often have difficulty communicating with store staff and navigating the store when shopping or using services in physical stores. Furthermore, language barriers and a lack of emotional support reduce the elderly's satisfaction. This reduces the opportunities for elderly people to enjoy shopping in physical stores and causes stress, which is an issue.

[1135] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1136] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. This enables elderly people to shop or use services in brick-and-mortar stores by recognizing faces and facial expressions, analyzing questions, generating language-compatible information, acquiring profiles and dialogue histories, and analyzing emotional states, thereby providing appropriate responses and in-store navigation based on their emotions.

[1137] "Image recognition means" refers to technology that uses a device such as a camera to recognize a user's face, facial expressions, and gestures.

[1138] "Natural language processing means" is a technology that analyzes questions and requests entered by users and understands their intentions.

[1139] "Multilingual translation means" refers to technology for translating and generating information in accordance with the language used by the user.

[1140] "Database connection means" refers to technology for connecting to a database that stores user profiles and interaction history, and obtaining the necessary information.

[1141] The "emotion engine" is a technology that analyzes the user's emotional state from their facial expressions and tone of voice to determine what emotions the user is feeling.

[1142] A "system" is a collection of components that integrate multiple technical means to provide specific functions or services.

[1143] "In-store navigation" is a function that guides users to find desired locations and products within a store.

[1144] "Support for the elderly" means providing support and services to help the elderly live their daily lives more conveniently and safely.

[1145] "Multilingual support" means providing services and information in different languages ​​to users who speak different languages.

[1146] "Appropriate response" means providing answers or guidance appropriate to the situation based on the user's question or condition.

[1147] A "profile" is data that records basic information and personal characteristics about a user.

[1148] System configuration

[1149] This invention is realized by a system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. The components of this system are described in detail below.

[1150] Hardware and Software

[1151] Image Recognition: The system recognizes the user's face and facial expressions using a smartphone with a camera or a head-mounted display. It uses the OpenCV (cv2) library and a model trained in TensorFlow.

[1152] Natural Language Processing: Natural language processing uses the Transformers library pipeline to analyze question answers, using generative AI models such as BERT and RoBERTa.

[1153] Multilingual translation method: For multilingual translation, the DeepTranslator library is used, and natural translation is provided in conjunction with the Google Translate API.

[1154] Database connection method: User profiles and interaction history are managed using SQLite, and the necessary information is obtained.

[1155] Emotion engine: Emotion analysis uses the Transformers library pipeline to analyze emotions from voice and facial expressions.

[1156] System initialization

[1157] Initialization takes place on the server, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing user profiles and interaction history.

[1158] User recognition and sentiment analysis

[1159] The device's camera captures the user's face, facial expressions, and gestures, and analyzes them using image recognition.The emotion engine analyzes the user's emotional state and determines how they are feeling.

[1160] Initiating conversations and answering questions

[1161] The system analyzes questions using natural language processing and generates an appropriate greeting based on the user's profile. It also adjusts the tone and content of the response based on the user's emotional state, using an emotion engine. The system offers conversations that are especially designed for seniors, ensuring their peace of mind.

[1162] In-store navigation

[1163] Based on the user's gestures, the system guides the user to the desired location or product within the store. The system also reflects the user's emotional state during navigation, adjusting the navigation to ensure a safe and secure journey.

[1164] Examples and prompts

[1165] For example, if an elderly person asks, "Do you have the shirt I'm looking for in stock?", the smartphone camera will recognize the user's face and perform image recognition and emotion analysis. Natural language processing will then be used to obtain the appropriate stock information, which will be translated into multiple languages ​​and displayed to the user.

[1166] Prompt Sentence Examples

[1167] "Do you have the shirt I'm looking for in stock?"

[1168] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1169] Step 1:

[1170] The server initializes the entire system, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing the user profile and interaction history.

[1171] Input: Image recognition model, natural language processing model, multilingual translation model, emotion engine, database

[1172] Output: Initialized models and database connections

[1173] Step 2:

[1174] The device uses a camera to capture the user's face, facial expressions, and gestures, and the captured image data is analyzed using image recognition tools.

[1175] Input: Video data of the user's face, expressions, and gestures

[1176] Output: Parsed user ID and emotional state

[1177] Step 3:

[1178] Based on the analysis results, the device further analyzes the user's emotional state using an emotion engine, thereby specifically grasping the user's emotional state.

[1179] Input: User's video data and analysis results

[1180] Output: Detailed emotional state data

[1181] Step 4:

[1182] The server retrieves the user profile and interaction history from the database and generates an appropriate greeting for the user. It uses natural language processing means to provide a greeting based on the user's emotional state.

[1183] Input: User ID, emotional state data, user profile and interaction history from database

[1184] Output: Emotion-based greeting

[1185] Step 5:

[1186] The user inputs a question into the terminal. The server analyzes the question using natural language processing to extract the intent of the question and related information. It then generates an appropriate answer based on the analysis results.

[1187] Input: User question

[1188] Output: Analyzed question intent and answer suggestions

[1189] Step 6:

[1190] The server uses a multilingual translation means to translate the generated answer into the user's language, and the translated result is displayed to the user.

[1191] Input: Analyzed question intent and answer suggestions

[1192] Output: The answer translated into the user's language

[1193] Step 7:

[1194] The device monitors the user's gestures and provides in-store navigation as needed, adjusting the tone and content of the guidance based on the user's emotional state.

[1195] Input: User gesture data and emotional state

[1196] Output: Navigation prompts and tuned tones

[1197] Through this series of steps, the system can provide an environment where seniors can shop and use services in physical stores in a natural and safe manner.

[1198] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1199] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1200] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1201] [Fourth embodiment]

[1202] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1203] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1204] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1205] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1206] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1207] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1208] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1209] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1210] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1211] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1212] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1213] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1214] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1215] The present invention relates to a multimodal AI system using a humanoid robot. This system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. Specific embodiments of the system are described below.

[1216] System initialization

[1217] The server performs the initial setup of the system, including loading the image recognition model, natural language processing (NLP) model, and multilingual translation model, and setting up the database, so that the entire system is ready to run efficiently and smoothly.

[1218] User Awareness

[1219] The robot uses a camera to recognize the user. During this process, the camera captures the user's face and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[1220] Start a conversation

[1221] The robot uses the user information obtained by the recognition means to initiate a conversation using natural language processing means. First, it retrieves the user profile from the database and generates an appropriate greeting based on that information. This allows the user to begin a natural dialogue with the system.

[1222] Answering questions

[1223] When a user types a question into the robot, it analyzes it using natural language processing. During the analysis, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. For example, if a user asks about the weather, the system retrieves weather information and provides it to the user.

[1224] Autonomous Mobility and Navigation

[1225] The robot can move autonomously to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path. During this process, it is designed to reach the destination while avoiding obstacles and other people. For example, if the user says, "Please guide me to the reception desk," the robot will calculate the route to the reception desk and actually begin guiding the user.

[1226] Specific examples

[1227] Example 1: Store directions

[1228] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expression, and analyzes it using image recognition. The robot then analyzes the intent of the question using natural language processing, and the system generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and actually guides the user to the reception desk.

[1229] Example 2: Bank balance inquiry

[1230] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot responds, "My current account balance is 100,000 yen." This allows the user to obtain information quickly and accurately.

[1231] Through these functions, the system of the present invention aims to achieve natural interaction with users and highly efficient service provision, thereby reducing workload and improving service quality.

[1232] The processing flow will be explained below.

[1233] Step 1:

[1234] The server initializes the entire system. First, it loads the image recognition, natural language processing, and multilingual translation models. Next, it sets up the database used by the system and prepares it to store user profiles and interaction history.

[1235] Step 2:

[1236] The robot uses a camera to capture images to recognize the user's face and facial expressions, and the captured images are passed to an image recognition means to analyze the user's face, emotions, and gestures.

[1237] Step 3:

[1238] Based on the analysis results, the robot retrieves the user's profile from the database, allowing it to refer to each user's past interaction history and specific information.

[1239] Step 4:

[1240] The robot uses natural language processing tools to initiate a conversation with the user, generating appropriate greetings and introductory phrases based on the user's profile, emotions, and gestures, and then delivering them to the user via voice or text.

[1241] Step 5:

[1242] When a user types a question into the robot, the question is analyzed using natural language processing tools, which extracts the intent of the question and related entities, thereby providing a clear understanding of the question.

[1243] Step 6:

[1244] Based on the analysis results, the robot retrieves the appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. The response is then provided to the user.

[1245] Step 7:

[1246] When a user requests guidance to a specific location, the robot begins navigation to the destination. First, it calculates the optimal path to the destination and then autonomously moves along that path, avoiding obstacles and other people along the way to safely reach the destination.

[1247] Step 8:

[1248] The robot records the results of its interactions with the user and their actions in a database, which can be used for future interactions. This record includes the interaction history, the user's reactions, and any new information acquired. This allows the system to continuously learn and improve the quality of its services.

[1249] In this way, the system of the present invention realizes natural interaction with the user and provides a variety of services.

[1250] Example 1

[1251] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1252] Conventional humanoid robot systems have had problems with natural interaction with users and efficient service provision. Specifically, they suffer from low recognition accuracy, delayed responses, and limited ability to support different languages. Furthermore, there is still room for improvement in terms of recognizing users' faces and gestures, identifying emotions, providing appropriate information, and navigation.

[1253] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1254] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and initial setting means. This allows the server to recognize the user's face and facial expressions, identify their ID and emotions, and provide appropriate information in multiple languages, thereby realizing natural interaction with the user and enabling efficient service provision.

[1255] "Image recognition means" refers to means that has the function of analyzing image data acquired using an imaging device such as a camera and recognizing a specific object (for example, a user's face or gesture).

[1256] "Natural language processing means" refers to means that include algorithms and models for analyzing text data entered by a user and understanding their intent and emotions.

[1257] A "multilingual translation means" is a means that has the function of translating text between different languages, and that understands input from a user in different languages ​​and generates corresponding information.

[1258] "Database connection means" refers to a means that has the function of communicating with a database server and acquiring and saving the necessary information.

[1259] The "means for performing initial settings" refers to a means for performing basic settings for system operation at the start, including loading various models and setting up a database.

[1260] "Means for recognizing a user's face and facial expressions" refers to means that has the function of extracting specific features from a user's facial image captured by a camera, identifying the user based on the results, and determining their emotional state.

[1261] "Means for identifying a user's ID and emotions" refers to a means for analyzing data obtained through image recognition to determine who the user is (ID) and their emotional state at the time.

[1262] "Means for obtaining a user profile" means means capable of obtaining information related to a user from a database, including past conversation history and preferences.

[1263] The "means for generating a greeting message" is a means having a function for automatically generating a greeting message suitable for a user based on the acquired user profile.

[1264] "Means for analyzing the intent of a question and related entities" refers to a means for analyzing the question entered by the user and identifying the intent of the question and related keywords (entities).

[1265] The "means for generating an appropriate answer to a question" is a means for obtaining relevant information based on the intent and entities of the analyzed question, and automatically generating an appropriate answer based on that information.

[1266] The "means for continuously capturing image data" refers to a means having a function for activating a camera and continuously acquiring image data at regular intervals.

[1267] The "means for preprocessing image data" refers to a means having a function for performing preprocessing such as resizing, noise removal, and color conversion on the acquired image data.

[1268] "Means for processing sensor information in real time and moving autonomously" refers to means that have the ability to analyze information from various sensors in real time, calculate a route according to the environment, and move autonomously.

[1269] The present invention relates to a multimodal AI system using a humanoid robot. The system includes image recognition means, natural language processing means, multilingual translation means, and database connection means. It also includes means for recognizing a user's face and facial expressions, identifying the user's ID and emotion, acquiring a user profile, generating a greeting, analyzing the intent of a question and related entities, generating appropriate answers to the question, continuously capturing image data, preprocessing the image data, and processing sensor information in real time to move autonomously.

[1270] 1. System initialization

[1271] The server performs the initial system configuration. First, for image recognition, it loads a ResNet50 model using TensorFlow using an NVIDIA GPU. Next, for natural language processing, it loads the BERT model using Hugging Face's Transformers library. It also configures the Google Translate API for multilingual translation, and starts a MySQL database to create the necessary tables and indexes.

[1272] 2. User Awareness

[1273] The robot captures the user's face and gestures using a built-in camera (e.g., a standard HD camera). The image data is preprocessed with OpenCV and then analyzed using a ResNet50 model running on NVIDIA GPUs, which identifies the user's identity, face, and facial expressions.

[1274] 3. Start a conversation

[1275] The robot starts a conversation using the data mentioned above. First, it uses the user's ID to retrieve the user profile from a MySQL database. Then, it uses that information to generate an appropriate greeting using Hugging Face's BERT model. For example, it might generate a greeting like, "Hello, Tanaka-san. How's your day?"

[1276] 4. Answering questions

[1277] Users input questions into the robot, which then uses voice input (e.g., a standard microphone) to receive the question and convert it into text using the Google Speech-to-Text API. It then uses Hugging Face's BERT model to analyze the intent of the question and relevant entities, and based on the analysis results, queries the appropriate API (e.g., the OpenWeatherMap API) or database to generate an answer to the question.

[1278] 5. Autonomous Movement and Navigation

[1279] The robot autonomously moves to a location specified by the user. First, the user specifies the destination via voice input or touchscreen. Next, the robot uses SLAM technology to generate a map of the environment and calculates the optimal route using algorithms such as the A algorithm. The robot autonomously moves along the calculated route using a LiDAR sensor and motor control unit, avoiding obstacles and reaching the destination.

[1280] Specific examples

[1281] In-store directions

[1282] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and first identifies the user using image recognition. It then analyzes the intent of the question using natural language processing and generates an appropriate response. The robot responds, "The reception desk is on the left side of this floor. I'll show you there," and autonomously guides the user to the reception desk using SLAM technology and a motor control unit.

[1283] Bank balance inquiry

[1284] When a user asks, "What is my account balance?", the robot first authenticates the user and retrieves account information from the database. After successful authentication, the robot responds, "Your current account balance is 100,000 yen."

[1285] Prompt Sentence Examples

[1286] "When a user asks for weather information, how do you get the information from the API and respond?"

[1287] "Please explain in detail the algorithm that recognizes user emotions."

[1288] As explained above, the system of the present invention realizes natural interaction with users and aims to reduce the workload and improve service quality through highly efficient service provision.

[1289] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1290] Step 1: Initialize the system

[1291] The server performs the initial system setup, inputting the image recognition model, natural language processing model, multilingual translation model, and database configuration information.

[1292] 1. As a deep learning model for image recognition, we load a ResNet50 model in TensorFlow using an NVIDIA GPU. The input is the TensorFlow model file, and the output is the model loaded in memory.

[1293] 2. Load the BERT model from Hugging Face's Transformers library as the natural language processing model. The input is the model file, and the output is the BERT model loaded in memory.

[1294] 3. Set up the Google Translate API as a multilingual translation model. The input is the API key and configuration information, and the output is the API availability status.

[1295] 4. Starts a MySQL database and creates the necessary tables and indexes. The input is database configuration information, and the output is a configured database ready to use.

[1296] Step 2: User Awareness

[1297] The robot uses a camera to recognize the user, and the input is image data obtained from the camera.

[1298] 1. Start a camera (e.g., HD camera) and continuously capture image data. The input is the camera sensor data, and the output is the captured image data.

[1299] 2. The acquired image data is preprocessed using the OpenCV library, which includes resizing, noise removal, color conversion, etc. The input is raw image data, and the output is preprocessed image data.

[1300] 3. The preprocessed image data is input to a ResNet50 model running on an NVIDIA GPU to identify the user's ID and emotion. The input is the preprocessed image data, and the output is the identified user's ID and emotion information.

[1301] Step 3: Start a conversation

[1302] The robot starts a conversation based on the recognized user information, including the user's ID and emotional information.

[1303] 1. The robot uses the user's ID to retrieve the user profile from the MySQL database. The input is the user's ID and the output is the retrieved user profile.

[1304] 2. Based on the retrieved user profile, the robot uses Hugging Face's BERT model to generate an appropriate greeting, such as "Hello, Tanaka-san. How are you today?" The input is the user profile, and the output is the generated greeting.

[1305] Step 4: Answer questions

[1306] The user inputs a question to the robot, and the input is voice data.

[1307] 1. Collect user questions using a voice input function (e.g., microphone). The input is voice data, and the output is the question data captured as voice.

[1308] 2. The robot converts voice data into text data using the Google Speech-to-Text API. The input is voice data and the output is text data.

[1309] 3. The text data is analyzed using Hugging Face's BERT model to extract the question intent and related entities. The input is the text data, and the output is the analysis results.

[1310] 4. Based on the analysis results, the robot queries the appropriate API (e.g., OpenWeatherMap API) or database to generate an answer to the question. The input is the analysis results and the necessary API information, and the output is the generated answer.

[1311] Step 5: Autonomous Movement and Navigation

[1312] The robot moves autonomously to a specified location, and the input is a destination specified by the user.

[1313] 1. The user specifies a destination using voice input or a touch screen. The input is the destination information, and the output is the specified destination.

[1314] 2. The robot uses SLAM technology to generate a map of the environment and calculate the optimal path. The input is sensor information and destination information, and the output is the calculated path information.

[1315] 3. The robot uses a LiDAR sensor and a motor control unit to move autonomously based on a calculated path. The input is path information and real-time sensor information, and the output is autonomous movement behavior.

[1316] (Application example 1)

[1317] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1318] Modern stores and public facilities require a means for visitors to quickly and accurately find their destination and obtain information. However, conventional guidance systems and information centers face problems such as labor shortages and language barriers. Furthermore, visitors often have to search for the information they need themselves, which is inconvenient. A system that solves these issues and allows visitors to easily reach their destination is needed.

[1319] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1320] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, and database connection means. This enables the server to recognize a user's face and facial expressions, analyze the user's questions, generate information corresponding to the user's language, and acquire a user profile and dialogue history. The server also includes connection means with a smart device, means for activating and controlling the robot via the smart device, autonomous robot movement means, and means for guiding the robot to a user-specified destination using the autonomous movement means. This allows visitors to easily call the robot from their smartphones, and the robot automatically guides the visitor to their destination, significantly improving visitor convenience.

[1321] "Image recognition means" refers to a device or system that uses a camera or sensor to recognize a user's face, facial expressions, and gestures.

[1322] "Natural language processing means" is a technology that analyzes questions and instructions uttered by users and understands their intent and content.

[1323] "Multilingual translation means" refers to technology that translates between different languages ​​and generates information corresponding to the language used by the user.

[1324] "Database connection means" refers to technology for accessing a database that stores user profiles, interaction history, and other necessary data, and acquiring that information.

[1325] "Means for connecting with smart devices" refers to technology that connects devices such as smartphones and tablets to the system, enabling two-way communication.

[1326] "Means for activating and controlling robots" refers to technology for activating robots via smart devices and controlling their movements and functions.

[1327] "Autonomous robot mobility means" refers to technology that enables a robot to move autonomously while recognizing its surrounding environment.

[1328] "Generating information" is the process of generating appropriate answers or guidance information based on the user's questions or instructions.

[1329] A "user profile" is data that compiles information about a user (e.g., the user's ID, past interaction history, preferences, etc.).

[1330] "Dialogue history" is data that records past conversations and interactions with a user.

[1331] "Autonomous movement" refers to a robot's ability to move while adapting to its surrounding environment based on its own judgment.

[1332] The present invention relates to a system that controls a robot in cooperation with a smart device to provide guidance and information to users in stores and public facilities. This system is realized using the following various means.

[1333] 1. Image Recognition Methods:

[1334] The server uses cameras and sensors to recognize the user's face, facial expressions, and gestures, using software libraries such as OpenCV and TensorFlow, which are known for their image processing technology.

[1335] 2. Natural Language Processing Tools:

[1336] The server uses natural language processing techniques to analyze the user's questions and instructions. This process uses various NLP (Natural Language Processing) libraries and tools (e.g., Google's Dialogflow, OpenAI's GPT-3).

[1337] 3. Multilingual translation tools:

[1338] The server uses a multilingual translation engine to translate between different languages ​​and generate information corresponding to the user's language, for example, the Google Translate API.

[1339] 4. Database connection method:

[1340] The server retrieves information such as the user's profile and interaction history from a database, using Firebase or AWS RDS to dynamically manage user information.

[1341] 5. Connectivity with smart devices:

[1342] The smartphone application communicates with the robot via Bluetooth or Wi-Fi to activate and control it, allowing users to summon the robot using their smartphone.

[1343] 6. Autonomous robotic mobility:

[1344] The robot autonomously navigates to its destination using SLAM (Simultaneous Localization and Mapping) technology with Lidar sensors and cameras, using ROS Lidar and Google Cartographer.

[1345] 7. Information generation:

[1346] The server generates appropriate answers and guidance information in response to user questions and instructions, using a generative AI model in the process.

[1347] 8. User Awareness and Guidance:

[1348] The robot recognizes the user with a camera, analyzes the intent of the question using natural language processing technology, and provides guidance. If the user asks, "Where is the dairy section?", the robot will use the store's map information to answer, "The dairy section is on the right side of this floor. I'll show you." and begin guiding the user. Using its autonomous mobility, the robot will accurately guide the user to their destination.

[1349] 9. Example prompt:

[1350] An example of a prompt sentence to input to the generative AI model is as follows:

[1351] "A user is in a store asking about the dairy section. The robot recognizes the user with a camera and analyzes the intent of the question using NLP technology. It then retrieves store map information from a database and provides directions. Please explain the specific processing flow in detail."

[1352] In this way, the system of the present invention can provide an environment in which the user can quickly and accurately obtain information regardless of the environment.

[1353] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1354] Step 1:

[1355] The user starts the smartphone app and presses the "Start Guide Robot" button, at which point the smartphone uses its camera and microphone to check the surrounding environment.

[1356] input:

[1357] User presses a button

[1358] output:

[1359] Camera and microphone ambient data

[1360] Specific behavior:

[1361] The camera captures image data of the surroundings, and the microphone captures audio data.

[1362] Step 2:

[1363] The device calls the robot via Bluetooth or Wi-Fi and sends commands for the robot to move to the user's location.

[1364] input:

[1365] Camera and microphone data

[1366] output:

[1367] Robot commands

[1368] Specific behavior:

[1369] It connects to the robot communication module via Bluetooth or Wi-Fi and sends the command "move to the user."

[1370] Step 3:

[1371] The robot recognizes the user using a camera and analyzes the user's face and gestures using image recognition means.

[1372] input:

[1373] Image data acquired by the camera

[1374] output:

[1375] User ID, facial expression, and gesture information

[1376] Specific behavior:

[1377] Image processing libraries such as OpenCV and TensorFlow are used to recognize and analyze the user's face and gestures and obtain user information.

[1378] Step 4:

[1379] The server uses natural language processing means to analyze the user's question and understand its intent.

[1380] input:

[1381] User utterances

[1382] output:

[1383] User questions and their intent

[1384] Specific behavior:

[1385] NLP tools such as Google's Dialogflow and OpenAI's GPT-3 are used to interpret the content of the speech and extract the intent of the question.

[1386] Step 5:

[1387] The server uses database connectivity to obtain the user's profile and interaction history and generates appropriate information.

[1388] input:

[1389] User ID, question content

[1390] output:

[1391] User profile and corresponding answer information

[1392] Specific behavior:

[1393] It connects to Firebase or AWS RDS, obtains the user's past interaction history and profile information, and then uses a generative AI model to generate appropriate answers and guidance information.

[1394] Step 6:

[1395] The server uses a multilingual translation means to translate the generated information into the language used by the user.

[1396] input:

[1397] Generated answer information

[1398] output:

[1399] Information translated into the user's language

[1400] Specific behavior:

[1401] Use the Google Translate API to translate the generated information into the user's language.

[1402] Step 7:

[1403] The robot presents appropriate information to the user via voice or text and begins providing guidance in response to the user's questions.

[1404] input:

[1405] Translated information

[1406] output:

[1407] User guidance information

[1408] Specific behavior:

[1409] The text information is converted into speech using a speech synthesis engine (e.g., Google TTS) and provided to the user. The autonomous mobility system is also used to calculate the route to the user's specified destination and begin providing guidance.

[1410] Step 8:

[1411] The robot uses autonomous mobility to guide the user to their destination while avoiding obstacles.

[1412] input:

[1413] User-specified destination information and surrounding environment information

[1414] output:

[1415] Guidance to the destination, reaching the final point

[1416] Specific behavior:

[1417] It uses Lidar sensors and SLAM technology (e.g., ROS Lidar, Google Cartographer) to map the surrounding environment, calculates routes and avoids obstacles in real time, and guides the user to a specified destination.

[1418] As described above, by acquiring and analyzing the necessary input data at each step and generating appropriate output, the system of the present invention can smoothly answer the user's questions and provide guidance. The system of the present invention provides visitors with fast and accurate guidance.

[1419] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1420] The present invention relates to a multimodal AI system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine to achieve natural interaction with users. Specific embodiments of the system are described below.

[1421] System initialization

[1422] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[1423] User Awareness

[1424] The robot uses a camera to recognize the user. During this process, the camera captures the user's face, facial expressions, and gestures, and then analyzes the data using image recognition. The analysis results in the user's ID, emotions, and gestures, allowing the system to understand who the user is and what emotional state they are in.

[1425] Emotion recognition

[1426] The robot uses an emotion engine to analyze the user's emotions from the captured facial expressions and voice. The emotion engine analyzes the data from each facial expression to determine whether the user is happy, confused, tired, etc.

[1427] Start a conversation

[1428] The robot uses the user information obtained by the recognition means and emotion engine to initiate a conversation using natural language processing means. First, it retrieves the user profile from a database and generates an appropriate greeting based on that information. This greeting is adjusted taking into account the user's emotional state. This allows the user to begin a natural dialogue with the system.

[1429] Answering questions

[1430] When a user inputs a question into the robot, the question is analyzed using natural language processing. During the analysis process, the intent of the question and related entities are extracted. Based on this information, an answer is generated by referencing the appropriate API or database. Furthermore, by incorporating the analysis results of the emotion engine into the response, the robot responds with an appropriate tone and content according to the user's emotional state.

[1431] Autonomous Mobility and Navigation

[1432] The robot can autonomously navigate to a destination specified by the user. Specifically, it calculates the optimal path to the destination and moves along that path, safely reaching the destination while avoiding obstacles and other people along the way. It can also adjust the tone and content of its guidance depending on the user's emotional state during the journey.

[1433] Specific examples

[1434] Example 1: Store directions

[1435] If a user asks, "Where is the reception desk?", the robot will capture the user's face and facial expressions with a camera and analyze them using image recognition and an emotion engine. It will then analyze the intent of the question using natural language processing, and the system will generate an appropriate response. The robot will respond, "The reception desk is on the left side of this floor. I'll show you," and actually guide the user to the reception desk. If the robot determines that the user appears confused, it can add a particularly detailed explanation.

[1436] Example 2: Bank balance inquiry

[1437] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from the database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and responds, "Your current account balance is 100,000 yen." If the user is nervous, the robot can also add additional comments to help them relax.

[1438] Through these functions, the system of the present invention realizes natural interaction with users and provides a variety of services. Utilizing the emotion engine enables advanced responses that take user emotions into consideration, significantly improving the quality of services.

[1439] The processing flow will be explained below.

[1440] Step 1:

[1441] The server initializes the entire system, first loading the image recognition model, natural language processing (NLP) model, multilingual translation model, and emotion engine, then setting up the database used by the system to store user profiles and interaction history.

[1442] Step 2:

[1443] The robot utilizes a camera to capture the user's face, facial expressions, and gestures, which are then analyzed using image recognition to obtain the user's identity, emotions, and gestures.

[1444] Step 3:

[1445] The robot uses an emotion engine to analyze the user's emotional state, extracting emotional data from facial expressions, voice tone, and gestures to determine whether the user is happy, confused, tired, etc.

[1446] Step 4:

[1447] The robot uses natural language processing to initiate a conversation based on the acquired user information. First, it retrieves the user profile from a database and generates a greeting based on that information. The greeting is adjusted to reflect the user's emotional state.

[1448] Step 5:

[1449] When a user types a question into the robot, the question is analyzed using natural language processing, which extracts the intent of the question and related entities to clearly understand the question.

[1450] Step 6:

[1451] Based on the analysis results, the robot retrieves appropriate information and generates a response. For example, if the question is about the weather, the robot will call a weather information API to retrieve the latest weather information and generate a response. At this time, the robot will take into account the user's emotional state obtained from the emotion engine and adjust the tone and content of the response.

[1452] Step 7:

[1453] When a user requests guidance to a specified location, the robot begins navigation. First, it calculates the optimal path to the destination and autonomously moves along that path. During the journey, it navigates safely while avoiding obstacles and other people. It also monitors the user's emotional state in real time and provides guidance accordingly.

[1454] Step 8:

[1455] The robot records the results of its interactions with the user in a database. Dialogue history and emotional data are accumulated and used for future interactions. This allows the system to continuously learn and improve the quality of its services.

[1456] Specific examples

[1457] Example 1: Store directions

[1458] When a user asks, "Where is the reception desk?", the robot uses a camera to capture the user's face and facial expressions, and analyzes them using image recognition and an emotion engine. Using the user's emotional state and profile information obtained based on the analysis, the robot generates a response using natural language processing. The robot responds in a friendly tone, saying, "The reception desk is on the left side of this floor. I'll show you there," and begins navigation.

[1459] Example 2: Bank balance inquiry

[1460] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves account information from a database. If authentication is successful, the robot uses its emotion engine to determine the user's emotional state and generate a response accordingly. It responds in a calm tone, saying, "Your current account balance is 100,000 yen," and can also add additional comments to ease the tension during the question.

[1461] This embodiment allows users to receive flexible responses that take their emotions into consideration, improving the quality of service.

[1462] Example 2

[1463] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1464] Conventional dialogue systems have limited interaction with users and lack the ability to respond with emotion or to handle diverse languages. Furthermore, their user recognition and autonomous movement capabilities are limited, making it difficult to achieve natural and effective dialogue. This results in a poor user experience and reduced system utilization efficiency.

[1465] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes an image recognition means, a natural language processing means, a multilingual translation means, a database connection means, an emotion analysis means, a user recognition means, and an autonomous movement means. This makes it possible to recognize the user's face and facial expression, analyze questions, generate information in multiple languages ​​while taking into account the emotional state, acquire a user profile and dialogue history, and autonomously move to a specified destination.

[1466] "Image recognition means" refers to a means of acquiring image data such as a user's face, facial expressions, and gestures using a camera or sensor, and analyzing that data.

[1467] A "natural language processing means" is a means for analyzing text or voice questions entered by a user, understanding their intent, and generating an appropriate response.

[1468] The "multilingual translation means" is a means for translating text input in different languages ​​and generating information corresponding to multiple languages.

[1469] "Database connection means" refers to a means for accessing a database to obtain and store related data such as user profiles and interaction history.

[1470] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data to identify their emotional state.

[1471] "User recognition means" refers to a means for identifying a user's ID and distinguishing between individual users.

[1472] An "autonomous vehicle" is a vehicle that moves autonomously to a specified destination while avoiding obstacles.

[1473] The present invention relates to a system that integrates image recognition means, natural language processing means, multilingual translation means, database connection means, emotion analysis means, user recognition means, and autonomous mobility means to realize natural interactions with users. Detailed embodiments of the system are described below.

[1474] System configuration

[1475] Hardware

[1476] This system uses the following hardware components:

[1477] Camera (e.g. high-resolution webcam)

[1478] microphone

[1479] speaker

[1480] Robot body (including motors and sensors for autonomous movement)

[1481] server

[1482] Database (e.g. MySQL, PostgreSQL)

[1483] software

[1484] The system consists of the following software components:

[1485] Image recognition models (e.g., TensorFlow, PyTorch)

[1486] Natural language processing models (e.g., GPT-4)

[1487] Multilingual translation models (e.g., Google Translate API)

[1488] Sentiment analysis engine (e.g. Affectiva SDK)

[1489] Database management systems (e.g., MySQL, PostgreSQL)

[1490] Autonomous movement algorithms (e.g., SLAM technology)

[1491] System Operation

[1492] System initialization

[1493] The server first initializes the entire system. This initialization includes loading image recognition models, natural language processing models, multilingual translation models, and a sentiment analysis engine, as well as setting up a database. Specifically, the server loads trained models into memory using TensorFlow or PyTorch, and initializes the sentiment analysis engine using the Affectiva SDK. The database uses MySQL or PostgreSQL, and is prepared to store user profiles and interaction history.

[1494] User Awareness

[1495] The robot recognizes the user through a camera. This involves capturing the user's face, facial expressions, and gestures, and then analyzing the data using libraries such as OpenCV. An emotion analysis engine is also integrated into this process to analyze the user's emotional state. The user's ID and emotional state are then used as the basis for the system to provide personalized responses.

[1496] Emotion recognition

[1497] The robot analyzes the user's emotions from the captured facial expressions and voice. It uses an emotion analysis engine to identify specific emotions (e.g., happiness, sadness, surprise), allowing the system to understand the user's current emotional state and prepare to generate a response accordingly.

[1498] Start a conversation

[1499] The robot initiates a conversation using an NLP model based on the user's profile and emotional state. The server queries a database to retrieve the user's profile. Using an NLP model (e.g., GPT-4), it generates an appropriate greeting for the user and speaks it through the speaker. The greeting is adjusted based on the user's emotional state.

[1500] Answering questions

[1501] When a user types a question into the robot, it is analyzed using a natural language processing model. Analysis includes extracting the intent of the question and relevant entities. The server then references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer, which the robot then speaks in a tone that best suits its emotional state.

[1502] Autonomous Mobility and Navigation

[1503] The robot autonomously navigates to a destination specified by the user. It uses an algorithm based on SLAM technology to calculate the optimal path, and moves along that path while avoiding obstacles. Along the way, it adjusts the tone and content of its guidance according to the user's emotional state.

[1504] Examples of concrete examples and prompts

[1505] Example 1: Store directions

[1506] When a user asks, "Where is the reception desk?", the robot captures the user's face and facial expressions with a camera and analyzes them using image recognition and an emotion engine. It then analyzes the intent of the question using natural language processing, generates and speaks an appropriate response, and guides the user to the reception desk, providing detailed explanations if the user appears confused.

[1507] Example prompt for a generative AI model: "The user is asking where the reception desk is. The robot will politely provide directions. Please respond especially politely if the user seems confused."

[1508] Example 2: Bank balance inquiry

[1509] When a user asks, "What is my account balance?", the robot authenticates the user and retrieves the account information from the database. If authentication is successful, the emotion engine analyzes the user's emotional state and responds with the current account balance. If the user is nervous, it adds a comment to relax them.

[1510] Example prompt for a generative AI model: "The user asks for their account balance. The robot responds by adding a relaxing comment based on the user's emotional state."

[1511] The above is a specific embodiment of the present invention, and the system allows users to experience natural conversations that take emotion into consideration.

[1512] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1513] Step 1: Initialize the system

[1514] The server initializes the entire system. First, it loads an image recognition model into memory using TensorFlow or PyTorch. Next, it loads a natural language processing (NLP) model such as OpenAI's GPT-4, and then initializes a multilingual translation model such as the Google Translate API and a sentiment analysis engine using the Affectiva SDK. It also sets up a database using MySQL or PostgreSQL to store user profiles and interaction history. This completes the initialization process, making the various models ready for processing. It receives model configuration files and database configuration information as input, and the model loading is complete as output.

[1515] Step 2: Getting started with user awareness

[1516] The robot recognizes the user through a camera. Specifically, it uses the camera to capture the user's face, facial expressions, and gestures, and analyzes the data using the OpenCV library. As a result of the analysis, the user's facial feature points are extracted, and the emotion analysis engine analyzes the user's emotional state. The input is the captured image data, and the output is the user's ID and emotional state.

[1517] Step 3: Detailed analysis of emotional state

[1518] The robot performs a detailed analysis of the user's emotions from the captured facial expressions and voice. Using the Affectiva SDK, it analyzes facial expression data to identify specific emotions (such as joy, sadness, or surprise). It also collects the user's voice data and analyzes its emotional nuances. It receives facial expression and voice data as input and outputs the identified emotional state.

[1519] Step 4: Start a conversation

[1520] The robot starts a conversation using an NLP model based on the user's profile and emotional state. The server executes an SQL query to retrieve the user profile from the database. Using the NLP model (GPT-4), it generates an appropriate greeting based on the user's profile and emotional state. It then speaks this to the user through the speaker. The input is the user profile and emotional state retrieved from the database, and the output is the generated greeting.

[1521] Step 5: Answer questions

[1522] The user inputs a question to the robot. The user uses a voice recognition system to input the question in text format. The server receives the text data and uses an NLP model to analyze the intent of the question and extract relevant entities. Based on this, it references the appropriate API (e.g., weather API, exchange rate API) or database to generate an answer. It then takes into account the results of the sentiment analysis engine to adjust the tone and content of the response and speaks the answer to the user through the speaker. The input is the user's question text, and the output is the generated answer.

[1523] Step 6: Autonomous Movement and Navigation

[1524] The robot moves autonomously to a destination specified by the user. First, the user inputs instructions for the destination. The robot uses SLAM technology to calculate the optimal path to the destination and moves along that path. During movement, it uses LiDAR and ultrasonic sensors to detect obstacles and avoid them to proceed safely. It also adjusts the tone and content of its guidance according to the user's emotional state. It receives destination instructions and sensor data as input, and outputs the movement results along the optimal path.

[1525] (Application example 2)

[1526] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1527] Today's elderly often have difficulty communicating with store staff and navigating the store when shopping or using services in physical stores. Furthermore, language barriers and a lack of emotional support reduce the elderly's satisfaction. This reduces the opportunities for elderly people to enjoy shopping in physical stores and causes stress, which is an issue.

[1528] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1529] In this invention, the server includes image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. This enables elderly people to shop or use services in brick-and-mortar stores by recognizing faces and facial expressions, analyzing questions, generating language-compatible information, acquiring profiles and dialogue histories, and analyzing emotional states, thereby providing appropriate responses and in-store navigation based on their emotions.

[1530] "Image recognition means" refers to technology that uses a device such as a camera to recognize a user's face, facial expressions, and gestures.

[1531] "Natural language processing means" is a technology that analyzes questions and requests entered by users and understands their intentions.

[1532] "Multilingual translation means" refers to technology for translating and generating information in accordance with the language used by the user.

[1533] "Database connection means" refers to technology for connecting to a database that stores user profiles and interaction history, and obtaining the necessary information.

[1534] The "emotion engine" is a technology that analyzes the user's emotional state from their facial expressions and tone of voice to determine what emotions the user is feeling.

[1535] A "system" is a collection of components that integrate multiple technical means to provide specific functions or services.

[1536] "In-store navigation" is a function that guides users to find desired locations and products within a store.

[1537] "Support for the elderly" means providing support and services to help the elderly live their daily lives more conveniently and safely.

[1538] "Multilingual support" means providing services and information in different languages ​​to users who speak different languages.

[1539] "Appropriate response" means providing answers or guidance appropriate to the situation based on the user's question or condition.

[1540] A "profile" is data that records basic information and personal characteristics about a user.

[1541] System configuration

[1542] This invention is realized by a system that combines image recognition means, natural language processing means, multilingual translation means, database connection means, and an emotion engine. The components of this system are described in detail below.

[1543] Hardware and Software

[1544] Image Recognition: The system recognizes the user's face and facial expressions using a smartphone with a camera or a head-mounted display. It uses the OpenCV (cv2) library and a model trained in TensorFlow.

[1545] Natural Language Processing: Natural language processing uses the Transformers library pipeline to analyze question answers, using generative AI models such as BERT and RoBERTa.

[1546] Multilingual translation method: For multilingual translation, the DeepTranslator library is used, and natural translation is provided in conjunction with the Google Translate API.

[1547] Database connection method: User profiles and interaction history are managed using SQLite, and the necessary information is obtained.

[1548] Emotion engine: Emotion analysis uses the Transformers library pipeline to analyze emotions from voice and facial expressions.

[1549] System initialization

[1550] Initialization takes place on the server, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing user profiles and interaction history.

[1551] User recognition and sentiment analysis

[1552] The device's camera captures the user's face, facial expressions, and gestures, and analyzes them using image recognition.The emotion engine analyzes the user's emotional state and determines how they are feeling.

[1553] Initiating conversations and answering questions

[1554] The system analyzes questions using natural language processing and generates an appropriate greeting based on the user's profile. It also adjusts the tone and content of the response based on the user's emotional state, using an emotion engine. The system offers conversations that are especially designed for seniors, ensuring their peace of mind.

[1555] In-store navigation

[1556] Based on the user's gestures, the system guides the user to the desired location or product within the store. The system also reflects the user's emotional state during navigation, adjusting the navigation to ensure a safe and secure journey.

[1557] Examples and prompts

[1558] For example, if an elderly person asks, "Do you have the shirt I'm looking for in stock?", the smartphone camera will recognize the user's face and perform image recognition and emotion analysis. Natural language processing will then be used to obtain the appropriate stock information, which will be translated into multiple languages ​​and displayed to the user.

[1559] Prompt Sentence Examples

[1560] "Do you have the shirt I'm looking for in stock?"

[1561] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1562] Step 1:

[1563] The server initializes the entire system, loading the image recognition model, natural language processing model, multilingual translation model, and emotion engine, as well as setting up the database and preparing the user profile and interaction history.

[1564] Input: Image recognition model, natural language processing model, multilingual translation model, emotion engine, database

[1565] Output: Initialized models and database connections

[1566] Step 2:

[1567] The device uses a camera to capture the user's face, facial expressions, and gestures, and the captured image data is analyzed using image recognition tools.

[1568] Input: Video data of the user's face, expressions, and gestures

[1569] Output: Parsed user ID and emotional state

[1570] Step 3:

[1571] Based on the analysis results, the device further analyzes the user's emotional state using an emotion engine, thereby specifically grasping the user's emotional state.

[1572] Input: User's video data and analysis results

[1573] Output: Detailed emotional state data

[1574] Step 4:

[1575] The server retrieves the user profile and interaction history from the database and generates an appropriate greeting for the user. It uses natural language processing means to provide a greeting based on the user's emotional state.

[1576] Input: User ID, emotional state data, user profile and interaction history from database

[1577] Output: Emotion-based greeting

[1578] Step 5:

[1579] The user inputs a question into the terminal. The server analyzes the question using natural language processing to extract the intent of the question and related information. It then generates an appropriate answer based on the analysis results.

[1580] Input: User question

[1581] Output: Analyzed question intent and answer suggestions

[1582] Step 6:

[1583] The server uses a multilingual translation means to translate the generated answer into the user's language, and the translated result is displayed to the user.

[1584] Input: Analyzed question intent and answer suggestions

[1585] Output: The answer translated into the user's language

[1586] Step 7:

[1587] The device monitors the user's gestures and provides in-store navigation as needed, adjusting the tone and content of the guidance based on the user's emotional state.

[1588] Input: User gesture data and emotional state

[1589] Output: Navigation prompts and tuned tones

[1590] Through this series of steps, the system can provide an environment where seniors can shop and use services in physical stores in a natural and safe manner.

[1591] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1592] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1593] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1594] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1595] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1596] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1597] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1598] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1599] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1600] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1601] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1602] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1603] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1604] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1605] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1606] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1607] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1608] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1609] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1610] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1611] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1612] The following is further disclosed regarding the above embodiment.

[1613] (Claim 1)

[1614] Image recognition means;

[1615] natural language processing means;

[1616] Multilingual translation means;

[1617] A database connection means;

[1618] The image recognition means recognizes the face and facial expression of the user,

[1619] Analyzing the user's question using the natural language processing means;

[1620] generating information corresponding to the user's language using the multilingual translation means;

[1621] means for acquiring a user profile and a dialogue history by said database connection means;

[1622] A system including:

[1623] (Claim 2)

[1624] 10. The system of claim 1, further comprising means for recognizing a user's gesture by said image recognition means.

[1625] (Claim 3)

[1626] 2. The system according to claim 1, further comprising means for carrying out a wide range of business processes based on user questions using said natural language processing means.

[1627] "Example 1"

[1628] (Claim 1)

[1629] Image recognition means;

[1630] natural language processing means;

[1631] Multilingual translation means;

[1632] A database connection means;

[1633] A means for performing an initial setting;

[1634] means for recognizing the face and facial expressions of the user;

[1635] a means of determining the identity and sentiment of a user;

[1636] A means for obtaining a user profile;

[1637] means for generating a greeting;

[1638] A means of analyzing the intent of the question and related entities;

[1639] a means for generating appropriate answers to questions;

[1640] means for continuously capturing image data;

[1641] means for preprocessing image data;

[1642] A means of processing sensor information in real time and moving autonomously;

[1643] A system including:

[1644] (Claim 2)

[1645] 10. The system of claim 1, further comprising means for recognizing a user gesture.

[1646] (Claim 3)

[1647] 10. The system of claim 1, further comprising means for obtaining relevant information and generating an answer based on a user's question.

[1648] "Application Example 1"

[1649] (Claim 1)

[1650] Image recognition means;

[1651] natural language processing means;

[1652] Multilingual translation means;

[1653] A database connection means;

[1654] The image recognition means recognizes the face and facial expression of the user,

[1655] Analyzing the user's question using the natural language processing means;

[1656] generating information corresponding to the user's language using the multilingual translation means;

[1657] means for acquiring a user profile and a dialogue history by said database connection means;

[1658] A means of connection with a smart device;

[1659] means for activating and controlling the robot via the smart device;

[1660] an autonomous means of movement for the robot;

[1661] a means for guiding the autonomous moving means to a destination designated by a user;

[1662] A system including:

[1663] (Claim 2)

[1664] 10. The system of claim 1, further comprising means for recognizing a user's gesture by said image recognition means.

[1665] (Claim 3)

[1666] 2. The system according to claim 1, further comprising a means for performing a wide range of business processes based on a user's question using the natural language processing means, and for guiding the robot to a destination through autonomous movement.

[1667] "Example 2: Combining Emotion Engines"

[1668] (Claim 1)

[1669] Image recognition means;

[1670] natural language processing means;

[1671] Multilingual translation means;

[1672] A database connection means;

[1673] A sentiment analysis means;

[1674] A user recognition means;

[1675] an autonomous means of transportation;

[1676] The image recognition means recognizes the face and facial expression of the user,

[1677] Analyzing the user's question using the natural language processing means;

[1678] generating information corresponding to the user's language using the multilingual translation means;

[1679] Acquiring a user profile and a dialogue history by the database connection means;

[1680] Analyzing the emotional state of the user by the emotion analysis means;

[1681] Identifying the user's ID by the user recognition means;

[1682] The autonomous moving means autonomously moves to a designated destination.

[1683] A system characterized by:

[1684] (Claim 2)

[1685] 10. The system of claim 1, further comprising means for recognizing a user's gesture by said image recognition means.

[1686] (Claim 3)

[1687] 2. The system according to claim 1, further comprising means for carrying out a wide range of business processes based on user questions using said natural language processing means.

[1688] "Application example 2 when combining emotion engines"

[1689] (Claim 1)

[1690] Image recognition means;

[1691] natural language processing means;

[1692] Multilingual translation means;

[1693] A database connection means;

[1694] Emotion engine and

[1695] The image recognition means recognizes the face and facial expression of the user,

[1696] Analyzing the user's question using the natural language processing means;

[1697] generating information corresponding to the user's language using the multilingual translation means;

[1698] Acquiring a user profile and a dialogue history by the database connection means;

[1699] Analyzing the user's emotional state using the emotion engine;

[1700] means for generating an appropriate response according to the emotional state of the user;

[1701] A system including:

[1702] (Claim 2)

[1703] means for recognizing a user's gesture by the image recognition means;

[1704] 10. The system of claim 1, further comprising means for providing in-store navigation based on user gestures.

[1705] (Claim 3)

[1706] providing appropriate in-store navigation information based on the user's question using the natural language processing means;

[1707] 10. The system of claim 1, further comprising means for providing multilingual and emotion-based assistance for seniors. [Explanation of symbols]

[1708] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. Image recognition means; natural language processing means; Multilingual translation means; A database connection means; The image recognition means recognizes the face and facial expression of the user, Analyzing the user's question using the natural language processing means; generating information corresponding to the user's language using the multilingual translation means; means for acquiring a user profile and a dialogue history from said database connection means; A system including:

2. The system of claim 1 further comprising means for recognizing a user's gesture by said image recognition means.

3. 2. The system according to claim 1, further comprising means for carrying out a wide range of business processes based on user questions using said natural language processing means.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A