System

The system addresses the limitations of IoT devices by offering user authentication, personalized data management, natural language processing, and emergency response capabilities, ensuring comprehensive support and rapid emergency assistance.

JP2026022330APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123847
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Current IoT devices and smart speakers lack the ability to handle complex use cases such as consultations, suggestions, and emergency responses, failing to provide adequate support for individuals who need social interaction, particularly the elderly, emotionally unstable, or reclusive people.

Method used

A system that includes user authentication, personalized data management, natural language processing, emergency response, and educational features, utilizing a server and terminal to provide tailored assistance and rapid emergency services.

Benefits of technology

Enables comprehensive support for users by managing personalized data, providing appropriate responses, and ensuring rapid emergency assistance, enhancing user convenience and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022330000001_ABST
    Figure 2026022330000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: Means for receiving authentication information of a user and performing authentication by collating the authentication information with a database, means for starting a session in a case where the authentication succeeds and returning an error message in a case where the authentication fails, means for displaying a screen for inputting personalized data and storing the input data in the database, means for converting a voice of the user into a text, analyzing the text on the basis of a natural language processing engine, and generating an appropriate answer, the system includes a means for arranging an emergency response when necessary, and a means for providing an appropriate teaching material or training program in response to the learning or training request of the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Current IoT devices and smart speakers can only perform pre-defined tasks, and are unable to handle use cases such as consultations and suggestions. Additionally, people who need someone to communicate with, such as elderly people living alone, emotionally unstable people, and people who are reclusive, face problems such as being unable to provide help in emergencies, making incorrect decisions, and not knowing how to interact with society. [Means for solving the problem]

[0005] The present invention provides a means for receiving a user's authentication information and verifying it against a database to perform authentication, and a means for starting a session if authentication is successful and returning an error message if authentication fails. It also includes a means for displaying a screen for inputting personalized data and saving the input data in a database. Furthermore, it provides a means for converting the user's speech into text, analyzing the text using a natural language processing engine, and generating an appropriate response. The system also includes a means for acquiring the user's location information and health data in an emergency and arranging for emergency response as necessary. It also includes a means for providing appropriate educational materials and training programs in response to the user's learning and training requests. This makes it possible to comprehensively support the user's lifestyle.

[0006] "User authentication" is the process of verifying a user's identity when accessing a system.

[0007] A "database" is a system for efficiently storing, searching, and managing large amounts of data.

[0008] A "session" refers to a unit of a series of communications and data exchanges between a user and a system.

[0009] "Personalized data" refers to individualized information such as a user's individual preferences, characteristics, and past history information.

[0010] A "natural language processing engine" refers to technology that enables computers to understand and analyze human language and generate appropriate responses.

[0011] "Emergency keywords" are specific words or phrases that users utter in an emergency situation, which trigger the system to immediately begin responding.

[0012] "Emergency response" is the process for providing prompt assistance to users in the event of an emergency.

[0013] "Instructional Materials" refers to materials and tools provided for learning and training purposes.

[0014] A "training program" is a program that encompasses a series of learning activities designed to acquire specific skills or knowledge.

[0015] "Location Information" means data that identifies a user's current geographic location and enables a system to represent that information on a map or in other formats.

[0016] "Health data" refers to information related to a user's health condition, including, for example, heart rate, blood pressure, and body temperature. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[0039] 1. User Authentication

[0040] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[0041] 2. Managing Personalized Data

[0042] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[0043] 3. Daily conversation and problem consultation

[0044] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text. For example, if the user says, "I've been feeling tired from work lately," the device will suggest ways to relax and take a break.

[0045] 4. Emergency Response

[0046] If a user experiences an emergency, the device will detect this by uttering an emergency keyword such as "Help!" and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also provide support by contacting the user's family and friends in an emergency. For example, if a user yells "My chest hurts! Help!", the server will use the user's location information to call an ambulance and notify their family.

[0047] 5. Education and Training Functions

[0048] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," a basic pronunciation practice program is provided and instruction begins.

[0049] As described above, this system can provide comprehensive support to users throughout their lives.

[0050] The processing flow will be explained below.

[0051] 1. User Authentication

[0052] Step 1:

[0053] The user enters their username and password into the login screen.

[0054] Step 2:

[0055] The terminal transmits the entered authentication information to the server.

[0056] Step 3:

[0057] The server compares the received authentication information with the database and generates an authentication result.

[0058] Step 4:

[0059] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[0060] Step 5:

[0061] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[0062] 2. Managing Personalized Data

[0063] Step 1:

[0064] The terminal displays a personalized data entry screen when you log in for the first time.

[0065] Step 2:

[0066] Users enter personal data such as name, age, health information, and hobbies.

[0067] Step 3:

[0068] The terminal transmits the input data to the server.

[0069] Step 4:

[0070] The server stores the received data in a database.

[0071] 3. Daily conversation and problem consultation

[0072] Step 1:

[0073] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[0074] Step 2:

[0075] The device converts the user's voice into text and sends it to the server.

[0076] Step 3:

[0077] The server analyzes the received text using a natural language processing engine.

[0078] Step 4:

[0079] The server generates an appropriate answer based on the user's past data and current situation.

[0080] Step 5:

[0081] The server generates a response and sends it to the terminal.

[0082] Step 6:

[0083] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[0084] 4. Emergency Response

[0085] Step 1:

[0086] The user utters an emergency keyword such as "Help!"

[0087] Step 2:

[0088] The device detects the emergency keyword and immediately sends an alert to the server.

[0089] Step 3:

[0090] The server collects the user's location and health data.

[0091] Step 4:

[0092] The server automatically contacts local emergency services.

[0093] Step 5:

[0094] The server will also contact the user's family and friends in an emergency.

[0095] Step 6:

[0096] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[0097] 5. Education and Training Functions

[0098] Step 1:

[0099] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[0100] Step 2:

[0101] The device sends a request to the server.

[0102] Step 3:

[0103] The server selects appropriate educational materials and training programs.

[0104] Step 4:

[0105] The server transmits the selected teaching materials and programs to the terminal.

[0106] Step 5:

[0107] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[0108] Example 1

[0109] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0110] Conventional AI assistant systems have had difficulty responding appropriately to diverse user needs. In particular, they lacked the ability to respond quickly in emergencies, manage personalized data based on individual user needs, and provide advanced dialogue functions that combine voice recognition and voice synthesis technologies. As a result, user convenience and safety were not sufficiently ensured.

[0111] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0112] In this invention, the server includes means for receiving user authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's speech into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as needed, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for analyzing prompt sentences based on the user's diverse needs and requests using a generative AI model and generating responses based thereon, means for analyzing the user's speech and responding vocally using speech recognition and synthesis technology, and means for detecting emergency keywords and sending immediate alerts. This allows for a wide range of support to be provided to users, enabling rapid response in emergencies and the provision of services tailored to individual needs.

[0113] "User authentication" is the process of verifying a user's identity when accessing a system, usually through a username and password.

[0114] A "database" is an electronic system for managing, storing, retrieving, and updating data efficiently and effectively.

[0115] A "session" refers to a series of operations and communications between when a user logs in to the system and when they log out.

[0116] "Personalized Data" refers to information about a user (e.g., name, age, health information, hobbies, etc.) that the system uses to provide the user with the most appropriate service.

[0117] A "natural language processing engine" refers to an algorithm or model for understanding and processing human language, and is used, for example, to analyze the intent and content of a conversation.

[0118] "Emergency response" refers to the process of responding quickly to a user's emergency situation and arranging for the necessary emergency services.

[0119] A "generative AI model" refers to an artificial intelligence algorithm that automatically generates appropriate responses or content based on user input.

[0120] A "prompt" refers to guided text used to elicit an appropriate response from a generative AI model.

[0121] "Speech recognition" refers to the technology that analyzes a user's speech and converts it into text data.

[0122] "Speech synthesis" refers to the technology of converting text data into voice data and providing information to users via voice.

[0123] "Emergency keywords" are specific words or phrases that the system uses to determine an emergency, and when detected, an emergency response is initiated.

[0124] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[0125] User authentication

[0126] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database (e.g., MySQL or PostgreSQL), and if authentication is successful, it starts a session and returns a login success message to the user. If authentication fails, it returns an error message instructing the user to enter their authentication information again.

[0127] Managing Personalization Data

[0128] When the user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database (e.g., SQLite or MongoDB). This makes it possible to provide services tailored to the user's needs and preferences.

[0129] Daily conversation and problem consultation

[0130] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine (such as GPT-3 or BERT) to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text.

[0131] For example, if a user says, "I've been feeling tired at work lately," the system will suggest ways to relax and take a break. The prompts in this case could be:

[0132] Generate an error message if a user fails to log in.

[0133] "Provide health advice if a user enters that running is a hobby."

[0134] "Generate relaxation suggestions for users who have recently been feeling tired from work."

[0135] Emergency response

[0136] If a user experiences an emergency, they can simply say "Help!" or another emergency keyword. The device will detect this and immediately send an alert to the server. The server will then collect the user's location and health data, and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends to provide support.

[0137] For example, if a user cries out, "My chest hurts, help me!", the server will call an ambulance based on the user's location and notify their family.

[0138] Education and Training Features

[0139] When a user wishes to learn or train, for example by requesting "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows users to efficiently self-study and train their skills.

[0140] As a concrete example, if a user requests, "I want to practice my English pronunciation," we will provide them with a basic pronunciation practice program and begin providing instruction.

[0141] As described above, this system can provide comprehensive support to users throughout their lives.

[0142] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0143] Step 1:

[0144] The user enters their username and password on the terminal, which is the authentication input.

[0145] Input: Username and Password

[0146] How it works: A user enters their username and password into the device's login screen.

[0147] Output: The terminal stores the entered information and prepares it for transmission to the next step.

[0148] Step 2:

[0149] The terminal transmits the entered authentication information to the server.

[0150] Input: The username and password entered by the user

[0151] What happens: The device sends authentication information to the server as an HTTPS request (for example, URL: https: / / example.com / api / login).

[0152] Output: An HTTPS request containing authentication information is sent to the server.

[0153] Step 3:

[0154] The server checks the received authentication information against a database.

[0155] Input: Authentication information sent from the device (username and password)

[0156] What it does: The server compares the authentication information with the user information in a database (e.g. MySQL) to see if they match.

[0157] Output: Authentication result (success or failure)

[0158] Step 4:

[0159] The server returns the authentication result to the terminal.

[0160] Input: Authentication result (success or failure)

[0161] Operation: The server sends the authentication result to the terminal as an HTTPS response.

[0162] Output: The terminal receives the authentication result.

[0163] Step 5:

[0164] The device displays the authentication result to the user.

[0165] Input: Authentication result received from the server (success or failure)

[0166] Operation: If successful, the device will transition to the next screen. If unsuccessful, an error message will be displayed saying "Username or password is incorrect."

[0167] Output: The authentication result is displayed to the user.

[0168] Step 6:

[0169] The terminal displays a personalization data entry screen to the user.

[0170] Input: User successfully logged in information

[0171] Action: The device renders and displays the personalization data entry screen.

[0172] Output: The user is presented with a data entry screen.

[0173] Step 7:

[0174] The user enters individual information.

[0175] Input: Personalized data input screen

[0176] How it works: User enters name, age, health information, hobbies, etc.

[0177] Output: Individual personalized data is input to the device.

[0178] Step 8:

[0179] The terminal transmits the input data to the server.

[0180] Input: Personalized data you enter (such as name, age, health information, hobbies, etc.)

[0181] What happens: The device sends this as an HTTPS request to the server (e.g., URL: https: / / example.com / api / personalize).

[0182] Output: An HTTPS request containing personalization data is sent to the server.

[0183] Step 9:

[0184] The server stores the data in a database.

[0185] Input: Personalization data sent from your device

[0186] How it works: The server stores data in a database (e.g., SQLite or MongoDB).

[0187] Output: A message confirming successful save is generated.

[0188] Step 10:

[0189] The server returns a message to the terminal confirming successful saving.

[0190] Input: Personalized data saving success information

[0191] How it works: The server sends a confirmation message to the device as an HTTPS response.

[0192] Output: You will receive a save confirmation message on your device.

[0193] Step 11:

[0194] The device displays a confirmation message to the user.

[0195] Input: Save confirmation message from the server

[0196] What it does: The device displays "Data saved" to the user.

[0197] Output: User receives confirmation to save data.

[0198] Step 12:

[0199] The user speaks into the device.

[0200] Input: Questions and inquiries by voice

[0201] Action: The user speaks into the device.

[0202] Output: The device receives the audio data.

[0203] Step 13:

[0204] The device converts the speech to text.

[0205] Input: Audio data

[0206] How it works: Your device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert speech to text.

[0207] Output: Text data is generated.

[0208] Step 14:

[0209] The terminal transmits the converted text data to the server.

[0210] Input: Text data

[0211] What happens: The device sends text data as an HTTPS request to the server (for example, URL: https: / / example.com / api / chat).

[0212] Output: An HTTPS request containing text data is sent to the server.

[0213] Step 15:

[0214] The server analyzes the text using a natural language processing engine.

[0215] Input: Text data sent from the terminal

[0216] How it works: The server uses a natural language processing engine (e.g., GPT-3) to parse the text and generate an appropriate answer.

[0217] Output: The answer text is generated.

[0218] Step 16:

[0219] The server generates a response text and sends it back to the terminal.

[0220] Input: Generated answer text

[0221] How it works: The server sends the answer text to the device as an HTTPS response.

[0222] Output: The answer text is sent to the terminal.

[0223] Step 17:

[0224] The device displays the answer to the user.

[0225] Input: Response text from the server

[0226] How it works: The device uses speech synthesis technology (e.g., Google Cloud Text-to-Speech API) to synthesize a response and provide it to the user in voice or text.

[0227] Output: The user receives the answer.

[0228] (Application example 1)

[0229] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0230] In modern society, users need personalized services and support based on their individual needs and circumstances. However, current systems struggle to properly manage detailed health information and data and provide optimal recommendations and support. Furthermore, there is a lack of systems that provide integrated support for multiple functions across all aspects of daily life, including diet, health management, and emergency response. Another challenge is providing a system that allows users to easily request and order food and that can respond quickly in emergencies.

[0231] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0232] In this invention, the server includes means for receiving a user's authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating appropriate responses, means for managing the user's dietary history and health information, means for accepting and analyzing the user's dietary requests and orders via voice or text, and providing personalized dietary suggestions and orders, means for acquiring the user's location information and health data in an emergency and arranging for emergency response as necessary, and means for providing appropriate educational materials and training programs in response to the user's learning and training requests. This allows users to receive appropriate dietary suggestions and orders based on their individual dietary requirements, and ensures rapid and appropriate responses in emergencies.

[0233] "User authentication" is the process of receiving a user's authentication information and verifying it against a database.

[0234] "Personalized Data" is information that is collected and stored about you based on your specific attributes and preferences.

[0235] "Natural language processing" is a technology for analyzing text, understanding its meaning, and generating appropriate answers.

[0236] "Diet history" is a record of meals a user has consumed in the past.

[0237] "Health information" refers to information that includes data related to health, such as the user's health condition and disease information.

[0238] "Meal suggestion" is the process of recommending appropriate meals based on the user's dietary history and health information.

[0239] "Speech recognition" is a technology that converts speech into text.

[0240] "Location information" is data that indicates the user's current location.

[0241] "Emergency response" is the process of providing users with the assistance and services they need in an emergency.

[0242] A "learning or training program" is a collection of materials and exercises provided to users to help them learn or improve their skills.

[0243] A system for implementing this invention is an advanced AI assistant system that includes user authentication, personalized data management, dialogue using natural language processing, emergency response, education and training functions, etc. Specific embodiments will be described in detail below.

[0244] User authentication

[0245] The server receives the user's authentication information and authenticates it by checking it against a database. Specifically, when the user logs in for the first time, they enter their username and password, which are sent to the server via their terminal. The server checks the received authentication information against its existing database, and if authentication is successful, it starts a session, or returns an error message if it fails.

[0246] Managing Personalization Data

[0247] The server displays a screen for the user to enter personalization data and stores the data entered by the user in a database. For example, the user enters personal information such as name, age, health information, hobbies, etc., which are then stored in the database. This allows the server to provide services based on the user's specific needs and preferences.

[0248] Dialogue using natural language processing

[0249] When a user speaks, the device converts the speech into text and sends it to the server. The server then analyzes the text using a natural language processing engine that uses a generative AI model to generate an appropriate answer to the user's question or request. The answer is then returned to the user via the device in voice or text.

[0250] Specific examples

[0251] For example, if a user says, "What's your lunch recommendation?", the server will suggest the perfect lunch based on the stored personalized data. Also, if a user requests, "I want to practice my English pronunciation," the server will provide appropriate learning materials.

[0252] Emergency response

[0253] In the event of an emergency, the server will obtain the user's location and health data and arrange for emergency response if necessary. For example, if the user yells, "My chest hurts! Help!", the device will detect this emergency keyword and immediately send an alert to the server. The server will then obtain the user's location, arrange for an ambulance, and notify their family.

[0254] Education and Training Features

[0255] When a user wishes to learn or train, the server selects appropriate learning materials and training programs and provides them via the terminal, allowing the user to efficiently self-study and train their skills.

[0256] Hardware and Software

[0257] The system uses the following hardware and software:

[0258] Hardware: Devices such as smartphones, smart glasses, and head-mounted displays

[0259] software:

[0260] Speech Recognition Library: speech_recognition

[0261] Natural language processing library: transformers

[0262] Database management library: SQLAlchemy

[0263] Generative AI model: "cl-tohoku / bert-base-japanese" as an example

[0264] Examples of prompt statements

[0265] For example, a prompt to a generative AI model might look like this:

[0266] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[0267] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0268] Step 1:

[0269] User authentication: When a user logs in for the first time, they enter their username and password. The device sends this information to the server. The server checks the received authentication information against an existing database and starts the session if authentication is successful. If authentication fails, it returns an error message to the device.

[0270] Input: Username, Password

[0271] Processing: Match against database

[0272] Output: Session started or error message

[0273] Step 2:

[0274] Input of personalized data: If authentication is successful, the device will display a personalized data input screen for the user. The user will input personal information such as name, age, health information, hobbies, etc. This data will be sent from the device to the server, where it will be stored in a database.

[0275] Input: Name, age, health information, hobbies

[0276] Processing: Save to database

[0277] Output: Personalized data

[0278] Step 3:

[0279] Speech recognition and natural language processing: The user speaks a question or request into the device. The device converts the speech into text and sends this text to the server. The server uses a natural language processing engine with a generative AI model to analyze the text and generate an appropriate response. This response is sent back to the device and returned to the user as voice or text.

[0280] Input: Voice input (questions and requests)

[0281] Processing: Speech-to-text conversion, text analysis, and generating appropriate answers

[0282] Output: Text or audio response

[0283] Step 4:

[0284] Management of dietary history and health information: When a user makes a dietary question or request, the server analyzes the stored personalized data (dietary history and health information) and generates appropriate dietary suggestions. The suggestions are notified to the user via the device.

[0285] Input: dietary requests, past dietary history, health information

[0286] Processing: Retrieving information from the database, analyzing it, and generating meal suggestions

[0287] Output: Meal suggestion notification

[0288] Step 5:

[0289] Emergency Response: When a user utters a keyword such as "Help!" in an emergency, the device detects this and immediately sends an alert to the server. The server obtains the user's location and health data and automatically arranges for local emergency response if necessary. The server also notifies the user's family and friends in an emergency.

[0290] Input: Emergency keywords, location information, health data

[0291] Processing: Keyword detection, location acquisition, emergency response arrangements, emergency contact

[0292] Output: Arrange emergency response, notify emergency contact

[0293] Step 6:

[0294] Education and Training: When a user wishes to learn or receive training, the user makes a request through the device, which is then sent to the server, which selects the appropriate educational material or training program and provides it through the device.

[0295] Input: Learning or training request

[0296] Processing: Analyzing your request and selecting educational materials and programs

[0297] Output: Providing educational materials and training programs

[0298] For example, if a user says to the device, "Tell me what lunch you recommend," the device converts the speech into text and sends it to the server. The server then takes into account the user's past eating history and health information to suggest the optimal lunch. The prompt to the generative AI model is as follows:

[0299] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[0300] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0301] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and an emotion engine, emergency response, and education and training. This system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[0302] 1. User Authentication

[0303] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[0304] 2. Managing Personalized Data

[0305] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[0306] 3. Daily conversation and problem consultation

[0307] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or concern. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and even the user's emotional state using an emotion engine. This answer is returned to the user via voice or text via the device. For example, if the user says, "I've been feeling tired lately," the device will not only suggest ways to relax and rest, but if the emotion engine detects that the user is feeling stressed, it will also provide measures that are particularly useful for reducing stress.

[0308] 4. Emergency Response

[0309] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends in an emergency. For example, if a user yells "My chest hurts, help!", the server will call an ambulance based on the user's location information and notify their family. The emotion engine can adjust the priority of emergency response according to the user's emotional state.

[0310] 5. Education and Training Functions

[0311] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," the server provides a basic pronunciation practice program and begins instruction. The emotion engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[0312] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[0313] The processing flow will be explained below.

[0314] 1. User Authentication

[0315] Step 1:

[0316] The user enters their username and password into the login screen.

[0317] Step 2:

[0318] The terminal transmits the entered authentication information to the server.

[0319] Step 3:

[0320] The server compares the received authentication information with the database and generates an authentication result.

[0321] Step 4:

[0322] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[0323] Step 5:

[0324] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[0325] 2. Managing Personalized Data

[0326] Step 1:

[0327] The terminal displays a personalized data entry screen when you log in for the first time.

[0328] Step 2:

[0329] Users enter personal data such as name, age, health information, hobbies, etc.

[0330] Step 3:

[0331] The terminal transmits the input data to the server.

[0332] Step 4:

[0333] The server stores the received data in a database.

[0334] 3. Daily conversation and problem consultation

[0335] Step 1:

[0336] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[0337] Step 2:

[0338] The device converts the user's voice into text and sends it to the server.

[0339] Step 3:

[0340] The server analyzes the received text using a natural language processing engine.

[0341] Step 4:

[0342] Based on the analysis results, the server analyzes the user's past data, current situation, and even the user's emotional state using an emotion engine.

[0343] Step 5:

[0344] The server generates the best answer and sends it to the device.

[0345] Step 6:

[0346] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[0347] 4. Emotional engine response

[0348] Step 1:

[0349] The device collects voice and text data when the user speaks.

[0350] Step 2:

[0351] The device sends voice data and text data to the server.

[0352] Step 3:

[0353] The server analyzes the voice and text data using an emotion engine.

[0354] Step 4:

[0355] The server stores the user's emotional state as an analysis result and reflects it in generating answers to everyday conversations and problem consultations.

[0356] Step 5:

[0357] The device will provide appropriate feedback and suggestions to the user based on their emotional state. For example, if the user is feeling stressed, it will suggest, "Why don't you take a short walk?"

[0358] 5. Emergency Response

[0359] Step 1:

[0360] The user utters an emergency keyword such as "Help!"

[0361] Step 2:

[0362] The device detects the emergency keyword and immediately sends an alert to the server.

[0363] Step 3:

[0364] The server collects the user's location and health data.

[0365] Step 4:

[0366] The server automatically contacts local emergency services.

[0367] Step 5:

[0368] The server will also contact the user's family and friends in an emergency.

[0369] Step 6:

[0370] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[0371] 6. Education and Training Functions

[0372] Step 1:

[0373] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[0374] Step 2:

[0375] The device sends a request to the server.

[0376] Step 3:

[0377] The server selects appropriate educational materials and training programs.

[0378] Step 4:

[0379] The server transmits the selected teaching materials and programs to the terminal.

[0380] Step 5:

[0381] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[0382] Step 6:

[0383] The server uses an emotion engine to analyze the user's emotional state and motivation for learning.

[0384] Step 7:

[0385] The device will suggest approaches based on the user's motivation and optimize the learning effect. For example, it will suggest specific steps such as, "Your English pronunciation is going well, so next let's practice using example sentences."

[0386] This allows users to receive not only everyday support but also individualized assistance based on their emotional state. By combining this with the emotion engine, even more personalized services can be realized.

[0387] Example 2

[0388] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0389] Conventional AI assistant systems often have separate functions for user authentication, personalized data management, dialogue using natural language processing and sentiment analysis technology, emergency response, and education and training, resulting in a lack of a comprehensive, integrated system. Furthermore, they lack the ability to provide personalized services that take the user's emotional state into account, making it difficult to provide detailed responses based on real-time sentiment analysis of the user.

[0390] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text and analyzing it with a natural language processing engine to generate an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for providing appropriate learning materials and training programs in accordance with the user's learning and training requirements, means for analyzing the user's emotional state using an emotion analysis engine and providing personalized services, means for detecting emergency keywords and immediately sending an alert to the server, means for automatically arranging local emergency response in an emergency and making emergency contact with the user's family and friends, and means for recording the user's progress and providing appropriate feedback. This enables comprehensive support and the provision of meticulous services tailored to the user's emotional state.

[0391] "Authentication Information" means information used to verify a user's identity, such as a username and password.

[0392] A "database" is a system for efficiently storing, managing, and retrieving data.

[0393] A "session" refers to the period during which a user is logged into a system and performs continuous operations.

[0394] "Personalized data" is information specific to each user, including name, age, health information, hobbies, etc.

[0395] A "natural language processing engine" is a general term for algorithms and software that understand, analyze, and generate responses to human language.

[0396] "Emotion analysis engine" is a general term for algorithms and software that analyze a user's emotional state from their statements and actions.

[0397] "Emergency keywords" are specific words or phrases that indicate an emergency, such as "Help!"

[0398] "Location information" refers to information that indicates the user's current location, including GPS data.

[0399] "Health data" refers to information about the user's health status, such as heart rate and blood pressure.

[0400] "Emergency response" refers to a series of procedures and actions to respond to an emergency, including, for example, arranging for an ambulance.

[0401] "Instructional Materials" means educational materials and content provided for User learning or training.

[0402] A "Training Program" is a planned sequence of study or practice designed to improve a User's specific skills or knowledge.

[0403] "Feedback" refers to evaluation and guidance of a user's behavior and performance, providing information that will help them improve their next actions.

[0404] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and emotion analysis technology, emergency response, and education and training. The system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[0405] Hardware and software used

[0406] The system uses the following key technologies:

[0407] Server: Equipped with database management, natural language processing engine, and sentiment analysis engine.

[0408] Database Management: MySQL, PostgreSQL

[0409] Natural Language Processing: OpenAI GPT

[0410] Sentiment analysis engine: IBM Watson Tone Analyzer

[0411] Terminal: Equipped with user interface, speech recognition and text generation functions.

[0412] User Interface: HTML, CSS, JavaScript

[0413] Speech Recognition: Google's Speech-to-Text

[0414] Text generation: OpenAI GPT

[0415] System Overview

[0416] 1. User Authentication

[0417] When a user logs in for the first time, they enter their username and password. This information is sent via the terminal to the server. The server checks the authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If it fails, it returns an error message instructing the user to enter their authentication information again.

[0418] 2. Managing Personalized Data

[0419] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters personal information such as name, age, health information, and hobbies. This data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[0420] 3. Daily conversation and problem consultation

[0421] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and emotional state using an emotion analysis engine. This answer is returned to the user via the device in voice or text.

[0422] 4. Emergency Response

[0423] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then collect the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends. The emotion analysis engine can adjust the priority of emergency response depending on the user's emotional state.

[0424] 5. Education and Training Functions

[0425] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. The emotion analysis engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[0426] Specific examples

[0427] 1. Example of a user authentication prompt

[0428] Username: example_user

[0429] Password: example_pass

[0430] 2. Examples of prompts for personalized data entry

[0431] Name: Taro

[0432] Age: 25

[0433] Health information: Good

[0434] Hobbies: Reading

[0435] 3. Examples of prompts for everyday conversation and problem-solving

[0436] I've been feeling tired lately

[0437] 4. Emergency Response Prompt Examples

[0438] help me!

[0439] 5. Examples of Prompt Sentences for Education and Training Functions

[0440] I want to practice my English pronunciation

[0441] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion analysis engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[0442] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0443] Step 1:

[0444] A user opens an application and enters their username and password on the login screen. The entered credentials are encrypted by the device and sent to the server, which becomes the input.

[0445] Step 2:

[0446] The server compares the received authentication information with the database (MySQL or PostgreSQL). The result of the comparison is determined to be either success or failure, and in either case the result is returned to the terminal in JSON format (output). This is data comparison and calculation.

[0447] Step 3:

[0448] The device receives the authentication result from the server, and displays the home screen if authentication is successful, or an error message if it is unsuccessful, allowing the user to perform the next operation.

[0449] Step 4:

[0450] When logging in for the first time, the device will display a personalized data entry screen, where the user will enter their name, age, health information, and hobbies. This will be their new input.

[0451] Step 5:

[0452] The device encrypts the entered personalized data and sends it to the server. The server receives the data and stores it in the database. After storage is complete, it sends a success message in JSON format to the device (output). This is the data storage operation.

[0453] Step 6:

[0454] The device receives the success message from the server and displays "Data saving completed" to the user. This step allows the user to confirm that the data was saved correctly.

[0455] Step 7:

[0456] When a user speaks (e.g., "I'm feeling tired these days"), the device converts the speech into text using Google's Speech-to-Text, which becomes the new input.

[0457] Step 8:

[0458] The device sends the converted text to the server. The server analyzes the text using OpenAI GPT to understand the user's intent. It also analyzes the emotional state using IBM Watson Tone Analyzer. Based on the analysis results, it generates an appropriate answer and sends it to the device in JSON format (output). This is how the data analysis and generative AI model works.

[0459] Step 9:

[0460] The device synthesizes the answer from the server as voice and responds to the user (e.g., "Try deep breathing and light exercise to relax"), allowing the user to receive specific advice.

[0461] Step 10:

[0462] If a user shouts "Help!" in an emergency, the device will detect the emergency keyword using the voice recognition system, which will become the new input.

[0463] Step 11:

[0464] The device generates an emergency alert and sends it to the server. The server collects the user's location and health data, automatically dispatches local emergency response, and also contacts family and friends in an emergency. This is data collection and dispatch of emergency response.

[0465] Step 12:

[0466] The progress status from the server is sent to the terminal in real time, and the terminal displays to the user, "The ambulance is scheduled to arrive. Please wait calmly." This allows the user to understand the progress of the emergency response.

[0467] Step 13:

[0468] If a user requests, "I want to practice my English pronunciation," the device sends that request to the server, which becomes the new input.

[0469] Step 14:

[0470] The server selects appropriate teaching materials and training programs and sends them to the terminal. The terminal displays them and instructs the user to "start practicing basic English pronunciation." This is the process of providing data and displaying the user interface.

[0471] Step 15:

[0472] The user follows the instructions to practice pronunciation. The device records the progress and sends it to the server. The server analyzes the progress and generates appropriate feedback and sends it to the device.

[0473] Step 16:

[0474] The device will display feedback to the user saying, "Your pronunciation is improving. Now let's practice the 'th' sound." This allows the user to effectively progress with their learning / training.

[0475] (Application example 2)

[0476] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0477] Security is becoming increasingly important in modern homes and offices. However, existing security systems are unable to provide users with adequately personalized measures for emergency response and daily security needs. Furthermore, it is difficult to take immediate and appropriate action in the event of an emergency, resulting in insufficient systems to ensure user safety. Furthermore, limited opportunities for security education and training mean that users lack support for effective self-learning. To address these challenges, an advanced AI assistant system that combines user authentication, personalized data management, natural language processing, and an emotion engine is needed.

[0478] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for immediately sending a notification to the user's family or emergency contacts, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for detecting suspicious individuals and comparing abnormal behavior with a learning database based on voice recognition to evaluate safety, means for automatically notifying a security company based on the safety evaluation results, means for using a generative AI model to generate security-related educational content, and means for creating prompt sentences for the generative AI model. This improves the user's emergency response capabilities and enables personalized responses to everyday security needs. It will also enable effective self-learning and training on security, which is expected to improve overall security.

[0479] "User authentication" is the process of receiving a user's authentication information via a terminal and verifying it against a database.

[0480] "Personalized Data" means individual information about a user, such as the user's name, age, health information, and hobbies, that is used to personalize the service.

[0481] A "natural language processing engine" is a software engine that analyzes a user's voice and text, understands their meaning, and generates appropriate answers.

[0482] The "emotion engine" is an engine that analyzes the user's emotional state and responds appropriately based on that state.

[0483] "Emergency response" is the process of obtaining location and health data and arranging the necessary response when a user is in an emergency.

[0484] A "generative AI model" is an artificial intelligence model that generates content and answers based on user requests and situations.

[0485] A "security assistant" is an application that provides user authentication, personalized data management, natural language processing, emotion engine, emergency response, and education and training capabilities to enhance home and office security.

[0486] "Speech recognition" is a technology that converts a user's voice into text and analyzes it.

[0487] "Suspicious Person Detection" is the process of using voice recognition and other sensors to identify suspicious behavior or people.

[0488] "Automatic notification to security companies" is a process in which the system automatically sends notifications to security companies when a suspicious person or emergency occurs.

[0489] "Educational content generation" is the process of using generative AI models to create appropriate educational content based on users' learning and training requirements.

[0490] "Prompt generation" is the process of creating textual instructions to input to a generative AI model.

[0491] The present invention provides a security assistant for improving home and office security using an advanced AI assistant system. Specific embodiments for carrying out the invention are described below.

[0492] 1. User Authentication

[0493] This system uses facial recognition using a smartphone camera for user authentication. The device acquires the user's facial image and performs facial recognition using OpenCV. If authentication is successful, a session is started and the user can access the application. If authentication fails, an error message is displayed.

[0494] 2. Managing Personalized Data

[0495] When logging in for the first time, the device displays a personalized data entry screen for the user, asking them to enter their name, age, health information, hobbies, etc. The entered data is then stored in an SQLite database via the device, allowing the device to provide services tailored to the user's needs and preferences.

[0496] 3. Daily conversation and problem consultation

[0497] When a user speaks to the device, the device converts the speech into text and analyzes the text using the Google Cloud Natural Language API. Based on the analysis results, the server takes into account the user's past data and current situation and generates an appropriate answer using a natural language processing engine. This answer is returned to the user via voice or text.

[0498] 4. Emergency Response

[0499] If a user experiences an emergency, the emergency response process will begin when emergency keywords such as "help" are detected through voice recognition. The server will then obtain the user's location and health data and notify security companies and emergency contacts using communication APIs such as Twilio. Notifications will also be sent simultaneously to the user's family and friends.

[0500] 5. Education and Training Functions

[0501] When a user requests security training or learning, the server uses a generative AI model (e.g., GPT-4) to generate prompts and provide appropriate learning materials and training programs to the device, allowing the user to efficiently self-study and train their skills.

[0502] Examples of concrete examples and prompts

[0503] For example, if a user is in an emergency situation, say, "There might be a suspicious person around! Help!" in the middle of the night, the device will recognize this voice and notify the server. The server will then obtain the user's location and immediately notify the security company automatically, while also sending a message to the user's family.

[0504] Also, here are some examples of learning and training prompts:

[0505] User: "I want to practice my English pronunciation."

[0506] AI: "Of course! I'll guide you. Let's start with the basic vowels. Say 'a'."

[0507] In this way, the system uses advanced AI technology to provide users with personalized security services.

[0508] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0509] Step 1:

[0510] The user starts up the device and performs facial authentication. The device uses the smartphone camera to capture the user's facial image and performs facial recognition using OpenCV. If facial authentication is successful, the device sends authentication information to the server and starts a session. If authentication fails, the device displays an error message.

[0511] Input: A face image taken from the device's camera

[0512] Output: Authentication successful or error message

[0513] Step 2:

[0514] When a user logs in for the first time, they enter their personalization data. The device displays a screen for entering their name, age, health information, hobbies, etc., and the user enters the information. The entered data is sent to the server via the device and stored in an SQLite database.

[0515] Input: Personalization data entered by the user

[0516] Output: Personalization data stored in a database

[0517] Step 3:

[0518] When a user speaks to the device, the device converts the speech into text. The device uses a microphone to capture the speech and converts it into text using the Google Cloud Natural Language API. The server receives this text, analyzes it with a natural language processing engine, and generates an appropriate response. The generated response is returned to the user via the device as voice or text.

[0519] Input: User's voice

[0520] Output: Voice and text responses

[0521] Step 4:

[0522] In an emergency, if the user utters an emergency keyword such as "help," the device recognizes the voice and initiates the emergency response process. The device analyzes the voice and, if it detects an emergency keyword, sends a notification to the server. The server obtains the user's location and health data and uses communication APIs such as Twilio to send notifications to security companies and emergency contacts. Notifications are also sent to the user's family at the same time.

[0523] Input: User's emergency voice

[0524] Output: Send emergency notification

[0525] Step 5:

[0526] When a user requests security education or training, the server uses a generative AI model to generate prompts and generate appropriate learning materials and training programs. The generated content is then provided to the user via their device, allowing the user to efficiently self-study and train their skills.

[0527] Input: User's learning request

[0528] Output: Teaching materials and training programs

[0529] Step 6:

[0530] When using voice recognition to detect suspicious individuals or recognize abnormal behavior, the device captures the voice and sends it to the server. The server analyzes the voice and compares it with a learning database to evaluate safety. Based on the evaluation results, it automatically notifies the security company.

[0531] Input: Audio data

[0532] Output: Safety assessment and notification to security company

[0533] Step 7:

[0534] When the user logs in again, the device will perform facial recognition again. If the authentication is successful, the device will provide the user with optimized services based on the personalization data entered last time.

[0535] Input: A face image taken from the device's camera

[0536] Output: Service provided after successful authentication

[0537] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0538] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0539] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0540] [Second embodiment]

[0541] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0542] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0543] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0544] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0545] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0546] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0547] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0548] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0549] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0550] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0551] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0552] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0553] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[0554] 1. User Authentication

[0555] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[0556] 2. Managing Personalized Data

[0557] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[0558] 3. Daily conversation and problem consultation

[0559] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text. For example, if the user says, "I've been feeling tired from work lately," the device will suggest ways to relax and take a break.

[0560] 4. Emergency Response

[0561] If a user experiences an emergency, the device will detect this by uttering an emergency keyword such as "Help!" and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also provide support by contacting the user's family and friends in an emergency. For example, if a user yells "My chest hurts! Help!", the server will use the user's location information to call an ambulance and notify their family.

[0562] 5. Education and Training Functions

[0563] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," a basic pronunciation practice program is provided and instruction begins.

[0564] As described above, this system can provide comprehensive support to users throughout their lives.

[0565] The processing flow will be explained below.

[0566] 1. User Authentication

[0567] Step 1:

[0568] The user enters their username and password into the login screen.

[0569] Step 2:

[0570] The terminal transmits the entered authentication information to the server.

[0571] Step 3:

[0572] The server compares the received authentication information with the database and generates an authentication result.

[0573] Step 4:

[0574] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[0575] Step 5:

[0576] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[0577] 2. Managing Personalized Data

[0578] Step 1:

[0579] The terminal displays a personalized data entry screen when you log in for the first time.

[0580] Step 2:

[0581] Users enter personal data such as name, age, health information, and hobbies.

[0582] Step 3:

[0583] The terminal transmits the input data to the server.

[0584] Step 4:

[0585] The server stores the received data in a database.

[0586] 3. Daily conversation and problem consultation

[0587] Step 1:

[0588] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[0589] Step 2:

[0590] The device converts the user's voice into text and sends it to the server.

[0591] Step 3:

[0592] The server analyzes the received text using a natural language processing engine.

[0593] Step 4:

[0594] The server generates an appropriate answer based on the user's past data and current situation.

[0595] Step 5:

[0596] The server generates a response and sends it to the terminal.

[0597] Step 6:

[0598] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[0599] 4. Emergency Response

[0600] Step 1:

[0601] The user utters an emergency keyword such as "Help!"

[0602] Step 2:

[0603] The device detects the emergency keyword and immediately sends an alert to the server.

[0604] Step 3:

[0605] The server collects the user's location and health data.

[0606] Step 4:

[0607] The server automatically contacts local emergency services.

[0608] Step 5:

[0609] The server will also contact the user's family and friends in an emergency.

[0610] Step 6:

[0611] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[0612] 5. Education and Training Functions

[0613] Step 1:

[0614] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[0615] Step 2:

[0616] The device sends a request to the server.

[0617] Step 3:

[0618] The server selects appropriate educational materials and training programs.

[0619] Step 4:

[0620] The server transmits the selected teaching materials and programs to the terminal.

[0621] Step 5:

[0622] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[0623] Example 1

[0624] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0625] Conventional AI assistant systems have had difficulty responding appropriately to diverse user needs. In particular, they lacked the ability to respond quickly in emergencies, manage personalized data based on individual user needs, and provide advanced dialogue functions that combine voice recognition and voice synthesis technologies. As a result, user convenience and safety were not sufficiently ensured.

[0626] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0627] In this invention, the server includes means for receiving user authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's speech into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as needed, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for analyzing prompt sentences based on the user's diverse needs and requests using a generative AI model and generating responses based thereon, means for analyzing the user's speech and responding vocally using speech recognition and synthesis technology, and means for detecting emergency keywords and sending immediate alerts. This allows for a wide range of support to be provided to users, enabling rapid response in emergencies and the provision of services tailored to individual needs.

[0628] "User authentication" is the process of verifying a user's identity when accessing a system, usually through a username and password.

[0629] A "database" is an electronic system for managing, storing, retrieving, and updating data efficiently and effectively.

[0630] A "session" refers to a series of operations and communications between when a user logs in to the system and when they log out.

[0631] "Personalized Data" refers to information about a user (e.g., name, age, health information, hobbies, etc.) that the system uses to provide the user with the most appropriate service.

[0632] A "natural language processing engine" refers to an algorithm or model for understanding and processing human language, and is used, for example, to analyze the intent and content of a conversation.

[0633] "Emergency response" refers to the process of responding quickly to a user's emergency situation and arranging for the necessary emergency services.

[0634] A "generative AI model" refers to an artificial intelligence algorithm that automatically generates appropriate responses or content based on user input.

[0635] A "prompt" refers to guided text used to elicit an appropriate response from a generative AI model.

[0636] "Speech recognition" refers to the technology that analyzes a user's speech and converts it into text data.

[0637] "Speech synthesis" refers to the technology of converting text data into voice data and providing information to users via voice.

[0638] "Emergency keywords" are specific words or phrases that the system uses to determine an emergency, and when detected, an emergency response is initiated.

[0639] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[0640] User authentication

[0641] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database (e.g., MySQL or PostgreSQL), and if authentication is successful, it starts a session and returns a login success message to the user. If authentication fails, it returns an error message instructing the user to enter their authentication information again.

[0642] Managing Personalization Data

[0643] When the user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database (e.g., SQLite or MongoDB). This makes it possible to provide services tailored to the user's needs and preferences.

[0644] Daily conversation and problem consultation

[0645] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine (such as GPT-3 or BERT) to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text.

[0646] For example, if a user says, "I've been feeling tired at work lately," the system will suggest ways to relax and take a break. The prompts in this case could be:

[0647] Generate an error message if a user fails to log in.

[0648] "Provide health advice if a user enters that running is a hobby."

[0649] "Generate relaxation suggestions for users who have recently been feeling tired from work."

[0650] Emergency response

[0651] If a user experiences an emergency, they can simply say "Help!" or another emergency keyword. The device will detect this and immediately send an alert to the server. The server will then collect the user's location and health data, and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends to provide support.

[0652] For example, if a user cries out, "My chest hurts, help me!", the server will call an ambulance based on the user's location and notify their family.

[0653] Education and Training Features

[0654] When a user wishes to learn or train, for example by requesting "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows users to efficiently self-study and train their skills.

[0655] As a concrete example, if a user requests, "I want to practice my English pronunciation," we will provide them with a basic pronunciation practice program and begin providing instruction.

[0656] As described above, this system can provide comprehensive support to users throughout their lives.

[0657] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0658] Step 1:

[0659] The user enters their username and password on the terminal, which is the authentication input.

[0660] Input: Username and Password

[0661] How it works: A user enters their username and password into the device's login screen.

[0662] Output: The terminal stores the entered information and prepares it for transmission to the next step.

[0663] Step 2:

[0664] The terminal transmits the entered authentication information to the server.

[0665] Input: The username and password entered by the user

[0666] What happens: The device sends authentication information to the server as an HTTPS request (for example, URL: https: / / example.com / api / login).

[0667] Output: An HTTPS request containing authentication information is sent to the server.

[0668] Step 3:

[0669] The server checks the received authentication information against a database.

[0670] Input: Authentication information sent from the device (username and password)

[0671] What it does: The server compares the authentication information with the user information in a database (e.g. MySQL) to see if they match.

[0672] Output: Authentication result (success or failure)

[0673] Step 4:

[0674] The server returns the authentication result to the terminal.

[0675] Input: Authentication result (success or failure)

[0676] Operation: The server sends the authentication result to the terminal as an HTTPS response.

[0677] Output: The terminal receives the authentication result.

[0678] Step 5:

[0679] The device displays the authentication result to the user.

[0680] Input: Authentication result received from the server (success or failure)

[0681] Operation: If successful, the device will transition to the next screen. If unsuccessful, an error message will be displayed saying "Username or password is incorrect."

[0682] Output: The authentication result is displayed to the user.

[0683] Step 6:

[0684] The terminal displays a personalization data entry screen to the user.

[0685] Input: User successfully logged in information

[0686] Action: The device renders and displays the personalization data entry screen.

[0687] Output: The user is presented with a data entry screen.

[0688] Step 7:

[0689] The user enters individual information.

[0690] Input: Personalized data input screen

[0691] How it works: User enters name, age, health information, hobbies, etc.

[0692] Output: Individual personalized data is input to the device.

[0693] Step 8:

[0694] The terminal transmits the input data to the server.

[0695] Input: Personalized data you enter (such as name, age, health information, hobbies, etc.)

[0696] What happens: The device sends this as an HTTPS request to the server (e.g., URL: https: / / example.com / api / personalize).

[0697] Output: An HTTPS request containing personalization data is sent to the server.

[0698] Step 9:

[0699] The server stores the data in a database.

[0700] Input: Personalization data sent from your device

[0701] How it works: The server stores data in a database (e.g., SQLite or MongoDB).

[0702] Output: A message confirming successful save is generated.

[0703] Step 10:

[0704] The server returns a message to the terminal confirming successful saving.

[0705] Input: Personalized data saving success information

[0706] How it works: The server sends a confirmation message to the device as an HTTPS response.

[0707] Output: You will receive a save confirmation message on your device.

[0708] Step 11:

[0709] The device displays a confirmation message to the user.

[0710] Input: Save confirmation message from the server

[0711] What it does: The device displays "Data saved" to the user.

[0712] Output: User receives confirmation to save data.

[0713] Step 12:

[0714] The user speaks into the device.

[0715] Input: Questions and inquiries by voice

[0716] Action: The user speaks into the device.

[0717] Output: The device receives the audio data.

[0718] Step 13:

[0719] The device converts the speech to text.

[0720] Input: Audio data

[0721] How it works: Your device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert speech to text.

[0722] Output: Text data is generated.

[0723] Step 14:

[0724] The terminal transmits the converted text data to the server.

[0725] Input: Text data

[0726] What happens: The device sends text data as an HTTPS request to the server (for example, URL: https: / / example.com / api / chat).

[0727] Output: An HTTPS request containing text data is sent to the server.

[0728] Step 15:

[0729] The server analyzes the text using a natural language processing engine.

[0730] Input: Text data sent from the terminal

[0731] How it works: The server uses a natural language processing engine (e.g., GPT-3) to parse the text and generate an appropriate answer.

[0732] Output: The answer text is generated.

[0733] Step 16:

[0734] The server generates a response text and sends it back to the terminal.

[0735] Input: Generated answer text

[0736] How it works: The server sends the answer text to the device as an HTTPS response.

[0737] Output: The answer text is sent to the terminal.

[0738] Step 17:

[0739] The device displays the answer to the user.

[0740] Input: Response text from the server

[0741] How it works: The device uses speech synthesis technology (e.g., Google Cloud Text-to-Speech API) to synthesize a response and provide it to the user in voice or text.

[0742] Output: The user receives the answer.

[0743] (Application example 1)

[0744] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0745] In modern society, users need personalized services and support based on their individual needs and circumstances. However, current systems struggle to properly manage detailed health information and data and provide optimal recommendations and support. Furthermore, there is a lack of systems that provide integrated support for multiple functions across all aspects of daily life, including diet, health management, and emergency response. Another challenge is providing a system that allows users to easily request and order food and that can respond quickly in emergencies.

[0746] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0747] In this invention, the server includes means for receiving a user's authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating appropriate responses, means for managing the user's dietary history and health information, means for accepting and analyzing the user's dietary requests and orders via voice or text, and providing personalized dietary suggestions and orders, means for acquiring the user's location information and health data in an emergency and arranging for emergency response as necessary, and means for providing appropriate educational materials and training programs in response to the user's learning and training requests. This allows users to receive appropriate dietary suggestions and orders based on their individual dietary requirements, and ensures rapid and appropriate responses in emergencies.

[0748] "User authentication" is the process of receiving a user's authentication information and verifying it against a database.

[0749] "Personalized Data" is information that is collected and stored about you based on your specific attributes and preferences.

[0750] "Natural language processing" is a technology for analyzing text, understanding its meaning, and generating appropriate answers.

[0751] "Diet history" is a record of meals a user has consumed in the past.

[0752] "Health information" refers to information that includes data related to health, such as the user's health condition and disease information.

[0753] "Meal suggestion" is the process of recommending appropriate meals based on the user's dietary history and health information.

[0754] "Speech recognition" is a technology that converts speech into text.

[0755] "Location information" is data that indicates the user's current location.

[0756] "Emergency response" is the process of providing users with the assistance and services they need in an emergency.

[0757] A "learning or training program" is a collection of materials and exercises provided to users to help them learn or improve their skills.

[0758] A system for implementing this invention is an advanced AI assistant system that includes user authentication, personalized data management, dialogue using natural language processing, emergency response, education and training functions, etc. Specific embodiments will be described in detail below.

[0759] User authentication

[0760] The server receives the user's authentication information and authenticates it by checking it against a database. Specifically, when the user logs in for the first time, they enter their username and password, which are sent to the server via their terminal. The server checks the received authentication information against its existing database, and if authentication is successful, it starts a session, or returns an error message if it fails.

[0761] Managing Personalization Data

[0762] The server displays a screen for the user to enter personalization data and stores the data entered by the user in a database. For example, the user enters personal information such as name, age, health information, hobbies, etc., which are then stored in the database. This allows the server to provide services based on the user's specific needs and preferences.

[0763] Dialogue using natural language processing

[0764] When a user speaks, the device converts the speech into text and sends it to the server. The server then analyzes the text using a natural language processing engine that uses a generative AI model to generate an appropriate answer to the user's question or request. The answer is then returned to the user via the device in voice or text.

[0765] Specific examples

[0766] For example, if a user says, "What's your lunch recommendation?", the server will suggest the perfect lunch based on the stored personalized data. Also, if a user requests, "I want to practice my English pronunciation," the server will provide appropriate learning materials.

[0767] Emergency response

[0768] In the event of an emergency, the server will obtain the user's location and health data and arrange for emergency response if necessary. For example, if the user yells, "My chest hurts! Help!", the device will detect this emergency keyword and immediately send an alert to the server. The server will then obtain the user's location, arrange for an ambulance, and notify their family.

[0769] Education and Training Features

[0770] When a user wishes to learn or train, the server selects appropriate learning materials and training programs and provides them via the terminal, allowing the user to efficiently self-study and train their skills.

[0771] Hardware and Software

[0772] The system uses the following hardware and software:

[0773] Hardware: Devices such as smartphones, smart glasses, and head-mounted displays

[0774] software:

[0775] Speech Recognition Library: speech_recognition

[0776] Natural language processing library: transformers

[0777] Database management library: SQLAlchemy

[0778] Generative AI model: "cl-tohoku / bert-base-japanese" as an example

[0779] Examples of prompt statements

[0780] For example, a prompt to a generative AI model might look like this:

[0781] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[0782] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0783] Step 1:

[0784] User authentication: When a user logs in for the first time, they enter their username and password. The device sends this information to the server. The server checks the received authentication information against an existing database and starts the session if authentication is successful. If authentication fails, it returns an error message to the device.

[0785] Input: Username, Password

[0786] Processing: Match against database

[0787] Output: Session started or error message

[0788] Step 2:

[0789] Input of personalized data: If authentication is successful, the device will display a personalized data input screen for the user. The user will input personal information such as name, age, health information, hobbies, etc. This data will be sent from the device to the server, where it will be stored in a database.

[0790] Input: Name, age, health information, hobbies

[0791] Processing: Save to database

[0792] Output: Personalized data

[0793] Step 3:

[0794] Speech recognition and natural language processing: The user speaks a question or request into the device. The device converts the speech into text and sends this text to the server. The server uses a natural language processing engine with a generative AI model to analyze the text and generate an appropriate response. This response is sent back to the device and returned to the user as voice or text.

[0795] Input: Voice input (questions and requests)

[0796] Processing: Speech-to-text conversion, text analysis, and generating appropriate answers

[0797] Output: Text or audio response

[0798] Step 4:

[0799] Management of dietary history and health information: When a user makes a dietary question or request, the server analyzes the stored personalized data (dietary history and health information) and generates appropriate dietary suggestions. The suggestions are notified to the user via the device.

[0800] Input: dietary requests, past dietary history, health information

[0801] Processing: Retrieving information from the database, analyzing it, and generating meal suggestions

[0802] Output: Meal suggestion notification

[0803] Step 5:

[0804] Emergency Response: When a user utters a keyword such as "Help!" in an emergency, the device detects this and immediately sends an alert to the server. The server obtains the user's location and health data and automatically arranges for local emergency response if necessary. The server also notifies the user's family and friends in an emergency.

[0805] Input: Emergency keywords, location information, health data

[0806] Processing: Keyword detection, location acquisition, emergency response arrangements, emergency contact

[0807] Output: Arrange emergency response, notify emergency contact

[0808] Step 6:

[0809] Education and Training: When a user wishes to learn or receive training, the user makes a request through the device, which is then sent to the server, which selects the appropriate educational material or training program and provides it through the device.

[0810] Input: Learning or training request

[0811] Processing: Analyzing your request and selecting educational materials and programs

[0812] Output: Providing educational materials and training programs

[0813] For example, if a user says to the device, "Tell me what lunch you recommend," the device converts the speech into text and sends it to the server. The server then takes into account the user's past eating history and health information to suggest the optimal lunch. The prompt to the generative AI model is as follows:

[0814] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[0815] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0816] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and an emotion engine, emergency response, and education and training. This system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[0817] 1. User Authentication

[0818] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[0819] 2. Managing Personalized Data

[0820] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[0821] 3. Daily conversation and problem consultation

[0822] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or concern. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and even the user's emotional state using an emotion engine. This answer is returned to the user via voice or text via the device. For example, if the user says, "I've been feeling tired lately," the device will not only suggest ways to relax and rest, but if the emotion engine detects that the user is feeling stressed, it will also provide measures that are particularly useful for reducing stress.

[0823] 4. Emergency Response

[0824] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends in an emergency. For example, if a user yells "My chest hurts, help!", the server will call an ambulance based on the user's location information and notify their family. The emotion engine can adjust the priority of emergency response according to the user's emotional state.

[0825] 5. Education and Training Functions

[0826] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," the server provides a basic pronunciation practice program and begins instruction. The emotion engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[0827] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[0828] The processing flow will be explained below.

[0829] 1. User Authentication

[0830] Step 1:

[0831] The user enters their username and password into the login screen.

[0832] Step 2:

[0833] The terminal transmits the entered authentication information to the server.

[0834] Step 3:

[0835] The server compares the received authentication information with the database and generates an authentication result.

[0836] Step 4:

[0837] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[0838] Step 5:

[0839] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[0840] 2. Managing Personalized Data

[0841] Step 1:

[0842] The terminal displays a personalized data entry screen when you log in for the first time.

[0843] Step 2:

[0844] Users enter personal data such as name, age, health information, hobbies, etc.

[0845] Step 3:

[0846] The terminal transmits the input data to the server.

[0847] Step 4:

[0848] The server stores the received data in a database.

[0849] 3. Daily conversation and problem consultation

[0850] Step 1:

[0851] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[0852] Step 2:

[0853] The device converts the user's voice into text and sends it to the server.

[0854] Step 3:

[0855] The server analyzes the received text using a natural language processing engine.

[0856] Step 4:

[0857] Based on the analysis results, the server analyzes the user's past data, current situation, and even the user's emotional state using an emotion engine.

[0858] Step 5:

[0859] The server generates the best answer and sends it to the device.

[0860] Step 6:

[0861] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[0862] 4. Emotional engine response

[0863] Step 1:

[0864] The device collects voice and text data when the user speaks.

[0865] Step 2:

[0866] The device sends voice data and text data to the server.

[0867] Step 3:

[0868] The server analyzes the voice and text data using an emotion engine.

[0869] Step 4:

[0870] The server stores the user's emotional state as an analysis result and reflects it in generating answers to everyday conversations and problem consultations.

[0871] Step 5:

[0872] The device will provide appropriate feedback and suggestions to the user based on their emotional state. For example, if the user is feeling stressed, it will suggest, "Why don't you take a short walk?"

[0873] 5. Emergency Response

[0874] Step 1:

[0875] The user utters an emergency keyword such as "Help!"

[0876] Step 2:

[0877] The device detects the emergency keyword and immediately sends an alert to the server.

[0878] Step 3:

[0879] The server collects the user's location and health data.

[0880] Step 4:

[0881] The server automatically contacts local emergency services.

[0882] Step 5:

[0883] The server will also contact the user's family and friends in an emergency.

[0884] Step 6:

[0885] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[0886] 6. Education and Training Functions

[0887] Step 1:

[0888] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[0889] Step 2:

[0890] The device sends a request to the server.

[0891] Step 3:

[0892] The server selects appropriate educational materials and training programs.

[0893] Step 4:

[0894] The server transmits the selected teaching materials and programs to the terminal.

[0895] Step 5:

[0896] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[0897] Step 6:

[0898] The server uses an emotion engine to analyze the user's emotional state and motivation for learning.

[0899] Step 7:

[0900] The device will suggest approaches based on the user's motivation and optimize the learning effect. For example, it will suggest specific steps such as, "Your English pronunciation is going well, so next let's practice using example sentences."

[0901] This allows users to receive not only everyday support but also individualized assistance based on their emotional state. By combining this with the emotion engine, even more personalized services can be realized.

[0902] Example 2

[0903] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0904] Conventional AI assistant systems often have separate functions for user authentication, personalized data management, dialogue using natural language processing and sentiment analysis technology, emergency response, and education and training, resulting in a lack of a comprehensive, integrated system. Furthermore, they lack the ability to provide personalized services that take the user's emotional state into account, making it difficult to provide detailed responses based on real-time sentiment analysis of the user.

[0905] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text and analyzing it with a natural language processing engine to generate an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for providing appropriate learning materials and training programs in accordance with the user's learning and training requirements, means for analyzing the user's emotional state using an emotion analysis engine and providing personalized services, means for detecting emergency keywords and immediately sending an alert to the server, means for automatically arranging local emergency response in an emergency and making emergency contact with the user's family and friends, and means for recording the user's progress and providing appropriate feedback. This enables comprehensive support and the provision of meticulous services tailored to the user's emotional state.

[0906] "Authentication Information" means information used to verify a user's identity, such as a username and password.

[0907] A "database" is a system for efficiently storing, managing, and retrieving data.

[0908] A "session" refers to the period during which a user is logged into a system and performs continuous operations.

[0909] "Personalized data" is information specific to each user, including name, age, health information, hobbies, etc.

[0910] A "natural language processing engine" is a general term for algorithms and software that understand, analyze, and generate responses to human language.

[0911] "Emotion analysis engine" is a general term for algorithms and software that analyze a user's emotional state from their statements and actions.

[0912] "Emergency keywords" are specific words or phrases that indicate an emergency, such as "Help!"

[0913] "Location information" refers to information that indicates the user's current location, including GPS data.

[0914] "Health data" refers to information about the user's health status, such as heart rate and blood pressure.

[0915] "Emergency response" refers to a series of procedures and actions to respond to an emergency, including, for example, arranging for an ambulance.

[0916] "Instructional Materials" means educational materials and content provided for User learning or training.

[0917] A "Training Program" is a planned sequence of study or practice designed to improve a User's specific skills or knowledge.

[0918] "Feedback" refers to evaluation and guidance of a user's behavior and performance, providing information that will help them improve their next actions.

[0919] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and emotion analysis technology, emergency response, and education and training. The system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[0920] Hardware and software used

[0921] The system uses the following key technologies:

[0922] Server: Equipped with database management, natural language processing engine, and sentiment analysis engine.

[0923] Database Management: MySQL, PostgreSQL

[0924] Natural Language Processing: OpenAI GPT

[0925] Sentiment analysis engine: IBM Watson Tone Analyzer

[0926] Terminal: Equipped with user interface, speech recognition and text generation functions.

[0927] User Interface: HTML, CSS, JavaScript

[0928] Speech Recognition: Google's Speech-to-Text

[0929] Text generation: OpenAI GPT

[0930] System Overview

[0931] 1. User Authentication

[0932] When a user logs in for the first time, they enter their username and password. This information is sent via the terminal to the server. The server checks the authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If it fails, it returns an error message instructing the user to enter their authentication information again.

[0933] 2. Managing Personalized Data

[0934] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters personal information such as name, age, health information, and hobbies. This data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[0935] 3. Daily conversation and problem consultation

[0936] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and emotional state using an emotion analysis engine. This answer is returned to the user via the device in voice or text.

[0937] 4. Emergency Response

[0938] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then collect the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends. The emotion analysis engine can adjust the priority of emergency response depending on the user's emotional state.

[0939] 5. Education and Training Functions

[0940] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. The emotion analysis engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[0941] Specific examples

[0942] 1. Example of a user authentication prompt

[0943] Username: example_user

[0944] Password: example_pass

[0945] 2. Examples of prompts for personalized data entry

[0946] Name: Taro

[0947] Age: 25

[0948] Health information: Good

[0949] Hobbies: Reading

[0950] 3. Examples of prompts for everyday conversation and problem-solving

[0951] I've been feeling tired lately

[0952] 4. Emergency Response Prompt Examples

[0953] help me!

[0954] 5. Examples of Prompt Sentences for Education and Training Functions

[0955] I want to practice my English pronunciation

[0956] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion analysis engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[0957] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0958] Step 1:

[0959] A user opens an application and enters their username and password on the login screen. The entered credentials are encrypted by the device and sent to the server, which becomes the input.

[0960] Step 2:

[0961] The server compares the received authentication information with the database (MySQL or PostgreSQL). The result of the comparison is determined to be either success or failure, and in either case the result is returned to the terminal in JSON format (output). This is data comparison and calculation.

[0962] Step 3:

[0963] The device receives the authentication result from the server, and displays the home screen if authentication is successful, or an error message if it is unsuccessful, allowing the user to perform the next operation.

[0964] Step 4:

[0965] When logging in for the first time, the device will display a personalized data entry screen, where the user will enter their name, age, health information, and hobbies. This will be their new input.

[0966] Step 5:

[0967] The device encrypts the entered personalized data and sends it to the server. The server receives the data and stores it in the database. After storage is complete, it sends a success message in JSON format to the device (output). This is the data storage operation.

[0968] Step 6:

[0969] The device receives the success message from the server and displays "Data saving completed" to the user. This step allows the user to confirm that the data was saved correctly.

[0970] Step 7:

[0971] When a user speaks (e.g., "I'm feeling tired these days"), the device converts the speech into text using Google's Speech-to-Text, which becomes the new input.

[0972] Step 8:

[0973] The device sends the converted text to the server. The server analyzes the text using OpenAI GPT to understand the user's intent. It also analyzes the emotional state using IBM Watson Tone Analyzer. Based on the analysis results, it generates an appropriate answer and sends it to the device in JSON format (output). This is how the data analysis and generative AI model works.

[0974] Step 9:

[0975] The device synthesizes the answer from the server as voice and responds to the user (e.g., "Try deep breathing and light exercise to relax"), allowing the user to receive specific advice.

[0976] Step 10:

[0977] If a user shouts "Help!" in an emergency, the device will detect the emergency keyword using the voice recognition system, which will become the new input.

[0978] Step 11:

[0979] The device generates an emergency alert and sends it to the server. The server collects the user's location and health data, automatically dispatches local emergency response, and also contacts family and friends in an emergency. This is data collection and dispatch of emergency response.

[0980] Step 12:

[0981] The progress status from the server is sent to the terminal in real time, and the terminal displays to the user, "The ambulance is scheduled to arrive. Please wait calmly." This allows the user to understand the progress of the emergency response.

[0982] Step 13:

[0983] If a user requests, "I want to practice my English pronunciation," the device sends that request to the server, which becomes the new input.

[0984] Step 14:

[0985] The server selects appropriate teaching materials and training programs and sends them to the terminal. The terminal displays them and instructs the user to "start practicing basic English pronunciation." This is the process of providing data and displaying the user interface.

[0986] Step 15:

[0987] The user follows the instructions to practice pronunciation. The device records the progress and sends it to the server. The server analyzes the progress and generates appropriate feedback and sends it to the device.

[0988] Step 16:

[0989] The device will display feedback to the user saying, "Your pronunciation is improving. Now let's practice the 'th' sound." This allows the user to effectively progress with their learning / training.

[0990] (Application example 2)

[0991] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0992] Security is becoming increasingly important in modern homes and offices. However, existing security systems are unable to provide users with adequately personalized measures for emergency response and daily security needs. Furthermore, it is difficult to take immediate and appropriate action in the event of an emergency, resulting in insufficient systems to ensure user safety. Furthermore, limited opportunities for security education and training mean that users lack support for effective self-learning. To address these challenges, an advanced AI assistant system that combines user authentication, personalized data management, natural language processing, and an emotion engine is needed.

[0993] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for immediately sending a notification to the user's family or emergency contacts, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for detecting suspicious individuals and comparing abnormal behavior with a learning database based on voice recognition to evaluate safety, means for automatically notifying a security company based on the safety evaluation results, means for using a generative AI model to generate security-related educational content, and means for creating prompt sentences for the generative AI model. This improves the user's emergency response capabilities and enables personalized responses to everyday security needs. It will also enable effective self-learning and training on security, which is expected to improve overall security.

[0994] "User authentication" is the process of receiving a user's authentication information via a terminal and verifying it against a database.

[0995] "Personalized Data" means individual information about a user, such as the user's name, age, health information, and hobbies, that is used to personalize the service.

[0996] A "natural language processing engine" is a software engine that analyzes a user's voice and text, understands their meaning, and generates appropriate answers.

[0997] The "emotion engine" is an engine that analyzes the user's emotional state and responds appropriately based on that state.

[0998] "Emergency response" is the process of obtaining location and health data and arranging the necessary response when a user is in an emergency.

[0999] A "generative AI model" is an artificial intelligence model that generates content and answers based on user requests and situations.

[1000] A "security assistant" is an application that provides user authentication, personalized data management, natural language processing, emotion engine, emergency response, and education and training capabilities to enhance home and office security.

[1001] "Speech recognition" is a technology that converts a user's voice into text and analyzes it.

[1002] "Suspicious Person Detection" is the process of using voice recognition and other sensors to identify suspicious behavior or people.

[1003] "Automatic notification to security companies" is a process in which the system automatically sends notifications to security companies when a suspicious person or emergency occurs.

[1004] "Educational content generation" is the process of using generative AI models to create appropriate educational content based on users' learning and training requirements.

[1005] "Prompt generation" is the process of creating textual instructions to input to a generative AI model.

[1006] The present invention provides a security assistant for improving home and office security using an advanced AI assistant system. Specific embodiments for carrying out the invention are described below.

[1007] 1. User Authentication

[1008] This system uses facial recognition using a smartphone camera for user authentication. The device acquires the user's facial image and performs facial recognition using OpenCV. If authentication is successful, a session is started and the user can access the application. If authentication fails, an error message is displayed.

[1009] 2. Managing Personalized Data

[1010] When logging in for the first time, the device displays a personalized data entry screen for the user, asking them to enter their name, age, health information, hobbies, etc. The entered data is then stored in an SQLite database via the device, allowing the device to provide services tailored to the user's needs and preferences.

[1011] 3. Daily conversation and problem consultation

[1012] When a user speaks to the device, the device converts the speech into text and analyzes the text using the Google Cloud Natural Language API. Based on the analysis results, the server takes into account the user's past data and current situation and generates an appropriate answer using a natural language processing engine. This answer is returned to the user via voice or text.

[1013] 4. Emergency Response

[1014] If a user experiences an emergency, the emergency response process will begin when emergency keywords such as "help" are detected through voice recognition. The server will then obtain the user's location and health data and notify security companies and emergency contacts using communication APIs such as Twilio. Notifications will also be sent simultaneously to the user's family and friends.

[1015] 5. Education and Training Functions

[1016] When a user requests security training or learning, the server uses a generative AI model (e.g., GPT-4) to generate prompts and provide appropriate learning materials and training programs to the device, allowing the user to efficiently self-study and train their skills.

[1017] Examples of concrete examples and prompts

[1018] For example, if a user is in an emergency situation, say, "There might be a suspicious person around! Help!" in the middle of the night, the device will recognize this voice and notify the server. The server will then obtain the user's location and immediately notify the security company automatically, while also sending a message to the user's family.

[1019] Also, here are some examples of learning and training prompts:

[1020] User: "I want to practice my English pronunciation."

[1021] AI: "Of course! I'll guide you. Let's start with the basic vowels. Say 'a'."

[1022] In this way, the system uses advanced AI technology to provide users with personalized security services.

[1023] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1024] Step 1:

[1025] The user starts up the device and performs facial authentication. The device uses the smartphone camera to capture the user's facial image and performs facial recognition using OpenCV. If facial authentication is successful, the device sends authentication information to the server and starts a session. If authentication fails, the device displays an error message.

[1026] Input: A face image taken from the device's camera

[1027] Output: Authentication successful or error message

[1028] Step 2:

[1029] When a user logs in for the first time, they enter their personalization data. The device displays a screen for entering their name, age, health information, hobbies, etc., and the user enters the information. The entered data is sent to the server via the device and stored in an SQLite database.

[1030] Input: Personalization data entered by the user

[1031] Output: Personalization data stored in a database

[1032] Step 3:

[1033] When a user speaks to the device, the device converts the speech into text. The device uses a microphone to capture the speech and converts it into text using the Google Cloud Natural Language API. The server receives this text, analyzes it with a natural language processing engine, and generates an appropriate response. The generated response is returned to the user via the device as voice or text.

[1034] Input: User's voice

[1035] Output: Voice and text responses

[1036] Step 4:

[1037] In an emergency, if the user utters an emergency keyword such as "help," the device recognizes the voice and initiates the emergency response process. The device analyzes the voice and, if it detects an emergency keyword, sends a notification to the server. The server obtains the user's location and health data and uses communication APIs such as Twilio to send notifications to security companies and emergency contacts. Notifications are also sent to the user's family at the same time.

[1038] Input: User's emergency voice

[1039] Output: Send emergency notification

[1040] Step 5:

[1041] When a user requests security education or training, the server uses a generative AI model to generate prompts and generate appropriate learning materials and training programs. The generated content is then provided to the user via their device, allowing the user to efficiently self-study and train their skills.

[1042] Input: User's learning request

[1043] Output: Teaching materials and training programs

[1044] Step 6:

[1045] When using voice recognition to detect suspicious individuals or recognize abnormal behavior, the device captures the voice and sends it to the server. The server analyzes the voice and compares it with a learning database to evaluate safety. Based on the evaluation results, it automatically notifies the security company.

[1046] Input: Audio data

[1047] Output: Safety assessment and notification to security company

[1048] Step 7:

[1049] When the user logs in again, the device will perform facial recognition again. If the authentication is successful, the device will provide the user with optimized services based on the personalization data entered last time.

[1050] Input: A face image taken from the device's camera

[1051] Output: Service provided after successful authentication

[1052] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1053] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1054] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1055] [Third embodiment]

[1056] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1057] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1058] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1059] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1060] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1061] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1062] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1063] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1064] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1065] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1066] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1067] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1068] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[1069] 1. User Authentication

[1070] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[1071] 2. Managing Personalized Data

[1072] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[1073] 3. Daily conversation and problem consultation

[1074] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text. For example, if the user says, "I've been feeling tired from work lately," the device will suggest ways to relax and take a break.

[1075] 4. Emergency Response

[1076] If a user experiences an emergency, the device will detect this by uttering an emergency keyword such as "Help!" and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also provide support by contacting the user's family and friends in an emergency. For example, if a user yells "My chest hurts! Help!", the server will use the user's location information to call an ambulance and notify their family.

[1077] 5. Education and Training Functions

[1078] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," a basic pronunciation practice program is provided and instruction begins.

[1079] As described above, this system can provide comprehensive support to users throughout their lives.

[1080] The processing flow will be explained below.

[1081] 1. User Authentication

[1082] Step 1:

[1083] The user enters their username and password into the login screen.

[1084] Step 2:

[1085] The terminal transmits the entered authentication information to the server.

[1086] Step 3:

[1087] The server compares the received authentication information with the database and generates an authentication result.

[1088] Step 4:

[1089] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[1090] Step 5:

[1091] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[1092] 2. Managing Personalized Data

[1093] Step 1:

[1094] The terminal displays a personalized data entry screen when you log in for the first time.

[1095] Step 2:

[1096] Users enter personal data such as name, age, health information, and hobbies.

[1097] Step 3:

[1098] The terminal transmits the input data to the server.

[1099] Step 4:

[1100] The server stores the received data in a database.

[1101] 3. Daily conversation and problem consultation

[1102] Step 1:

[1103] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[1104] Step 2:

[1105] The device converts the user's voice into text and sends it to the server.

[1106] Step 3:

[1107] The server analyzes the received text using a natural language processing engine.

[1108] Step 4:

[1109] The server generates an appropriate answer based on the user's past data and current situation.

[1110] Step 5:

[1111] The server generates a response and sends it to the terminal.

[1112] Step 6:

[1113] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[1114] 4. Emergency Response

[1115] Step 1:

[1116] The user utters an emergency keyword such as "Help!"

[1117] Step 2:

[1118] The device detects the emergency keyword and immediately sends an alert to the server.

[1119] Step 3:

[1120] The server collects the user's location and health data.

[1121] Step 4:

[1122] The server automatically contacts local emergency services.

[1123] Step 5:

[1124] The server will also contact the user's family and friends in an emergency.

[1125] Step 6:

[1126] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[1127] 5. Education and Training Functions

[1128] Step 1:

[1129] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[1130] Step 2:

[1131] The device sends a request to the server.

[1132] Step 3:

[1133] The server selects appropriate educational materials and training programs.

[1134] Step 4:

[1135] The server transmits the selected teaching materials and programs to the terminal.

[1136] Step 5:

[1137] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[1138] Example 1

[1139] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1140] Conventional AI assistant systems have had difficulty responding appropriately to diverse user needs. In particular, they lacked the ability to respond quickly in emergencies, manage personalized data based on individual user needs, and provide advanced dialogue functions that combine voice recognition and voice synthesis technologies. As a result, user convenience and safety were not sufficiently ensured.

[1141] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1142] In this invention, the server includes means for receiving user authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's speech into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as needed, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for analyzing prompt sentences based on the user's diverse needs and requests using a generative AI model and generating responses based thereon, means for analyzing the user's speech and responding vocally using speech recognition and synthesis technology, and means for detecting emergency keywords and sending immediate alerts. This allows for a wide range of support to be provided to users, enabling rapid response in emergencies and the provision of services tailored to individual needs.

[1143] "User authentication" is the process of verifying a user's identity when accessing a system, usually through a username and password.

[1144] A "database" is an electronic system for managing, storing, retrieving, and updating data efficiently and effectively.

[1145] A "session" refers to a series of operations and communications between when a user logs in to the system and when they log out.

[1146] "Personalized Data" refers to information about a user (e.g., name, age, health information, hobbies, etc.) that the system uses to provide the user with the most appropriate service.

[1147] A "natural language processing engine" refers to an algorithm or model for understanding and processing human language, and is used, for example, to analyze the intent and content of a conversation.

[1148] "Emergency response" refers to the process of responding quickly to a user's emergency situation and arranging for the necessary emergency services.

[1149] A "generative AI model" refers to an artificial intelligence algorithm that automatically generates appropriate responses or content based on user input.

[1150] A "prompt" refers to guided text used to elicit an appropriate response from a generative AI model.

[1151] "Speech recognition" refers to the technology that analyzes a user's speech and converts it into text data.

[1152] "Speech synthesis" refers to the technology of converting text data into voice data and providing information to users via voice.

[1153] "Emergency keywords" are specific words or phrases that the system uses to determine an emergency, and when detected, an emergency response is initiated.

[1154] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[1155] User authentication

[1156] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database (e.g., MySQL or PostgreSQL), and if authentication is successful, it starts a session and returns a login success message to the user. If authentication fails, it returns an error message instructing the user to enter their authentication information again.

[1157] Managing Personalization Data

[1158] When the user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database (e.g., SQLite or MongoDB). This makes it possible to provide services tailored to the user's needs and preferences.

[1159] Daily conversation and problem consultation

[1160] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine (such as GPT-3 or BERT) to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text.

[1161] For example, if a user says, "I've been feeling tired at work lately," the system will suggest ways to relax and take a break. The prompts in this case could be:

[1162] Generate an error message if a user fails to log in.

[1163] "Provide health advice if a user enters that running is a hobby."

[1164] "Generate relaxation suggestions for users who have recently been feeling tired from work."

[1165] Emergency response

[1166] If a user experiences an emergency, they can simply say "Help!" or another emergency keyword. The device will detect this and immediately send an alert to the server. The server will then collect the user's location and health data, and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends to provide support.

[1167] For example, if a user cries out, "My chest hurts, help me!", the server will call an ambulance based on the user's location and notify their family.

[1168] Education and Training Features

[1169] When a user wishes to learn or train, for example by requesting "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows users to efficiently self-study and train their skills.

[1170] As a concrete example, if a user requests, "I want to practice my English pronunciation," we will provide them with a basic pronunciation practice program and begin providing instruction.

[1171] As described above, this system can provide comprehensive support to users throughout their lives.

[1172] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1173] Step 1:

[1174] The user enters their username and password on the terminal, which is the authentication input.

[1175] Input: Username and Password

[1176] How it works: A user enters their username and password into the device's login screen.

[1177] Output: The terminal stores the entered information and prepares it for transmission to the next step.

[1178] Step 2:

[1179] The terminal transmits the entered authentication information to the server.

[1180] Input: The username and password entered by the user

[1181] What happens: The device sends authentication information to the server as an HTTPS request (for example, URL: https: / / example.com / api / login).

[1182] Output: An HTTPS request containing authentication information is sent to the server.

[1183] Step 3:

[1184] The server checks the received authentication information against a database.

[1185] Input: Authentication information sent from the device (username and password)

[1186] What it does: The server compares the authentication information with the user information in a database (e.g. MySQL) to see if they match.

[1187] Output: Authentication result (success or failure)

[1188] Step 4:

[1189] The server returns the authentication result to the terminal.

[1190] Input: Authentication result (success or failure)

[1191] Operation: The server sends the authentication result to the terminal as an HTTPS response.

[1192] Output: The terminal receives the authentication result.

[1193] Step 5:

[1194] The device displays the authentication result to the user.

[1195] Input: Authentication result received from the server (success or failure)

[1196] Operation: If successful, the device will transition to the next screen. If unsuccessful, an error message will be displayed saying "Username or password is incorrect."

[1197] Output: The authentication result is displayed to the user.

[1198] Step 6:

[1199] The terminal displays a personalization data entry screen to the user.

[1200] Input: User successfully logged in information

[1201] Action: The device renders and displays the personalization data entry screen.

[1202] Output: The user is presented with a data entry screen.

[1203] Step 7:

[1204] The user enters individual information.

[1205] Input: Personalized data input screen

[1206] How it works: User enters name, age, health information, hobbies, etc.

[1207] Output: Individual personalized data is input to the device.

[1208] Step 8:

[1209] The terminal transmits the input data to the server.

[1210] Input: Personalized data you enter (such as name, age, health information, hobbies, etc.)

[1211] What happens: The device sends this as an HTTPS request to the server (e.g., URL: https: / / example.com / api / personalize).

[1212] Output: An HTTPS request containing personalization data is sent to the server.

[1213] Step 9:

[1214] The server stores the data in a database.

[1215] Input: Personalization data sent from your device

[1216] How it works: The server stores data in a database (e.g., SQLite or MongoDB).

[1217] Output: A message confirming successful save is generated.

[1218] Step 10:

[1219] The server returns a message to the terminal confirming successful saving.

[1220] Input: Personalized data saving success information

[1221] How it works: The server sends a confirmation message to the device as an HTTPS response.

[1222] Output: You will receive a save confirmation message on your device.

[1223] Step 11:

[1224] The device displays a confirmation message to the user.

[1225] Input: Save confirmation message from the server

[1226] What it does: The device displays "Data saved" to the user.

[1227] Output: User receives confirmation to save data.

[1228] Step 12:

[1229] The user speaks into the device.

[1230] Input: Questions and inquiries by voice

[1231] Action: The user speaks into the device.

[1232] Output: The device receives the audio data.

[1233] Step 13:

[1234] The device converts the speech to text.

[1235] Input: Audio data

[1236] How it works: Your device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert speech to text.

[1237] Output: Text data is generated.

[1238] Step 14:

[1239] The terminal transmits the converted text data to the server.

[1240] Input: Text data

[1241] What happens: The device sends text data as an HTTPS request to the server (for example, URL: https: / / example.com / api / chat).

[1242] Output: An HTTPS request containing text data is sent to the server.

[1243] Step 15:

[1244] The server analyzes the text using a natural language processing engine.

[1245] Input: Text data sent from the terminal

[1246] How it works: The server uses a natural language processing engine (e.g., GPT-3) to parse the text and generate an appropriate answer.

[1247] Output: The answer text is generated.

[1248] Step 16:

[1249] The server generates a response text and sends it back to the terminal.

[1250] Input: Generated answer text

[1251] How it works: The server sends the answer text to the device as an HTTPS response.

[1252] Output: The answer text is sent to the terminal.

[1253] Step 17:

[1254] The device displays the answer to the user.

[1255] Input: Response text from the server

[1256] How it works: The device uses speech synthesis technology (e.g., Google Cloud Text-to-Speech API) to synthesize a response and provide it to the user in voice or text.

[1257] Output: The user receives the answer.

[1258] (Application example 1)

[1259] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1260] In modern society, users need personalized services and support based on their individual needs and circumstances. However, current systems struggle to properly manage detailed health information and data and provide optimal recommendations and support. Furthermore, there is a lack of systems that provide integrated support for multiple functions across all aspects of daily life, including diet, health management, and emergency response. Another challenge is providing a system that allows users to easily request and order food and that can respond quickly in emergencies.

[1261] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1262] In this invention, the server includes means for receiving a user's authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating appropriate responses, means for managing the user's dietary history and health information, means for accepting and analyzing the user's dietary requests and orders via voice or text, and providing personalized dietary suggestions and orders, means for acquiring the user's location information and health data in an emergency and arranging for emergency response as necessary, and means for providing appropriate educational materials and training programs in response to the user's learning and training requests. This allows users to receive appropriate dietary suggestions and orders based on their individual dietary requirements, and ensures rapid and appropriate responses in emergencies.

[1263] "User authentication" is the process of receiving a user's authentication information and verifying it against a database.

[1264] "Personalized Data" is information that is collected and stored about you based on your specific attributes and preferences.

[1265] "Natural language processing" is a technology for analyzing text, understanding its meaning, and generating appropriate answers.

[1266] "Diet history" is a record of meals a user has consumed in the past.

[1267] "Health information" refers to information that includes data related to health, such as the user's health condition and disease information.

[1268] "Meal suggestion" is the process of recommending appropriate meals based on the user's dietary history and health information.

[1269] "Speech recognition" is a technology that converts speech into text.

[1270] "Location information" is data that indicates the user's current location.

[1271] "Emergency response" is the process of providing users with the assistance and services they need in an emergency.

[1272] A "learning or training program" is a collection of materials and exercises provided to users to help them learn or improve their skills.

[1273] A system for implementing this invention is an advanced AI assistant system that includes user authentication, personalized data management, dialogue using natural language processing, emergency response, education and training functions, etc. Specific embodiments will be described in detail below.

[1274] User authentication

[1275] The server receives the user's authentication information and authenticates it by checking it against a database. Specifically, when the user logs in for the first time, they enter their username and password, which are sent to the server via their terminal. The server checks the received authentication information against its existing database, and if authentication is successful, it starts a session, or returns an error message if it fails.

[1276] Managing Personalization Data

[1277] The server displays a screen for the user to enter personalization data and stores the data entered by the user in a database. For example, the user enters personal information such as name, age, health information, hobbies, etc., which are then stored in the database. This allows the server to provide services based on the user's specific needs and preferences.

[1278] Dialogue using natural language processing

[1279] When a user speaks, the device converts the speech into text and sends it to the server. The server then analyzes the text using a natural language processing engine that uses a generative AI model to generate an appropriate answer to the user's question or request. The answer is then returned to the user via the device in voice or text.

[1280] Specific examples

[1281] For example, if a user says, "What's your lunch recommendation?", the server will suggest the perfect lunch based on the stored personalized data. Also, if a user requests, "I want to practice my English pronunciation," the server will provide appropriate learning materials.

[1282] Emergency response

[1283] In the event of an emergency, the server will obtain the user's location and health data and arrange for emergency response if necessary. For example, if the user yells, "My chest hurts! Help!", the device will detect this emergency keyword and immediately send an alert to the server. The server will then obtain the user's location, arrange for an ambulance, and notify their family.

[1284] Education and Training Features

[1285] When a user wishes to learn or train, the server selects appropriate learning materials and training programs and provides them via the terminal, allowing the user to efficiently self-study and train their skills.

[1286] Hardware and Software

[1287] The system uses the following hardware and software:

[1288] Hardware: Devices such as smartphones, smart glasses, and head-mounted displays

[1289] software:

[1290] Speech Recognition Library: speech_recognition

[1291] Natural language processing library: transformers

[1292] Database management library: SQLAlchemy

[1293] Generative AI model: "cl-tohoku / bert-base-japanese" as an example

[1294] Examples of prompt statements

[1295] For example, a prompt to a generative AI model might look like this:

[1296] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[1297] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1298] Step 1:

[1299] User authentication: When a user logs in for the first time, they enter their username and password. The device sends this information to the server. The server checks the received authentication information against an existing database and starts the session if authentication is successful. If authentication fails, it returns an error message to the device.

[1300] Input: Username, Password

[1301] Processing: Match against database

[1302] Output: Session started or error message

[1303] Step 2:

[1304] Input of personalized data: If authentication is successful, the device will display a personalized data input screen for the user. The user will input personal information such as name, age, health information, hobbies, etc. This data will be sent from the device to the server, where it will be stored in a database.

[1305] Input: Name, age, health information, hobbies

[1306] Processing: Save to database

[1307] Output: Personalized data

[1308] Step 3:

[1309] Speech recognition and natural language processing: The user speaks a question or request into the device. The device converts the speech into text and sends this text to the server. The server uses a natural language processing engine with a generative AI model to analyze the text and generate an appropriate response. This response is sent back to the device and returned to the user as voice or text.

[1310] Input: Voice input (questions and requests)

[1311] Processing: Speech-to-text conversion, text analysis, and generating appropriate answers

[1312] Output: Text or audio response

[1313] Step 4:

[1314] Management of dietary history and health information: When a user makes a dietary question or request, the server analyzes the stored personalized data (dietary history and health information) and generates appropriate dietary suggestions. The suggestions are notified to the user via the device.

[1315] Input: dietary requests, past dietary history, health information

[1316] Processing: Retrieving information from the database, analyzing it, and generating meal suggestions

[1317] Output: Meal suggestion notification

[1318] Step 5:

[1319] Emergency Response: When a user utters a keyword such as "Help!" in an emergency, the device detects this and immediately sends an alert to the server. The server obtains the user's location and health data and automatically arranges for local emergency response if necessary. The server also notifies the user's family and friends in an emergency.

[1320] Input: Emergency keywords, location information, health data

[1321] Processing: Keyword detection, location acquisition, emergency response arrangements, emergency contact

[1322] Output: Arrange emergency response, notify emergency contact

[1323] Step 6:

[1324] Education and Training: When a user wishes to learn or receive training, the user makes a request through the device, which is then sent to the server, which selects the appropriate educational material or training program and provides it through the device.

[1325] Input: Learning or training request

[1326] Processing: Analyzing your request and selecting educational materials and programs

[1327] Output: Providing educational materials and training programs

[1328] For example, if a user says to the device, "Tell me what lunch you recommend," the device converts the speech into text and sends it to the server. The server then takes into account the user's past eating history and health information to suggest the optimal lunch. The prompt to the generative AI model is as follows:

[1329] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[1330] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1331] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and an emotion engine, emergency response, and education and training. This system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[1332] 1. User Authentication

[1333] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[1334] 2. Managing Personalized Data

[1335] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[1336] 3. Daily conversation and problem consultation

[1337] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or concern. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and even the user's emotional state using an emotion engine. This answer is returned to the user via voice or text via the device. For example, if the user says, "I've been feeling tired lately," the device will not only suggest ways to relax and rest, but if the emotion engine detects that the user is feeling stressed, it will also provide measures that are particularly useful for reducing stress.

[1338] 4. Emergency Response

[1339] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends in an emergency. For example, if a user yells "My chest hurts, help!", the server will call an ambulance based on the user's location information and notify their family. The emotion engine can adjust the priority of emergency response according to the user's emotional state.

[1340] 5. Education and Training Functions

[1341] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," the server provides a basic pronunciation practice program and begins instruction. The emotion engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[1342] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[1343] The processing flow will be explained below.

[1344] 1. User Authentication

[1345] Step 1:

[1346] The user enters their username and password into the login screen.

[1347] Step 2:

[1348] The terminal transmits the entered authentication information to the server.

[1349] Step 3:

[1350] The server compares the received authentication information with the database and generates an authentication result.

[1351] Step 4:

[1352] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[1353] Step 5:

[1354] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[1355] 2. Managing Personalized Data

[1356] Step 1:

[1357] The terminal displays a personalized data entry screen when you log in for the first time.

[1358] Step 2:

[1359] Users enter personal data such as name, age, health information, hobbies, etc.

[1360] Step 3:

[1361] The terminal transmits the input data to the server.

[1362] Step 4:

[1363] The server stores the received data in a database.

[1364] 3. Daily conversation and problem consultation

[1365] Step 1:

[1366] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[1367] Step 2:

[1368] The device converts the user's voice into text and sends it to the server.

[1369] Step 3:

[1370] The server analyzes the received text using a natural language processing engine.

[1371] Step 4:

[1372] Based on the analysis results, the server analyzes the user's past data, current situation, and even the user's emotional state using an emotion engine.

[1373] Step 5:

[1374] The server generates the best answer and sends it to the device.

[1375] Step 6:

[1376] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[1377] 4. Emotional engine response

[1378] Step 1:

[1379] The device collects voice and text data when the user speaks.

[1380] Step 2:

[1381] The device sends voice data and text data to the server.

[1382] Step 3:

[1383] The server analyzes the voice and text data using an emotion engine.

[1384] Step 4:

[1385] The server stores the user's emotional state as an analysis result and reflects it in generating answers to everyday conversations and problem consultations.

[1386] Step 5:

[1387] The device will provide appropriate feedback and suggestions to the user based on their emotional state. For example, if the user is feeling stressed, it will suggest, "Why don't you take a short walk?"

[1388] 5. Emergency Response

[1389] Step 1:

[1390] The user utters an emergency keyword such as "Help!"

[1391] Step 2:

[1392] The device detects the emergency keyword and immediately sends an alert to the server.

[1393] Step 3:

[1394] The server collects the user's location and health data.

[1395] Step 4:

[1396] The server automatically contacts local emergency services.

[1397] Step 5:

[1398] The server will also contact the user's family and friends in an emergency.

[1399] Step 6:

[1400] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[1401] 6. Education and Training Functions

[1402] Step 1:

[1403] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[1404] Step 2:

[1405] The device sends a request to the server.

[1406] Step 3:

[1407] The server selects appropriate educational materials and training programs.

[1408] Step 4:

[1409] The server transmits the selected teaching materials and programs to the terminal.

[1410] Step 5:

[1411] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[1412] Step 6:

[1413] The server uses an emotion engine to analyze the user's emotional state and motivation for learning.

[1414] Step 7:

[1415] The device will suggest approaches based on the user's motivation and optimize the learning effect. For example, it will suggest specific steps such as, "Your English pronunciation is going well, so next let's practice using example sentences."

[1416] This allows users to receive not only everyday support but also individualized assistance based on their emotional state. By combining this with the emotion engine, even more personalized services can be realized.

[1417] Example 2

[1418] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1419] Conventional AI assistant systems often have separate functions for user authentication, personalized data management, dialogue using natural language processing and sentiment analysis technology, emergency response, and education and training, resulting in a lack of a comprehensive, integrated system. Furthermore, they lack the ability to provide personalized services that take the user's emotional state into account, making it difficult to provide detailed responses based on real-time sentiment analysis of the user.

[1420] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text and analyzing it with a natural language processing engine to generate an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for providing appropriate learning materials and training programs in accordance with the user's learning and training requirements, means for analyzing the user's emotional state using an emotion analysis engine and providing personalized services, means for detecting emergency keywords and immediately sending an alert to the server, means for automatically arranging local emergency response in an emergency and making emergency contact with the user's family and friends, and means for recording the user's progress and providing appropriate feedback. This enables comprehensive support and the provision of meticulous services tailored to the user's emotional state.

[1421] "Authentication Information" means information used to verify a user's identity, such as a username and password.

[1422] A "database" is a system for efficiently storing, managing, and retrieving data.

[1423] A "session" refers to the period during which a user is logged into a system and performs continuous operations.

[1424] "Personalized data" is information specific to each user, including name, age, health information, hobbies, etc.

[1425] A "natural language processing engine" is a general term for algorithms and software that understand, analyze, and generate responses to human language.

[1426] "Emotion analysis engine" is a general term for algorithms and software that analyze a user's emotional state from their statements and actions.

[1427] "Emergency keywords" are specific words or phrases that indicate an emergency, such as "Help!"

[1428] "Location information" refers to information that indicates the user's current location, including GPS data.

[1429] "Health data" refers to information about the user's health status, such as heart rate and blood pressure.

[1430] "Emergency response" refers to a series of procedures and actions to respond to an emergency, including, for example, arranging for an ambulance.

[1431] "Instructional Materials" means educational materials and content provided for User learning or training.

[1432] A "Training Program" is a planned sequence of study or practice designed to improve a User's specific skills or knowledge.

[1433] "Feedback" refers to evaluation and guidance of a user's behavior and performance, providing information that will help them improve their next actions.

[1434] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and emotion analysis technology, emergency response, and education and training. The system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[1435] Hardware and software used

[1436] The system uses the following key technologies:

[1437] Server: Equipped with database management, natural language processing engine, and sentiment analysis engine.

[1438] Database Management: MySQL, PostgreSQL

[1439] Natural Language Processing: OpenAI GPT

[1440] Sentiment analysis engine: IBM Watson Tone Analyzer

[1441] Terminal: Equipped with user interface, speech recognition and text generation functions.

[1442] User Interface: HTML, CSS, JavaScript

[1443] Speech Recognition: Google's Speech-to-Text

[1444] Text generation: OpenAI GPT

[1445] System Overview

[1446] 1. User Authentication

[1447] When a user logs in for the first time, they enter their username and password. This information is sent via the terminal to the server. The server checks the authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If it fails, it returns an error message instructing the user to enter their authentication information again.

[1448] 2. Managing Personalized Data

[1449] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters personal information such as name, age, health information, and hobbies. This data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[1450] 3. Daily conversation and problem consultation

[1451] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and emotional state using an emotion analysis engine. This answer is returned to the user via the device in voice or text.

[1452] 4. Emergency Response

[1453] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then collect the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends. The emotion analysis engine can adjust the priority of emergency response depending on the user's emotional state.

[1454] 5. Education and Training Functions

[1455] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. The emotion analysis engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[1456] Specific examples

[1457] 1. Example of a user authentication prompt

[1458] Username: example_user

[1459] Password: example_pass

[1460] 2. Examples of prompts for personalized data entry

[1461] Name: Taro

[1462] Age: 25

[1463] Health information: Good

[1464] Hobbies: Reading

[1465] 3. Examples of prompts for everyday conversation and problem-solving

[1466] I've been feeling tired lately

[1467] 4. Emergency Response Prompt Examples

[1468] help me!

[1469] 5. Examples of Prompt Sentences for Education and Training Functions

[1470] I want to practice my English pronunciation

[1471] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion analysis engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[1472] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1473] Step 1:

[1474] A user opens an application and enters their username and password on the login screen. The entered credentials are encrypted by the device and sent to the server, which becomes the input.

[1475] Step 2:

[1476] The server compares the received authentication information with the database (MySQL or PostgreSQL). The result of the comparison is determined to be either success or failure, and in either case the result is returned to the terminal in JSON format (output). This is data comparison and calculation.

[1477] Step 3:

[1478] The device receives the authentication result from the server, and displays the home screen if authentication is successful, or an error message if it is unsuccessful, allowing the user to perform the next operation.

[1479] Step 4:

[1480] When logging in for the first time, the device will display a personalized data entry screen, where the user will enter their name, age, health information, and hobbies. This will be their new input.

[1481] Step 5:

[1482] The device encrypts the entered personalized data and sends it to the server. The server receives the data and stores it in the database. After storage is complete, it sends a success message in JSON format to the device (output). This is the data storage operation.

[1483] Step 6:

[1484] The device receives the success message from the server and displays "Data saving completed" to the user. This step allows the user to confirm that the data was saved correctly.

[1485] Step 7:

[1486] When a user speaks (e.g., "I'm feeling tired these days"), the device converts the speech into text using Google's Speech-to-Text, which becomes the new input.

[1487] Step 8:

[1488] The device sends the converted text to the server. The server analyzes the text using OpenAI GPT to understand the user's intent. It also analyzes the emotional state using IBM Watson Tone Analyzer. Based on the analysis results, it generates an appropriate answer and sends it to the device in JSON format (output). This is how the data analysis and generative AI model works.

[1489] Step 9:

[1490] The device synthesizes the answer from the server as voice and responds to the user (e.g., "Try deep breathing and light exercise to relax"), allowing the user to receive specific advice.

[1491] Step 10:

[1492] If a user shouts "Help!" in an emergency, the device will detect the emergency keyword using the voice recognition system, which will become the new input.

[1493] Step 11:

[1494] The device generates an emergency alert and sends it to the server. The server collects the user's location and health data, automatically dispatches local emergency response, and also contacts family and friends in an emergency. This is data collection and dispatch of emergency response.

[1495] Step 12:

[1496] The progress status from the server is sent to the terminal in real time, and the terminal displays to the user, "The ambulance is scheduled to arrive. Please wait calmly." This allows the user to understand the progress of the emergency response.

[1497] Step 13:

[1498] If a user requests, "I want to practice my English pronunciation," the device sends that request to the server, which becomes the new input.

[1499] Step 14:

[1500] The server selects appropriate teaching materials and training programs and sends them to the terminal. The terminal displays them and instructs the user to "start practicing basic English pronunciation." This is the process of providing data and displaying the user interface.

[1501] Step 15:

[1502] The user follows the instructions to practice pronunciation. The device records the progress and sends it to the server. The server analyzes the progress and generates appropriate feedback and sends it to the device.

[1503] Step 16:

[1504] The device will display feedback to the user saying, "Your pronunciation is improving. Now let's practice the 'th' sound." This allows the user to effectively progress with their learning / training.

[1505] (Application example 2)

[1506] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1507] Security is becoming increasingly important in modern homes and offices. However, existing security systems are unable to provide users with adequately personalized measures for emergency response and daily security needs. Furthermore, it is difficult to take immediate and appropriate action in the event of an emergency, resulting in insufficient systems to ensure user safety. Furthermore, limited opportunities for security education and training mean that users lack support for effective self-learning. To address these challenges, an advanced AI assistant system that combines user authentication, personalized data management, natural language processing, and an emotion engine is needed.

[1508] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for immediately sending a notification to the user's family or emergency contacts, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for detecting suspicious individuals and comparing abnormal behavior with a learning database based on voice recognition to evaluate safety, means for automatically notifying a security company based on the safety evaluation results, means for using a generative AI model to generate security-related educational content, and means for creating prompt sentences for the generative AI model. This improves the user's emergency response capabilities and enables personalized responses to everyday security needs. It will also enable effective self-learning and training on security, which is expected to improve overall security.

[1509] "User authentication" is the process of receiving a user's authentication information via a terminal and verifying it against a database.

[1510] "Personalized Data" means individual information about a user, such as the user's name, age, health information, and hobbies, that is used to personalize the service.

[1511] A "natural language processing engine" is a software engine that analyzes a user's voice and text, understands their meaning, and generates appropriate answers.

[1512] The "emotion engine" is an engine that analyzes the user's emotional state and responds appropriately based on that state.

[1513] "Emergency response" is the process of obtaining location and health data and arranging the necessary response when a user is in an emergency.

[1514] A "generative AI model" is an artificial intelligence model that generates content and answers based on user requests and situations.

[1515] A "security assistant" is an application that provides user authentication, personalized data management, natural language processing, emotion engine, emergency response, and education and training capabilities to enhance home and office security.

[1516] "Speech recognition" is a technology that converts a user's voice into text and analyzes it.

[1517] "Suspicious Person Detection" is the process of using voice recognition and other sensors to identify suspicious behavior or people.

[1518] "Automatic notification to security companies" is a process in which the system automatically sends notifications to security companies when a suspicious person or emergency occurs.

[1519] "Educational content generation" is the process of using generative AI models to create appropriate educational content based on users' learning and training requirements.

[1520] "Prompt generation" is the process of creating textual instructions to input to a generative AI model.

[1521] The present invention provides a security assistant for improving home and office security using an advanced AI assistant system. Specific embodiments for carrying out the invention are described below.

[1522] 1. User Authentication

[1523] This system uses facial recognition using a smartphone camera for user authentication. The device acquires the user's facial image and performs facial recognition using OpenCV. If authentication is successful, a session is started and the user can access the application. If authentication fails, an error message is displayed.

[1524] 2. Managing Personalized Data

[1525] When logging in for the first time, the device displays a personalized data entry screen for the user, asking them to enter their name, age, health information, hobbies, etc. The entered data is then stored in an SQLite database via the device, allowing the device to provide services tailored to the user's needs and preferences.

[1526] 3. Daily conversation and problem consultation

[1527] When a user speaks to the device, the device converts the speech into text and analyzes the text using the Google Cloud Natural Language API. Based on the analysis results, the server takes into account the user's past data and current situation and generates an appropriate answer using a natural language processing engine. This answer is returned to the user via voice or text.

[1528] 4. Emergency Response

[1529] If a user experiences an emergency, the emergency response process will begin when emergency keywords such as "help" are detected through voice recognition. The server will then obtain the user's location and health data and notify security companies and emergency contacts using communication APIs such as Twilio. Notifications will also be sent simultaneously to the user's family and friends.

[1530] 5. Education and Training Functions

[1531] When a user requests security training or learning, the server uses a generative AI model (e.g., GPT-4) to generate prompts and provide appropriate learning materials and training programs to the device, allowing the user to efficiently self-study and train their skills.

[1532] Examples of concrete examples and prompts

[1533] For example, if a user is in an emergency situation, say, "There might be a suspicious person around! Help!" in the middle of the night, the device will recognize this voice and notify the server. The server will then obtain the user's location and immediately notify the security company automatically, while also sending a message to the user's family.

[1534] Also, here are some examples of learning and training prompts:

[1535] User: "I want to practice my English pronunciation."

[1536] AI: "Of course! I'll guide you. Let's start with the basic vowels. Say 'a'."

[1537] In this way, the system uses advanced AI technology to provide users with personalized security services.

[1538] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1539] Step 1:

[1540] The user starts up the device and performs facial authentication. The device uses the smartphone camera to capture the user's facial image and performs facial recognition using OpenCV. If facial authentication is successful, the device sends authentication information to the server and starts a session. If authentication fails, the device displays an error message.

[1541] Input: A face image taken from the device's camera

[1542] Output: Authentication successful or error message

[1543] Step 2:

[1544] When a user logs in for the first time, they enter their personalization data. The device displays a screen for entering their name, age, health information, hobbies, etc., and the user enters the information. The entered data is sent to the server via the device and stored in an SQLite database.

[1545] Input: Personalization data entered by the user

[1546] Output: Personalization data stored in a database

[1547] Step 3:

[1548] When a user speaks to the device, the device converts the speech into text. The device uses a microphone to capture the speech and converts it into text using the Google Cloud Natural Language API. The server receives this text, analyzes it with a natural language processing engine, and generates an appropriate response. The generated response is returned to the user via the device as voice or text.

[1549] Input: User's voice

[1550] Output: Voice and text responses

[1551] Step 4:

[1552] In an emergency, if the user utters an emergency keyword such as "help," the device recognizes the voice and initiates the emergency response process. The device analyzes the voice and, if it detects an emergency keyword, sends a notification to the server. The server obtains the user's location and health data and uses communication APIs such as Twilio to send notifications to security companies and emergency contacts. Notifications are also sent to the user's family at the same time.

[1553] Input: User's emergency voice

[1554] Output: Send emergency notification

[1555] Step 5:

[1556] When a user requests security education or training, the server uses a generative AI model to generate prompts and generate appropriate learning materials and training programs. The generated content is then provided to the user via their device, allowing the user to efficiently self-study and train their skills.

[1557] Input: User's learning request

[1558] Output: Teaching materials and training programs

[1559] Step 6:

[1560] When using voice recognition to detect suspicious individuals or recognize abnormal behavior, the device captures the voice and sends it to the server. The server analyzes the voice and compares it with a learning database to evaluate safety. Based on the evaluation results, it automatically notifies the security company.

[1561] Input: Audio data

[1562] Output: Safety assessment and notification to security company

[1563] Step 7:

[1564] When the user logs in again, the device will perform facial recognition again. If the authentication is successful, the device will provide the user with optimized services based on the personalization data entered last time.

[1565] Input: A face image taken from the device's camera

[1566] Output: Service provided after successful authentication

[1567] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1568] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1569] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1570] [Fourth embodiment]

[1571] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1572] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1573] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1574] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1575] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1576] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1577] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1578] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1579] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1580] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1581] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1582] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1583] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1584] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[1585] 1. User Authentication

[1586] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[1587] 2. Managing Personalized Data

[1588] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[1589] 3. Daily conversation and problem consultation

[1590] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text. For example, if the user says, "I've been feeling tired from work lately," the device will suggest ways to relax and take a break.

[1591] 4. Emergency Response

[1592] If a user experiences an emergency, the device will detect this by uttering an emergency keyword such as "Help!" and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also provide support by contacting the user's family and friends in an emergency. For example, if a user yells "My chest hurts! Help!", the server will use the user's location information to call an ambulance and notify their family.

[1593] 5. Education and Training Functions

[1594] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," a basic pronunciation practice program is provided and instruction begins.

[1595] As described above, this system can provide comprehensive support to users throughout their lives.

[1596] The processing flow will be explained below.

[1597] 1. User Authentication

[1598] Step 1:

[1599] The user enters their username and password into the login screen.

[1600] Step 2:

[1601] The terminal transmits the entered authentication information to the server.

[1602] Step 3:

[1603] The server compares the received authentication information with the database and generates an authentication result.

[1604] Step 4:

[1605] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[1606] Step 5:

[1607] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[1608] 2. Managing Personalized Data

[1609] Step 1:

[1610] The terminal displays a personalized data entry screen when you log in for the first time.

[1611] Step 2:

[1612] Users enter personal data such as name, age, health information, and hobbies.

[1613] Step 3:

[1614] The terminal transmits the input data to the server.

[1615] Step 4:

[1616] The server stores the received data in a database.

[1617] 3. Daily conversation and problem consultation

[1618] Step 1:

[1619] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[1620] Step 2:

[1621] The device converts the user's voice into text and sends it to the server.

[1622] Step 3:

[1623] The server analyzes the received text using a natural language processing engine.

[1624] Step 4:

[1625] The server generates an appropriate answer based on the user's past data and current situation.

[1626] Step 5:

[1627] The server generates a response and sends it to the terminal.

[1628] Step 6:

[1629] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[1630] 4. Emergency Response

[1631] Step 1:

[1632] The user utters an emergency keyword such as "Help!"

[1633] Step 2:

[1634] The device detects the emergency keyword and immediately sends an alert to the server.

[1635] Step 3:

[1636] The server collects the user's location and health data.

[1637] Step 4:

[1638] The server automatically contacts local emergency services.

[1639] Step 5:

[1640] The server will also contact the user's family and friends in an emergency.

[1641] Step 6:

[1642] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[1643] 5. Education and Training Functions

[1644] Step 1:

[1645] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[1646] Step 2:

[1647] The device sends a request to the server.

[1648] Step 3:

[1649] The server selects appropriate educational materials and training programs.

[1650] Step 4:

[1651] The server transmits the selected teaching materials and programs to the terminal.

[1652] Step 5:

[1653] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[1654] Example 1

[1655] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1656] Conventional AI assistant systems have had difficulty responding appropriately to diverse user needs. In particular, they lacked the ability to respond quickly in emergencies, manage personalized data based on individual user needs, and provide advanced dialogue functions that combine voice recognition and voice synthesis technologies. As a result, user convenience and safety were not sufficiently ensured.

[1657] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1658] In this invention, the server includes means for receiving user authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's speech into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as needed, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for analyzing prompt sentences based on the user's diverse needs and requests using a generative AI model and generating responses based thereon, means for analyzing the user's speech and responding vocally using speech recognition and synthesis technology, and means for detecting emergency keywords and sending immediate alerts. This allows for a wide range of support to be provided to users, enabling rapid response in emergencies and the provision of services tailored to individual needs.

[1659] "User authentication" is the process of verifying a user's identity when accessing a system, usually through a username and password.

[1660] A "database" is an electronic system for managing, storing, retrieving, and updating data efficiently and effectively.

[1661] A "session" refers to a series of operations and communications between when a user logs in to the system and when they log out.

[1662] "Personalized Data" refers to information about a user (e.g., name, age, health information, hobbies, etc.) that the system uses to provide the user with the most appropriate service.

[1663] A "natural language processing engine" refers to an algorithm or model for understanding and processing human language, and is used, for example, to analyze the intent and content of a conversation.

[1664] "Emergency response" refers to the process of responding quickly to a user's emergency situation and arranging for the necessary emergency services.

[1665] A "generative AI model" refers to an artificial intelligence algorithm that automatically generates appropriate responses or content based on user input.

[1666] A "prompt" refers to guided text used to elicit an appropriate response from a generative AI model.

[1667] "Speech recognition" refers to the technology that analyzes a user's speech and converts it into text data.

[1668] "Speech synthesis" refers to the technology of converting text data into voice data and providing information to users via voice.

[1669] "Emergency keywords" are specific words or phrases that the system uses to determine an emergency, and when detected, an emergency response is initiated.

[1670] The system for implementing this invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using a natural language processing engine, emergency response, and education and training. This system exchanges information between a server, a terminal, and a user to meet the needs of the user.

[1671] User authentication

[1672] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database (e.g., MySQL or PostgreSQL), and if authentication is successful, it starts a session and returns a login success message to the user. If authentication fails, it returns an error message instructing the user to enter their authentication information again.

[1673] Managing Personalization Data

[1674] When the user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database (e.g., SQLite or MongoDB). This makes it possible to provide services tailored to the user's needs and preferences.

[1675] Daily conversation and problem consultation

[1676] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine (such as GPT-3 or BERT) to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, it generates an answer that best suits the user's past data and current situation. This answer is returned to the user via the device in voice or text.

[1677] For example, if a user says, "I've been feeling tired at work lately," the system will suggest ways to relax and take a break. The prompts in this case could be:

[1678] Generate an error message if a user fails to log in.

[1679] "Provide health advice if a user enters that running is a hobby."

[1680] "Generate relaxation suggestions for users who have recently been feeling tired from work."

[1681] Emergency response

[1682] If a user experiences an emergency, they can simply say "Help!" or another emergency keyword. The device will detect this and immediately send an alert to the server. The server will then collect the user's location and health data, and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends to provide support.

[1683] For example, if a user cries out, "My chest hurts, help me!", the server will call an ambulance based on the user's location and notify their family.

[1684] Education and Training Features

[1685] When a user wishes to learn or train, for example by requesting "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows users to efficiently self-study and train their skills.

[1686] As a concrete example, if a user requests, "I want to practice my English pronunciation," we will provide them with a basic pronunciation practice program and begin providing instruction.

[1687] As described above, this system can provide comprehensive support to users throughout their lives.

[1688] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1689] Step 1:

[1690] The user enters their username and password on the terminal, which is the authentication input.

[1691] Input: Username and Password

[1692] How it works: A user enters their username and password into the device's login screen.

[1693] Output: The terminal stores the entered information and prepares it for transmission to the next step.

[1694] Step 2:

[1695] The terminal transmits the entered authentication information to the server.

[1696] Input: The username and password entered by the user

[1697] What happens: The device sends authentication information to the server as an HTTPS request (for example, URL: https: / / example.com / api / login).

[1698] Output: An HTTPS request containing authentication information is sent to the server.

[1699] Step 3:

[1700] The server checks the received authentication information against a database.

[1701] Input: Authentication information sent from the device (username and password)

[1702] What it does: The server compares the authentication information with the user information in a database (e.g. MySQL) to see if they match.

[1703] Output: Authentication result (success or failure)

[1704] Step 4:

[1705] The server returns the authentication result to the terminal.

[1706] Input: Authentication result (success or failure)

[1707] Operation: The server sends the authentication result to the terminal as an HTTPS response.

[1708] Output: The terminal receives the authentication result.

[1709] Step 5:

[1710] The device displays the authentication result to the user.

[1711] Input: Authentication result received from the server (success or failure)

[1712] Operation: If successful, the device will transition to the next screen. If unsuccessful, an error message will be displayed saying "Username or password is incorrect."

[1713] Output: The authentication result is displayed to the user.

[1714] Step 6:

[1715] The terminal displays a personalization data entry screen to the user.

[1716] Input: User successfully logged in information

[1717] Action: The device renders and displays the personalization data entry screen.

[1718] Output: The user is presented with a data entry screen.

[1719] Step 7:

[1720] The user enters individual information.

[1721] Input: Personalized data input screen

[1722] How it works: User enters name, age, health information, hobbies, etc.

[1723] Output: Individual personalized data is input to the device.

[1724] Step 8:

[1725] The terminal transmits the input data to the server.

[1726] Input: Personalized data you enter (such as name, age, health information, hobbies, etc.)

[1727] What happens: The device sends this as an HTTPS request to the server (e.g., URL: https: / / example.com / api / personalize).

[1728] Output: An HTTPS request containing personalization data is sent to the server.

[1729] Step 9:

[1730] The server stores the data in a database.

[1731] Input: Personalization data sent from your device

[1732] How it works: The server stores data in a database (e.g., SQLite or MongoDB).

[1733] Output: A message confirming successful save is generated.

[1734] Step 10:

[1735] The server returns a message to the terminal confirming successful saving.

[1736] Input: Personalized data saving success information

[1737] How it works: The server sends a confirmation message to the device as an HTTPS response.

[1738] Output: You will receive a save confirmation message on your device.

[1739] Step 11:

[1740] The device displays a confirmation message to the user.

[1741] Input: Save confirmation message from the server

[1742] What it does: The device displays "Data saved" to the user.

[1743] Output: User receives confirmation to save data.

[1744] Step 12:

[1745] The user speaks into the device.

[1746] Input: Questions and inquiries by voice

[1747] Action: The user speaks into the device.

[1748] Output: The device receives the audio data.

[1749] Step 13:

[1750] The device converts the speech to text.

[1751] Input: Audio data

[1752] How it works: Your device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert speech to text.

[1753] Output: Text data is generated.

[1754] Step 14:

[1755] The terminal transmits the converted text data to the server.

[1756] Input: Text data

[1757] What happens: The device sends text data as an HTTPS request to the server (for example, URL: https: / / example.com / api / chat).

[1758] Output: An HTTPS request containing text data is sent to the server.

[1759] Step 15:

[1760] The server analyzes the text using a natural language processing engine.

[1761] Input: Text data sent from the terminal

[1762] How it works: The server uses a natural language processing engine (e.g., GPT-3) to parse the text and generate an appropriate answer.

[1763] Output: The answer text is generated.

[1764] Step 16:

[1765] The server generates a response text and sends it back to the terminal.

[1766] Input: Generated answer text

[1767] How it works: The server sends the answer text to the device as an HTTPS response.

[1768] Output: The answer text is sent to the terminal.

[1769] Step 17:

[1770] The device displays the answer to the user.

[1771] Input: Response text from the server

[1772] How it works: The device uses speech synthesis technology (e.g., Google Cloud Text-to-Speech API) to synthesize a response and provide it to the user in voice or text.

[1773] Output: The user receives the answer.

[1774] (Application example 1)

[1775] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1776] In modern society, users need personalized services and support based on their individual needs and circumstances. However, current systems struggle to properly manage detailed health information and data and provide optimal recommendations and support. Furthermore, there is a lack of systems that provide integrated support for multiple functions across all aspects of daily life, including diet, health management, and emergency response. Another challenge is providing a system that allows users to easily request and order food and that can respond quickly in emergencies.

[1777] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1778] In this invention, the server includes means for receiving a user's authentication information and verifying it against a database for authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating appropriate responses, means for managing the user's dietary history and health information, means for accepting and analyzing the user's dietary requests and orders via voice or text, and providing personalized dietary suggestions and orders, means for acquiring the user's location information and health data in an emergency and arranging for emergency response as necessary, and means for providing appropriate educational materials and training programs in response to the user's learning and training requests. This allows users to receive appropriate dietary suggestions and orders based on their individual dietary requirements, and ensures rapid and appropriate responses in emergencies.

[1779] "User authentication" is the process of receiving a user's authentication information and verifying it against a database.

[1780] "Personalized Data" is information that is collected and stored about you based on your specific attributes and preferences.

[1781] "Natural language processing" is a technology for analyzing text, understanding its meaning, and generating appropriate answers.

[1782] "Diet history" is a record of meals a user has consumed in the past.

[1783] "Health information" refers to information that includes data related to health, such as the user's health condition and disease information.

[1784] "Meal suggestion" is the process of recommending appropriate meals based on the user's dietary history and health information.

[1785] "Speech recognition" is a technology that converts speech into text.

[1786] "Location information" is data that indicates the user's current location.

[1787] "Emergency response" is the process of providing users with the assistance and services they need in an emergency.

[1788] A "learning or training program" is a collection of materials and exercises provided to users to help them learn or improve their skills.

[1789] A system for implementing this invention is an advanced AI assistant system that includes user authentication, personalized data management, dialogue using natural language processing, emergency response, education and training functions, etc. Specific embodiments will be described in detail below.

[1790] User authentication

[1791] The server receives the user's authentication information and authenticates it by checking it against a database. Specifically, when the user logs in for the first time, they enter their username and password, which are sent to the server via their terminal. The server checks the received authentication information against its existing database, and if authentication is successful, it starts a session, or returns an error message if it fails.

[1792] Managing Personalization Data

[1793] The server displays a screen for the user to enter personalization data and stores the data entered by the user in a database. For example, the user enters personal information such as name, age, health information, hobbies, etc., which are then stored in the database. This allows the server to provide services based on the user's specific needs and preferences.

[1794] Dialogue using natural language processing

[1795] When a user speaks, the device converts the speech into text and sends it to the server. The server then analyzes the text using a natural language processing engine that uses a generative AI model to generate an appropriate answer to the user's question or request. The answer is then returned to the user via the device in voice or text.

[1796] Specific examples

[1797] For example, if a user says, "What's your lunch recommendation?", the server will suggest the perfect lunch based on the stored personalized data. Also, if a user requests, "I want to practice my English pronunciation," the server will provide appropriate learning materials.

[1798] Emergency response

[1799] In the event of an emergency, the server will obtain the user's location and health data and arrange for emergency response if necessary. For example, if the user yells, "My chest hurts! Help!", the device will detect this emergency keyword and immediately send an alert to the server. The server will then obtain the user's location, arrange for an ambulance, and notify their family.

[1800] Education and Training Features

[1801] When a user wishes to learn or train, the server selects appropriate learning materials and training programs and provides them via the terminal, allowing the user to efficiently self-study and train their skills.

[1802] Hardware and Software

[1803] The system uses the following hardware and software:

[1804] Hardware: Devices such as smartphones, smart glasses, and head-mounted displays

[1805] software:

[1806] Speech Recognition Library: speech_recognition

[1807] Natural language processing library: transformers

[1808] Database management library: SQLAlchemy

[1809] Generative AI model: "cl-tohoku / bert-base-japanese" as an example

[1810] Examples of prompt statements

[1811] For example, a prompt to a generative AI model might look like this:

[1812] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[1813] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1814] Step 1:

[1815] User authentication: When a user logs in for the first time, they enter their username and password. The device sends this information to the server. The server checks the received authentication information against an existing database and starts the session if authentication is successful. If authentication fails, it returns an error message to the device.

[1816] Input: Username, Password

[1817] Processing: Match against database

[1818] Output: Session started or error message

[1819] Step 2:

[1820] Input of personalized data: If authentication is successful, the device will display a personalized data input screen for the user. The user will input personal information such as name, age, health information, hobbies, etc. This data will be sent from the device to the server, where it will be stored in a database.

[1821] Input: Name, age, health information, hobbies

[1822] Processing: Save to database

[1823] Output: Personalized data

[1824] Step 3:

[1825] Speech recognition and natural language processing: The user speaks a question or request into the device. The device converts the speech into text and sends this text to the server. The server uses a natural language processing engine with a generative AI model to analyze the text and generate an appropriate response. This response is sent back to the device and returned to the user as voice or text.

[1826] Input: Voice input (questions and requests)

[1827] Processing: Speech-to-text conversion, text analysis, and generating appropriate answers

[1828] Output: Text or audio response

[1829] Step 4:

[1830] Management of dietary history and health information: When a user makes a dietary question or request, the server analyzes the stored personalized data (dietary history and health information) and generates appropriate dietary suggestions. The suggestions are notified to the user via the device.

[1831] Input: dietary requests, past dietary history, health information

[1832] Processing: Retrieving information from the database, analyzing it, and generating meal suggestions

[1833] Output: Meal suggestion notification

[1834] Step 5:

[1835] Emergency Response: When a user utters a keyword such as "Help!" in an emergency, the device detects this and immediately sends an alert to the server. The server obtains the user's location and health data and automatically arranges for local emergency response if necessary. The server also notifies the user's family and friends in an emergency.

[1836] Input: Emergency keywords, location information, health data

[1837] Processing: Keyword detection, location acquisition, emergency response arrangements, emergency contact

[1838] Output: Arrange emergency response, notify emergency contact

[1839] Step 6:

[1840] Education and Training: When a user wishes to learn or receive training, the user makes a request through the device, which is then sent to the server, which selects the appropriate educational material or training program and provides it through the device.

[1841] Input: Learning or training request

[1842] Processing: Analyzing your request and selecting educational materials and programs

[1843] Output: Providing educational materials and training programs

[1844] For example, if a user says to the device, "Tell me what lunch you recommend," the device converts the speech into text and sends it to the server. The server then takes into account the user's past eating history and health information to suggest the optimal lunch. The prompt to the generative AI model is as follows:

[1845] Generate a response when a user says, "What's your lunch recommendation?" taking into account the user's past eating history and health information. The generated response should include meal suggestions and provide specific menu items.

[1846] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1847] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and an emotion engine, emergency response, and education and training. This system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[1848] 1. User Authentication

[1849] When a user logs in for the first time, they enter their username and password. This information is sent to the server via the terminal. The server checks the received authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If authentication is unsuccessful, it returns an error message instructing the user to enter their authentication information again.

[1850] 2. Managing Personalized Data

[1851] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters individual information such as name, age, health information, and hobbies. The entered data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[1852] 3. Daily conversation and problem consultation

[1853] When a user speaks, the device converts the voice into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or concern. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and even the user's emotional state using an emotion engine. This answer is returned to the user via voice or text via the device. For example, if the user says, "I've been feeling tired lately," the device will not only suggest ways to relax and rest, but if the emotion engine detects that the user is feeling stressed, it will also provide measures that are particularly useful for reducing stress.

[1854] 4. Emergency Response

[1855] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then obtain the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends in an emergency. For example, if a user yells "My chest hurts, help!", the server will call an ambulance based on the user's location information and notify their family. The emotion engine can adjust the priority of emergency response according to the user's emotional state.

[1856] 5. Education and Training Functions

[1857] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. For example, in response to a request such as "I want to practice my English pronunciation," the server provides a basic pronunciation practice program and begins instruction. The emotion engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[1858] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[1859] The processing flow will be explained below.

[1860] 1. User Authentication

[1861] Step 1:

[1862] The user enters their username and password into the login screen.

[1863] Step 2:

[1864] The terminal transmits the entered authentication information to the server.

[1865] Step 3:

[1866] The server compares the received authentication information with the database and generates an authentication result.

[1867] Step 4:

[1868] If the server is successful in authenticating, it generates a session ID and returns it to the terminal. If it fails, it returns an error message.

[1869] Step 5:

[1870] The device receives the session ID and returns the user's home screen, prompting them to re-enter the session if an error message is returned.

[1871] 2. Managing Personalized Data

[1872] Step 1:

[1873] The terminal displays a personalized data entry screen when you log in for the first time.

[1874] Step 2:

[1875] Users enter personal data such as name, age, health information, hobbies, etc.

[1876] Step 3:

[1877] The terminal transmits the input data to the server.

[1878] Step 4:

[1879] The server stores the received data in a database.

[1880] 3. Daily conversation and problem consultation

[1881] Step 1:

[1882] The user speaks to the AI ​​concierge, for example, "I've been feeling tired lately. What should I do?"

[1883] Step 2:

[1884] The device converts the user's voice into text and sends it to the server.

[1885] Step 3:

[1886] The server analyzes the received text using a natural language processing engine.

[1887] Step 4:

[1888] Based on the analysis results, the server analyzes the user's past data, current situation, and even the user's emotional state using an emotion engine.

[1889] Step 5:

[1890] The server generates the best answer and sends it to the device.

[1891] Step 6:

[1892] The device will then communicate the server's response to the user via voice or text, for example, "To help you relax, try some light stretches."

[1893] 4. Emotional engine response

[1894] Step 1:

[1895] The device collects voice and text data when the user speaks.

[1896] Step 2:

[1897] The device sends voice data and text data to the server.

[1898] Step 3:

[1899] The server analyzes the voice and text data using an emotion engine.

[1900] Step 4:

[1901] The server stores the user's emotional state as an analysis result and reflects it in generating answers to everyday conversations and problem consultations.

[1902] Step 5:

[1903] The device will provide appropriate feedback and suggestions to the user based on their emotional state. For example, if the user is feeling stressed, it will suggest, "Why don't you take a short walk?"

[1904] 5. Emergency Response

[1905] Step 1:

[1906] The user utters an emergency keyword such as "Help!"

[1907] Step 2:

[1908] The device detects the emergency keyword and immediately sends an alert to the server.

[1909] Step 3:

[1910] The server collects the user's location and health data.

[1911] Step 4:

[1912] The server automatically contacts local emergency services.

[1913] Step 5:

[1914] The server will also contact the user's family and friends in an emergency.

[1915] Step 6:

[1916] The device notifies the user, "An ambulance has been called. Please remain calm as they will arrive shortly."

[1917] 6. Education and Training Functions

[1918] Step 1:

[1919] The user requests what they would like to learn or train, for example, "I want to practice my English pronunciation."

[1920] Step 2:

[1921] The device sends a request to the server.

[1922] Step 3:

[1923] The server selects appropriate educational materials and training programs.

[1924] Step 4:

[1925] The server transmits the selected teaching materials and programs to the terminal.

[1926] Step 5:

[1927] The device provides guidance to the user, saying, "Today, let's practice basic English pronunciation."

[1928] Step 6:

[1929] The server uses an emotion engine to analyze the user's emotional state and motivation for learning.

[1930] Step 7:

[1931] The device will suggest approaches based on the user's motivation and optimize the learning effect. For example, it will suggest specific steps such as, "Your English pronunciation is going well, so next let's practice using example sentences."

[1932] This allows users to receive not only everyday support but also individualized assistance based on their emotional state. By combining this with the emotion engine, even more personalized services can be realized.

[1933] Example 2

[1934] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1935] Conventional AI assistant systems often have separate functions for user authentication, personalized data management, dialogue using natural language processing and sentiment analysis technology, emergency response, and education and training, resulting in a lack of a comprehensive, integrated system. Furthermore, they lack the ability to provide personalized services that take the user's emotional state into account, making it difficult to provide detailed responses based on real-time sentiment analysis of the user.

[1936] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text and analyzing it with a natural language processing engine to generate an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for providing appropriate learning materials and training programs in accordance with the user's learning and training requirements, means for analyzing the user's emotional state using an emotion analysis engine and providing personalized services, means for detecting emergency keywords and immediately sending an alert to the server, means for automatically arranging local emergency response in an emergency and making emergency contact with the user's family and friends, and means for recording the user's progress and providing appropriate feedback. This enables comprehensive support and the provision of meticulous services tailored to the user's emotional state.

[1937] "Authentication Information" means information used to verify a user's identity, such as a username and password.

[1938] A "database" is a system for efficiently storing, managing, and retrieving data.

[1939] A "session" refers to the period during which a user is logged into a system and performs continuous operations.

[1940] "Personalized data" is information specific to each user, including name, age, health information, hobbies, etc.

[1941] A "natural language processing engine" is a general term for algorithms and software that understand, analyze, and generate responses to human language.

[1942] "Emotion analysis engine" is a general term for algorithms and software that analyze a user's emotional state from their statements and actions.

[1943] "Emergency keywords" are specific words or phrases that indicate an emergency, such as "Help!"

[1944] "Location information" refers to information that indicates the user's current location, including GPS data.

[1945] "Health data" refers to information about the user's health status, such as heart rate and blood pressure.

[1946] "Emergency response" refers to a series of procedures and actions to respond to an emergency, including, for example, arranging for an ambulance.

[1947] "Instructional Materials" means educational materials and content provided for User learning or training.

[1948] A "Training Program" is a planned sequence of study or practice designed to improve a User's specific skills or knowledge.

[1949] "Feedback" refers to evaluation and guidance of a user's behavior and performance, providing information that will help them improve their next actions.

[1950] This invention is an advanced AI assistant system that includes functions for user authentication, personalized data management, dialogue using natural language processing and emotion analysis technology, emergency response, and education and training. The system exchanges information between the server, terminals, and users to respond to user needs. Furthermore, it can recognize the user's emotions and respond based on their state, providing even more personalized services.

[1951] Hardware and software used

[1952] The system uses the following key technologies:

[1953] Server: Equipped with database management, natural language processing engine, and sentiment analysis engine.

[1954] Database Management: MySQL, PostgreSQL

[1955] Natural Language Processing: OpenAI GPT

[1956] Sentiment analysis engine: IBM Watson Tone Analyzer

[1957] Terminal: Equipped with user interface, speech recognition and text generation functions.

[1958] User Interface: HTML, CSS, JavaScript

[1959] Speech Recognition: Google's Speech-to-Text

[1960] Text generation: OpenAI GPT

[1961] System Overview

[1962] 1. User Authentication

[1963] When a user logs in for the first time, they enter their username and password. This information is sent via the terminal to the server. The server checks the authentication information against a database, and if authentication is successful, it starts a session and returns a login success message to the user. If it fails, it returns an error message instructing the user to enter their authentication information again.

[1964] 2. Managing Personalized Data

[1965] When a user logs in for the first time, the device displays a personalized data entry screen. The user enters personal information such as name, age, health information, and hobbies. This data is sent to the server via the device and stored in a database. This makes it possible to provide services tailored to the user's needs and preferences.

[1966] 3. Daily conversation and problem consultation

[1967] When a user speaks, the device converts the speech into text and sends it to the server. The server then uses a natural language processing engine to analyze the text and understand the user's question or inquiry. Based on the results of the analysis, the device generates the most appropriate answer by analyzing the user's past data, current situation, and emotional state using an emotion analysis engine. This answer is returned to the user via the device in voice or text.

[1968] 4. Emergency Response

[1969] If a user experiences an emergency, they can utter an emergency keyword such as "Help!", which the device will detect and immediately send an alert to the server. The server will then collect the user's location and health data and automatically arrange for local emergency response if necessary. It will also contact the user's family and friends. The emotion analysis engine can adjust the priority of emergency response depending on the user's emotional state.

[1970] 5. Education and Training Functions

[1971] When a user wishes to learn or train, for example by making a request such as "I want to practice my English pronunciation," the device sends the request to the server. The server then selects appropriate learning materials and training programs and provides them to the user via the device. This allows the user to efficiently self-study and train their skills. The emotion analysis engine can analyze the user's motivation and emotional state regarding learning and suggest an appropriate approach.

[1972] Specific examples

[1973] 1. Example of a user authentication prompt

[1974] Username: example_user

[1975] Password: example_pass

[1976] 2. Examples of prompts for personalized data entry

[1977] Name: Taro

[1978] Age: 25

[1979] Health information: Good

[1980] Hobbies: Reading

[1981] 3. Examples of prompts for everyday conversation and problem-solving

[1982] I've been feeling tired lately

[1983] 4. Emergency Response Prompt Examples

[1984] help me!

[1985] 5. Examples of Prompt Sentences for Education and Training Functions

[1986] I want to practice my English pronunciation

[1987] As described above, this system can provide comprehensive support for users across all aspects of their lives. By combining it with an emotion analysis engine, the level of personalization can be further increased, making it possible to provide detailed services tailored to the user's emotional state.

[1988] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1989] Step 1:

[1990] A user opens an application and enters their username and password on the login screen. The entered credentials are encrypted by the device and sent to the server, which becomes the input.

[1991] Step 2:

[1992] The server compares the received authentication information with the database (MySQL or PostgreSQL). The result of the comparison is determined to be either success or failure, and in either case the result is returned to the terminal in JSON format (output). This is data comparison and calculation.

[1993] Step 3:

[1994] The device receives the authentication result from the server, and displays the home screen if authentication is successful, or an error message if it is unsuccessful, allowing the user to perform the next operation.

[1995] Step 4:

[1996] When logging in for the first time, the device will display a personalized data entry screen, where the user will enter their name, age, health information, and hobbies. This will be their new input.

[1997] Step 5:

[1998] The device encrypts the entered personalized data and sends it to the server. The server receives the data and stores it in the database. After storage is complete, it sends a success message in JSON format to the device (output). This is the data storage operation.

[1999] Step 6:

[2000] The device receives the success message from the server and displays "Data saving completed" to the user. This step allows the user to confirm that the data was saved correctly.

[2001] Step 7:

[2002] When a user speaks (e.g., "I'm feeling tired these days"), the device converts the speech into text using Google's Speech-to-Text, which becomes the new input.

[2003] Step 8:

[2004] The device sends the converted text to the server. The server analyzes the text using OpenAI GPT to understand the user's intent. It also analyzes the emotional state using IBM Watson Tone Analyzer. Based on the analysis results, it generates an appropriate answer and sends it to the device in JSON format (output). This is how the data analysis and generative AI model works.

[2005] Step 9:

[2006] The device synthesizes the answer from the server as voice and responds to the user (e.g., "Try deep breathing and light exercise to relax"), allowing the user to receive specific advice.

[2007] Step 10:

[2008] If a user shouts "Help!" in an emergency, the device will detect the emergency keyword using the voice recognition system, which will become the new input.

[2009] Step 11:

[2010] The device generates an emergency alert and sends it to the server. The server collects the user's location and health data, automatically dispatches local emergency response, and also contacts family and friends in an emergency. This is data collection and dispatch of emergency response.

[2011] Step 12:

[2012] The progress status from the server is sent to the terminal in real time, and the terminal displays to the user, "The ambulance is scheduled to arrive. Please wait calmly." This allows the user to understand the progress of the emergency response.

[2013] Step 13:

[2014] If a user requests, "I want to practice my English pronunciation," the device sends that request to the server, which becomes the new input.

[2015] Step 14:

[2016] The server selects appropriate teaching materials and training programs and sends them to the terminal. The terminal displays them and instructs the user to "start practicing basic English pronunciation." This is the process of providing data and displaying the user interface.

[2017] Step 15:

[2018] The user follows the instructions to practice pronunciation. The device records the progress and sends it to the server. The server analyzes the progress and generates appropriate feedback and sends it to the device.

[2019] Step 16:

[2020] The device will display feedback to the user saying, "Your pronunciation is improving. Now let's practice the 'th' sound." This allows the user to effectively progress with their learning / training.

[2021] (Application example 2)

[2022] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2023] Security is becoming increasingly important in modern homes and offices. However, existing security systems are unable to provide users with adequately personalized measures for emergency response and daily security needs. Furthermore, it is difficult to take immediate and appropriate action in the event of an emergency, resulting in insufficient systems to ensure user safety. Furthermore, limited opportunities for security education and training mean that users lack support for effective self-learning. To address these challenges, an advanced AI assistant system that combines user authentication, personalized data management, natural language processing, and an emotion engine is needed.

[2024] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving user authentication information and comparing it with a database to perform authentication, means for starting a session if authentication is successful and returning an error message if authentication is unsuccessful, means for displaying a screen for inputting personalized data and saving the input data in a database, means for converting the user's voice into text, analyzing the text using a natural language processing engine, and generating an appropriate response, means for acquiring the user's location information and health data in an emergency and arranging emergency response as necessary, means for immediately sending a notification to the user's family or emergency contacts, means for providing appropriate learning materials and training programs in response to the user's learning and training requests, means for detecting suspicious individuals and comparing abnormal behavior with a learning database based on voice recognition to evaluate safety, means for automatically notifying a security company based on the safety evaluation results, means for using a generative AI model to generate security-related educational content, and means for creating prompt sentences for the generative AI model. This improves the user's emergency response capabilities and enables personalized responses to everyday security needs. It will also enable effective self-learning and training on security, which is expected to improve overall security.

[2025] "User authentication" is the process of receiving a user's authentication information via a terminal and verifying it against a database.

[2026] "Personalized Data" means individual information about a user, such as the user's name, age, health information, and hobbies, that is used to personalize the service.

[2027] A "natural language processing engine" is a software engine that analyzes a user's voice and text, understands their meaning, and generates appropriate answers.

[2028] The "emotion engine" is an engine that analyzes the user's emotional state and responds appropriately based on that state.

[2029] "Emergency response" is the process of obtaining location and health data and arranging the necessary response when a user is in an emergency.

[2030] A "generative AI model" is an artificial intelligence model that generates content and answers based on user requests and situations.

[2031] A "security assistant" is an application that provides user authentication, personalized data management, natural language processing, emotion engine, emergency response, and education and training capabilities to enhance home and office security.

[2032] "Speech recognition" is a technology that converts a user's voice into text and analyzes it.

[2033] "Suspicious Person Detection" is the process of using voice recognition and other sensors to identify suspicious behavior or people.

[2034] "Automatic notification to security companies" is a process in which the system automatically sends notifications to security companies when a suspicious person or emergency occurs.

[2035] "Educational content generation" is the process of using generative AI models to create appropriate educational content based on users' learning and training requirements.

[2036] "Prompt generation" is the process of creating textual instructions to input to a generative AI model.

[2037] The present invention provides a security assistant for improving home and office security using an advanced AI assistant system. Specific embodiments for carrying out the invention are described below.

[2038] 1. User Authentication

[2039] This system uses facial recognition using a smartphone camera for user authentication. The device acquires the user's facial image and performs facial recognition using OpenCV. If authentication is successful, a session is started and the user can access the application. If authentication fails, an error message is displayed.

[2040] 2. Managing Personalized Data

[2041] When logging in for the first time, the device displays a personalized data entry screen for the user, asking them to enter their name, age, health information, hobbies, etc. The entered data is then stored in an SQLite database via the device, allowing the device to provide services tailored to the user's needs and preferences.

[2042] 3. Daily conversation and problem consultation

[2043] When a user speaks to the device, the device converts the speech into text and analyzes the text using the Google Cloud Natural Language API. Based on the analysis results, the server takes into account the user's past data and current situation and generates an appropriate answer using a natural language processing engine. This answer is returned to the user via voice or text.

[2044] 4. Emergency Response

[2045] If a user experiences an emergency, the emergency response process will begin when emergency keywords such as "help" are detected through voice recognition. The server will then obtain the user's location and health data and notify security companies and emergency contacts using communication APIs such as Twilio. Notifications will also be sent simultaneously to the user's family and friends.

[2046] 5. Education and Training Functions

[2047] When a user requests security training or learning, the server uses a generative AI model (e.g., GPT-4) to generate prompts and provide appropriate learning materials and training programs to the device, allowing the user to efficiently self-study and train their skills.

[2048] Examples of concrete examples and prompts

[2049] For example, if a user is in an emergency situation, say, "There might be a suspicious person around! Help!" in the middle of the night, the device will recognize this voice and notify the server. The server will then obtain the user's location and immediately notify the security company automatically, while also sending a message to the user's family.

[2050] Also, here are some examples of learning and training prompts:

[2051] User: "I want to practice my English pronunciation."

[2052] AI: "Of course! I'll guide you. Let's start with the basic vowels. Say 'a'."

[2053] In this way, the system uses advanced AI technology to provide users with personalized security services.

[2054] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2055] Step 1:

[2056] The user starts up the device and performs facial authentication. The device uses the smartphone camera to capture the user's facial image and performs facial recognition using OpenCV. If facial authentication is successful, the device sends authentication information to the server and starts a session. If authentication fails, the device displays an error message.

[2057] Input: A face image taken from the device's camera

[2058] Output: Authentication successful or error message

[2059] Step 2:

[2060] When a user logs in for the first time, they enter their personalization data. The device displays a screen for entering their name, age, health information, hobbies, etc., and the user enters the information. The entered data is sent to the server via the device and stored in an SQLite database.

[2061] Input: Personalization data entered by the user

[2062] Output: Personalization data stored in a database

[2063] Step 3:

[2064] When a user speaks to the device, the device converts the speech into text. The device uses a microphone to capture the speech and converts it into text using the Google Cloud Natural Language API. The server receives this text, analyzes it with a natural language processing engine, and generates an appropriate response. The generated response is returned to the user via the device as voice or text.

[2065] Input: User's voice

[2066] Output: Voice and text responses

[2067] Step 4:

[2068] In an emergency, if the user utters an emergency keyword such as "help," the device recognizes the voice and initiates the emergency response process. The device analyzes the voice and, if it detects an emergency keyword, sends a notification to the server. The server obtains the user's location and health data and uses communication APIs such as Twilio to send notifications to security companies and emergency contacts. Notifications are also sent to the user's family at the same time.

[2069] Input: User's emergency voice

[2070] Output: Send emergency notification

[2071] Step 5:

[2072] When a user requests security education or training, the server uses a generative AI model to generate prompts and generate appropriate learning materials and training programs. The generated content is then provided to the user via their device, allowing the user to efficiently self-study and train their skills.

[2073] Input: User's learning request

[2074] Output: Teaching materials and training programs

[2075] Step 6:

[2076] When using voice recognition to detect suspicious individuals or recognize abnormal behavior, the device captures the voice and sends it to the server. The server analyzes the voice and compares it with a learning database to evaluate safety. Based on the evaluation results, it automatically notifies the security company.

[2077] Input: Audio data

[2078] Output: Safety assessment and notification to security company

[2079] Step 7:

[2080] When the user logs in again, the device will perform facial recognition again. If the authentication is successful, the device will provide the user with optimized services based on the personalization data entered last time.

[2081] Input: A face image taken from the device's camera

[2082] Output: Service provided after successful authentication

[2083] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2084] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2085] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2086] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2087] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2088] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2089] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2090] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2091] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2092] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2093] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2094] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2095] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2096] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2097] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2098] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each...

Claims

1. a means of receiving user credentials and authenticating them against a database; A means to initiate a session if authentication is successful and return an error message if authentication is unsuccessful; a means for displaying a screen for inputting personalization data and storing the input data in a database; A means for converting the user's speech into text, analyzing the text based on a natural language processing engine, and generating appropriate answers; A means to obtain user location and health data in an emergency and arrange emergency response if necessary; and A system that includes a means for providing appropriate learning materials and training programs in response to user learning and training requirements.

2. 10. The system of claim 1, further comprising a natural language processing engine that generates appropriate answers based on the user's past data and current situation.

3. 2. The system of claim 1, further comprising means for detecting an emergency keyword and immediately sending an alert to a server.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A