System

The system addresses the lack of personalized support and security in existing systems by using input, profile generation, response, and authentication devices to offer tailored responses and secure user data, improving user experience and security.

JP2026025500APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128309
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Existing systems fail to provide personalized experiences for individuals, particularly for elderly people, children, and busy housewives, and lack sufficient security measures to protect personal information, leading to inadequate support and potential leaks.

Method used

A system comprising an input device for personal information, a profile generation device, a response generation device, a voice output device, a learning device for behavioral history, an update device, and an authentication device using biometric methods like corneal identification and fingerprint authentication to create personalized responses and secure user data.

Benefits of technology

The system provides personalized responses to user inquiries and concerns while ensuring high-security authentication, enhancing user convenience and protecting personal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025500000001_ABST
    Figure 2026025500000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: input means for inputting a user's personality, hobbies and preferences, and daily behavior patterns; profile generating means for processing information acquired from said input means; response generating means for responding to user's questions and worries based on a profile generated by said profile generating means; voice output means for returning a response generated by said response generating means by voice; learning means for recording and learning a user's behavior history; updating means for updating said profile generating means based on information acquired by said learning means; and authentication means for performing advanced security authentication.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, many people use smartphones, but it is difficult to provide a personalized experience that suits each individual user. Furthermore, there is a problem that elderly people, children, busy housewives, and others in particular cannot receive appropriate support when they need consultation or advice in their daily lives. Furthermore, there is a risk of personal information leaks, so advanced security measures are required. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems with a system that includes an input device for inputting a user's personality, hobbies, preferences, and daily behavioral patterns; a profile generation device for processing the information acquired from the input device; a response generation device for responding to the user's questions and concerns based on the profile generated by the profile generation device; a voice output device for providing the responses generated by the response generation device in audio; a learning device for recording and learning the user's behavioral history; an update device for updating the profile generation device based on the information acquired by the learning device; and an authentication device for performing high-security authentication. This system generates an individual profile based on the information input by the user and can provide appropriate responses to various inquiries and concerns in daily life. Furthermore, the use of high-security measures such as corneal identification, fingerprint authentication, and DNA authentication reduces the risk of personal information leaks.

[0006] The "input means" is a means by which a user inputs personal information such as personality, hobbies, preferences, and daily behavior patterns into a smartphone.

[0007] The "profile generating means" is a means for creating a user profile based on personal information acquired from the input means.

[0008] The "answer generation means" is a means for generating an appropriate answer to a user's question or concern based on the generated profile.

[0009] The "voice output means" is a means for converting the text data generated by the response generation means into voice and returning a response to the user.

[0010] The "learning means" is a means for recording the user's behavior history and updating the user's profile based on that history.

[0011] The "update means" is a means for continuously updating the profile generation means based on new information acquired by the learning means.

[0012] "Authentication means" refers to a means of performing high-security authentication such as corneal identification, fingerprint authentication, or DNA authentication in order to protect the user's personal information.

[0013] The "voice recognition means" is a means for analyzing the user's voice and converting it into text data.

[0014] The "transmission means" is a means for transmitting the converted text data to the server. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The system of the present invention is designed to provide users with personalized responses to their questions and concerns in everyday life. Specific embodiments for carrying out the present invention will be described below.

[0037] System configuration

[0038] This system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[0039] 1. Input Method

[0040] 2. Profile Generation Method

[0041] 3. Response Generation Method

[0042] 4. Audio output means

[0043] 5. Learning Methods

[0044] 6. Update method

[0045] 7. Authentication Methods

[0046] 8. Voice Recognition Methods

[0047] 9. Means of Transmission

[0048] Program processing flow

[0049] Initial Setup Phase

[0050] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0051] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[0052] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[0053] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[0054] Daily Use Phase

[0055] 1. User: Speaks to their smartphone. Example: "What do you recommend for dinner tonight?"

[0056] 2. Terminal: The voice input module recognizes the voice and converts it into text data, which is then sent to the server.

[0057] 3. Server: The received text data is analyzed using a natural language processing engine to extract the intent. The server then references the user profile and past data to generate the optimal response.

[0058] 4. Server: The generated response is sent to the terminal in text format.

[0059] 5. Terminal: The text data is converted into speech using a speech synthesis engine, and a response is provided to the user. Example: "I think curry would be good for dinner tonight. The ingredients needed are..."

[0060] Specific examples

[0061] Example 1: Menu suggestion

[0062] 1. User: "What should I make today?"

[0063] 2. Device: Converts speech into text and sends it to the server.

[0064] 3. Server: Analyzes the text data and checks the user's past meal history, preferences, and refrigerator inventory information. Then, it generates an appropriate menu.

[0065] 4. Server: Sends the generated menu suggestions in text format to the terminal.

[0066] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Today, I recommend teriyaki chicken."

[0067] Example 2: Advice

[0068] 1. User: "I've been having trouble at work lately. What should I do?"

[0069] 2. Device: Converts speech into text and sends it to the server.

[0070] 3. Server: Analyzes the text data and generates advice based on the user's work situation and past consultation history.

[0071] 4. Server: Sends the generated advice in text format to the terminal.

[0072] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Try to complete your daily tasks little by little, at your own pace, without rushing."

[0073] Advanced Security Settings

[0074] 1. User: Open the security settings from the device settings screen. Select from the setting options of Corneal Identification, Fingerprint Identification, and DNA Identification.

[0075] 2. Device: According to the selected authentication method, the necessary biometric data is collected. The data is stored in a temporary storage area and encrypted for security purposes.

[0076] 3. Server: Receives encrypted biometric data from the device and stores it in a secure database. Builds an authentication system and applies it to each user.

[0077] 4. Terminal: Installs and activates the authentication system received from the server locally, granting the user access whenever authentication is successful.

[0078] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching daily life. In addition, the system's advanced security functions can safely protect the user's personal information.

[0079] The processing flow will be explained below.

[0080] Initial Setup Phase

[0081] Step 1:

[0082] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0083] Step 2:

[0084] Terminal: Temporarily stores personal information entered by the user. Checks the completeness and format of the input data and determines whether the next step can be processed.

[0085] Step 3:

[0086] Terminal: The personal information entered is sent to the server. When sending, the data is encrypted to ensure security.

[0087] Step 4:

[0088] Server: Deserializes the personal information received from the device and stores it in a database. Based on the stored data, an initial profile is generated.

[0089] Step 5:

[0090] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[0091] Step 6:

[0092] Terminal: Saves the initial profile and learning model received from the server locally. Notifies the user that the setup is complete.

[0093] Daily Use Phase

[0094] Step 1:

[0095] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[0096] Step 2:

[0097] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[0098] Step 3:

[0099] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[0100] Step 4:

[0101] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[0102] Step 5:

[0103] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[0104] Step 6:

[0105] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[0106] Advanced Security Settings

[0107] Step 1:

[0108] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[0109] Step 2:

[0110] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[0111] Step 3:

[0112] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[0113] Step 4:

[0114] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[0115] Step 5:

[0116] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[0117] Step 6:

[0118] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[0119] Continuous learning phase

[0120] Step 1:

[0121] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[0122] Step 2:

[0123] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[0124] Step 3:

[0125] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[0126] Step 4:

[0127] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[0128] Step 5:

[0129] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[0130] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[0131] Example 1

[0132] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0133] Current systems that provide personalized responses lack the ability to quickly and accurately process user input and provide appropriate responses to individual questions and concerns in everyday life. Furthermore, continuous learning and profile updates based on user behavioral history are insufficient, resulting in issues with the accuracy and relevance of responses. Furthermore, in terms of security, measures to safely protect personal information are insufficient, and the authentication process is cumbersome, resulting in a lack of user convenience.

[0134] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0135] In this invention, the server includes: input means for inputting a user's personal information, personality, hobbies, preferences, and behavioral patterns; profile generation means for processing the information acquired from the input means; response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; voice output means for returning the responses generated by the response generation means by voice; learning means for recording and learning the user's behavioral history; update means for updating the profile generation means based on the information acquired by the learning means; voice recognition means for analyzing the user's voice and converting it into text data; transmission means for transmitting the text data converted by the voice recognition means to the response generation means; authentication means for performing high-security authentication; and initial learning model generation means for generating an individual learning model using the user profile and storing it on the server. This makes it possible to comprehensively process a user's diverse information to provide highly accurate personalized responses, and to improve the accuracy of responses through continuous learning, thereby realizing a safe and convenient authentication process.

[0136] The "input means" is a means for inputting data such as the user's personal information, personality, hobbies, preferences, and behavioral patterns.

[0137] The "profile generating means" is a means for generating a user profile based on information acquired from the input means.

[0138] The "answer generation means" is a means for generating an appropriate answer to a user's question or concern based on the generated user profile.

[0139] The "audio output means" is a means for returning the response generated by the response generating means to the user by voice.

[0140] The "learning means" is a means for recording the user's behavior history and for the system to continuously learn based on that data.

[0141] The "update means" is a means for updating the profile generation means based on new information acquired by the learning means.

[0142] The "voice recognition means" is a means having the function of analyzing the user's voice and converting it into text data.

[0143] The "transmitting means" is a means for transmitting the text data converted by the speech recognition means to the response generating means.

[0144] "Authentication means" refers to a means for performing an advanced authentication process to ensure user security.

[0145] The "initial learning model generating means" is a means for generating an individual learning model based on a user profile and storing it on a server.

[0146] The "voice synthesis means" is a means for converting the text data generated by the response generation means into voice.

[0147] The system of the present invention is designed to provide users with personalized responses to their everyday questions and concerns. The components and operation of this system will now be described in detail.

[0148] System configuration

[0149] The system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system consists of the following main components:

[0150] 1. Input Method

[0151] A user starts up a new smartphone and inputs personal information such as name, date of birth, personality, blood type, hobbies, and preferences. This information is provided to the system via input means.

[0152] 2. Profile Generation Method

[0153] The terminal stores the user information acquired from the input means in a temporary storage area. After verifying the integrity of the data, it sends it to the server. The server updates the database based on the received data and creates a user profile. Database software (e.g., MySQL or PostgreSQL) is used to generate the profile.

[0154] 3. Response Generation Method

[0155] Once the profile is created, the system generates responses to the user's questions and concerns. This response generation process uses a natural language processing engine (for example, IBM Watson's NLP capabilities) to generate the best possible response based on the user profile and past data.

[0156] 4. Audio output means

[0157] The generated response is sent in text format to the user device, which then uses a speech synthesis engine (e.g., Microsoft Azure's Text-to-Speech service) to convert the text into speech and provide the response to the user.

[0158] 5. Learning Methods

[0159] The system records the user's behavior history and continuously learns from it. This learning process is implemented using a machine learning framework (e.g., TensorFlow).

[0160] 6. Update method

[0161] The information obtained by the learning means is reflected in the profile generation means as appropriate, and the profile is updated, thereby improving the accuracy of the system over time.

[0162] 7. Voice Recognition Methods

[0163] When a user speaks a question, the device's voice input module recognizes the speech and converts it into text data. This process uses a speech recognition API (for example, Google's Speech-to-Text API).

[0164] 8. Means of Transmission

[0165] The text data converted by the speech recognition means is sent to the server via the transmission means.

[0166] 9. Authentication Methods

[0167] The system utilizes biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide a high level of security, ensuring that users' personal information is kept safe.

[0168] 10. Initial learning model generation method

[0169] The server generates an initial learning model based on the user profile and stores it on the server. This initial learning model is generated using a machine learning framework such as TensorFlow.

[0170] Specific examples

[0171] Example 1: Menu suggestion

[0172] When a user speaks to their smartphone, "What should I make today?", the voice input module recognizes the speech and converts it into text data. This text data is sent to the server and analyzed by a natural language processing engine. An appropriate menu is generated by referring to the user's past meal history, preferences, and refrigerator inventory information. The generated menu is sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and suggested to the user.

[0173] Example 2: Advice

[0174] When a user asks, "Work hasn't been going well lately. What should I do?", the voice input module converts the voice into text data. This data is sent to the server, where a natural language processing engine analyzes the user's work situation and past consultation history. The most appropriate advice is generated and sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and conveyed to the user.

[0175] AI chatbot-generated prompts

[0176] Example prompt 1: "Generate a response when a user asks verbally, 'What's for dinner tonight?'"

[0177] Example prompt 2: "Generate advice for when a user says, 'I've been having trouble at work lately. What should I do?'"

[0178] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching the user's daily life and protecting it safely.

[0179] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0180] Step 1:

[0181] The user starts up the new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0182] Input: Personal information that users enter into input forms.

[0183] Output: The personal information data entered.

[0184] Step 2:

[0185] The device receives user input and stores the entered data in a temporary storage area (the smartphone's RAM). At this point, the data is verified for integrity and sent to the server. The data is encrypted.

[0186] Input: Personal information data provided by the user.

[0187] Data processing: Data validation and encryption (e.g., AES encryption).

[0188] Output: The encrypted personal information data is sent to the server.

[0189] Step 3:

[0190] The server receives the data sent from the device and stores it in a database (MySQL, PostgreSQL, etc.). It generates a user profile based on the received initial data and creates an initial learning model.

[0191] Input: Encrypted personal information data.

[0192] Data processing: Decrypting data and saving it to the database, generating user profiles (executing SQL queries).

[0193] Output: An initial learning model and a user profile are generated and stored in a database.

[0194] Step 4:

[0195] The device receives the initial learning model from the server and saves it locally (in the smartphone's internal storage). A message is displayed to the user indicating that the setup is complete.

[0196] Input: The initial training model sent from the server.

[0197] Data processing: Saving the initial training model.

[0198] Output: The initial training model is saved locally and a message appears saying that the setup is complete.

[0199] Step 5:

[0200] The user speaks to the smartphone using voice. For example, "What do you recommend for dinner tonight?"

[0201] Input: User's voice input.

[0202] Step 6:

[0203] The device's voice input module recognizes speech and converts it into text data. This process uses a speech recognition API (such as Google's Speech-to-Text API). The converted text data is then sent to the server.

[0204] Input: Audio data.

[0205] Data processing: Recognizing voice data and converting it to text.

[0206] Output: Text data is sent to the server.

[0207] Step 7:

[0208] The server analyzes the received text data using a natural language processing engine (such as IBM Watson's NLP function) to extract the intent, and generates the optimal response by referencing the user profile and past data.

[0209] Input: Text data.

[0210] Data processing: Natural language processing, intent extraction, and response generation.

[0211] Output: The generated response text data.

[0212] Step 8:

[0213] The server sends the generated response in text format to the terminal.

[0214] Input: The generated response text data.

[0215] Output: The response text data is sent to the terminal.

[0216] Step 9:

[0217] The device uses a speech synthesis engine (such as Microsoft Azure's Text-to-Speech service) to convert the text data into speech and provide a response to the user. For example: "I think curry would be good for dinner tonight. The ingredients you need are..."

[0218] Input: Response text data.

[0219] Data processing: Converting text data into audio.

[0220] Output: A spoken response.

[0221] Step 10:

[0222] The learning method records the user's behavior history and the system continuously learns based on that data. This is implemented using a machine learning framework (such as TensorFlow).

[0223] Input: User behavior history data.

[0224] Data processing: Preprocessing data and training machine learning models.

[0225] Output: The trained model and updated model parameters.

[0226] Step 11:

[0227] The server updates the profile generation means based on the information obtained by the learning means, thereby improving the profile and increasing the accuracy of responses.

[0228] Input: trained model parameters.

[0229] Data processing: Profile update work.

[0230] Output: The updated user profile.

[0231] Step 12:

[0232] The authentication method uses biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide high-security authentication.

[0233] Input: Biometric data.

[0234] Data Processing: Biometric data collection and authentication.

[0235] Output: Authentication result.

[0236] Step 13:

[0237] The initial learning model generating means generates an individual learning model based on the user profile and stores it in the server.

[0238] Input: User profile data.

[0239] Data processing: generating and saving learning models.

[0240] Output: Initial training model.

[0241] (Application example 1)

[0242] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0243] Customer service in traditional brick-and-mortar stores relies on the knowledge and experience of store staff, making it difficult to provide consistent, high-quality service to customers. In particular, responses to customer questions and concerns are not personalized, which leads to lower customer satisfaction. It is also difficult to accurately grasp daily customer behavior patterns and preferences, making it difficult to recommend optimal products. Furthermore, there are also issues with in-store security authentication, and personal information may not be adequately protected.

[0244] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0245] In this invention, the server includes: an input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns; a profile generation means for processing the information acquired from the input means; a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; a voice output means for returning the responses generated by the response generation means by voice; a learning means for recording and learning the user's behavioral history; an update means for updating the profile generation means based on the information acquired by the learning means; an authentication means for performing high-security authentication; a voice recognition means for analyzing the voice input and converting it into text data; a natural language processing means for analyzing the text data using a natural language processing engine and generating an optimal response by referring to the user's profile and past data; a voice synthesis means for converting the text data generated by the natural language processing means into voice; a robot that operates as a customer service assistant to support operations in a physical store; and a profile generation means for suggesting products in the physical store. This enables personalized responses to customer questions and concerns, thereby contributing not only to improved customer satisfaction but also to improved security.

[0246] "Input means" refers to a device or function for inputting the user's personality, hobbies, preferences, and daily behavior patterns.

[0247] The "profile generating means" refers to a device or function for processing information acquired from the input means and generating a user profile.

[0248] The "response generating means" refers to a device or function for responding to a user's questions or concerns based on the profile generated by the profile generating means.

[0249] The "audio output means" refers to a device or function for returning the response generated by the response generation means by voice.

[0250] The "learning means" refers to a device or function for recording and learning from the user's behavior history.

[0251] The "update means" refers to a device or function for updating the profile generation means based on the information acquired by the learning means.

[0252] "Authentication means" means a device or function for performing high-security authentication.

[0253] "Speech recognition means" refers to a device or function for analyzing voice input and converting it into text data.

[0254] "Natural language processing means" refers to a device or function that uses a natural language processing engine to analyze text data and generate optimal responses by referring to the user's profile and past data.

[0255] The "speech synthesis means" refers to a device or function for converting text data generated by the natural language processing means into speech.

[0256] A "customer service assistant" is a robot that assists with work in physical stores.

[0257] The "profile generation means for making product suggestions" refers to a device or function that generates a profile for making product suggestions in a physical store.

[0258] The system of the present invention aims to provide personalized responses to questions and concerns that users have about their daily lives in brick-and-mortar stores. Specific embodiments for carrying out the invention will now be described.

[0259] System configuration

[0260] The system consists of the following major components:

[0261] 1. Input Method

[0262] 2. Profile Generation Method

[0263] 3. Response Generation Method

[0264] 4. Audio output means

[0265] 5. Learning Methods

[0266] 6. Update method

[0267] 7. Authentication Methods

[0268] 8. Voice Recognition Methods

[0269] 9. Natural Language Processing Tools

[0270] 10. Speech synthesis means

[0271] 11. Customer Service Assistant

[0272] 12. Product proposal profile generation means

[0273] Details of each means are as follows.

[0274] Program Generation and Hardware / Software Usage

[0275] 1. Initial Setup Phase

[0276] User: Launches the application at the store entrance and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[0277] Terminal: The entered information is temporarily stored, encrypted, and sent to the server.

[0278] Server: Generates a user profile based on the received data and creates a personalized product recommendation model.

[0279] Terminal: Saves the initial model received from the server locally and displays a notification that the setup is complete.

[0280] 2. Operational Phase

[0281] User: Asks questions by voice in-store. Example: "What wines do you recommend these days?"

[0282] Terminal: The speech is converted into text using a speech recognition tool (Google Speech-to-Text API) and sent to the server.

[0283] Server: Analyzes the text data using a natural language processing engine (spaCy or BERT) and generates an appropriate response by referencing the user's profile and past data.

[0284] Server: Sends the generated response in text format to the terminal.

[0285] Terminal: Converts text to speech using a speech synthesis engine (Amazon Polly) and provides responses to the user.

[0286] Specific examples

[0287] Example 1

[0288] User: "What wines have you recommended recently?"

[0289] Device: Converts speech to text and sends it to the server.

[0290] Server: Generates a response based on the user's profile and past data, and sends text data to the device saying, "This time, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[0291] Terminal: A speech synthesis engine is used to convert text into speech and provide guidance to the user.

[0292] Prompt Sentence Examples

[0293] Please enter your name, gender, age, preferences, purchase history, and product interest information.

[0294] What wines do you recommend these days?

[0295] "Today, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[0296] Hardware and software used

[0297] Hardware: Smartphones, smart glasses, in-store robots

[0298] software:

[0299] Speech recognition engine: Google Speech-to-Text API

[0300] Natural language processing engine: spaCy, BERT-based model

[0301] Speech synthesis engine: Amazon Polly

[0302] Database: MySQL, Firebase

[0303] Communication protocol: HTTPS, WebSocket

[0304] In this way, the system of the present invention can provide personalized responses to users' questions and concerns, thereby improving customer satisfaction in brick-and-mortar stores. Furthermore, advanced security features can protect users' personal information.

[0305] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0306] Step 1:

[0307] User: Launches the application at the entrance of a physical store and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[0308] Input: Personal information

[0309] Output: Temporarily saved personal information

[0310] Specific operation: By entering information into the input form and pressing the "Submit" button, the data is stored in a temporary storage area on the device.

[0311] Step 2:

[0312] Terminal: The personal information entered is temporarily stored, encrypted for security purposes, and then sent to the server.

[0313] Input:Temporarily saved personal information

[0314] Output: Encrypted personal information data

[0315] Specific operation: Personal information is encrypted inside the device and sent to the server using the HTTPS communication protocol.

[0316] Step 3:

[0317] Server: Decrypts the received encrypted data and stores it in a database. Generates a user profile and creates a personalized product recommendation model.

[0318] Input: Encrypted personal information data

[0319] Output: User profile and proposed model

[0320] Specific operations: Decrypting encrypted data, running the user profile generation algorithm, and saving the generated profile and proposed model to a database.

[0321] Step 4:

[0322] Terminal: The terminal locally stores the user profile and proposed model received from the server and displays a notification to the user that the setup is complete.

[0323] Input: User profile and proposed model data

[0324] Output: Locally saved profile and proposed model, notification that setup is complete

[0325] Specific behavior: Save received data locally, display completion notification

[0326] Step 5:

[0327] User: Asks questions by dictating as they walk through the store. Example: "What wines do you recommend these days?"

[0328] Input: Voice data (question)

[0329] Output: Start speech recognition via speech input interface

[0330] Specific actions: Press the voice input button on your smartphone or smart glasses and dictate your question.

[0331] Step 6:

[0332] Terminal: Using a voice recognition tool (Google Speech-to-Text API), the voice data is converted into text data and sent to the server.

[0333] Input: Audio data

[0334] Output: Text data

[0335] Specific operations: Calling the speech recognition API, generating converted text data, and sending the text data to the server.

[0336] Step 7:

[0337] Server: Analyzes the text data using natural language processing tools (spaCy or BERT engine) and generates an appropriate response by referencing the user's profile and past data.

[0338] Input: Text data, user profile

[0339] Output: Response text data

[0340] What it does: Runs a natural language analysis engine, consults a user profile database, and generates the best possible response.

[0341] Step 8:

[0342] Server: Sends the generated response in text format to the terminal.

[0343] Input: Response text data

[0344] Output: Text data sent to the terminal

[0345] Specific operation: Processing to send text data to the terminal.

[0346] Step 9:

[0347] Terminal: The response text data is converted into speech using a speech synthesis engine (Amazon Polly) and a response is provided to the user.

[0348] Input: Response text data

[0349] Output: Audio data

[0350] Specific operations: Calling the speech synthesis API, generating converted speech data, and outputting speech from the speaker.

[0351] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0352] The system of the present invention is designed in combination with an emotion engine to provide personalized responses to users' questions and concerns in their daily lives. Specific embodiments for implementing the present invention will be described below.

[0353] System configuration

[0354] This system generates an individual profile based on the user's input and emotional information, and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[0355] 1. Input Method

[0356] 2. Profile Generation Method

[0357] 3. Response Generation Method

[0358] 4. Audio output means

[0359] 5. Learning Methods

[0360] 6. Update method

[0361] 7. Authentication Methods

[0362] 8. Voice Recognition Methods

[0363] 9. Means of Transmission

[0364] 10. Emotion Recognition Engine

[0365] Program processing flow

[0366] Initial Setup Phase

[0367] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0368] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[0369] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[0370] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[0371] Daily Use Phase

[0372] 1. User: Talks to the smartphone and asks questions or asks for advice. For example, "What do you recommend for dinner tonight?"

[0373] 2. Terminal: The voice input module captures the spoken voice and converts it into text data using the voice recognition engine.

[0374] 3. Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[0375] 4. Server: The received text data is analyzed using a natural language processing engine to extract the user's intent. The server then generates the optimal response data by referencing the user profile and past data.

[0376] 5. Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[0377] 6. Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it may reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[0378] Processing using an emotion recognition engine

[0379] 1. User: Speaks, makes facial expressions and shows movements.

[0380] 2. On the device: The emotion recognition engine detects the user's tone of voice, facial expressions, movements, and language patterns to analyze emotions in real time.

[0381] 3. Terminal: Temporarily stores the analyzed emotion data and sends it to the server.

[0382] 4. Server: Receives emotion data and uses it in the profile generation means and response generation means.

[0383] 5. Server: Based on the emotional data, generate a personalized response that is more suited to the user's current situation.

[0384] 6. Device: Provides the user with a voice response based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[0385] Advanced Security Settings

[0386] 1. User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[0387] 2. Terminal: According to the selected authentication method, the necessary biometric authentication data is acquired, temporarily stored, and encrypted.

[0388] 3. Terminal: Sends encrypted authentication data to the server, which checks the data for consistency and integrity upon transmission.

[0389] 4. Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[0390] 5. Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[0391] 6. Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[0392] Continuous learning phase

[0393] 1. User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[0394] 2. Device: Automatically collects user behavior history and usage data, temporarily stores the collected data, and sends it to the server as needed.

[0395] 3. Server: Analyzes the received behavioral history data and updates the user profile. Machine learning algorithms are used to learn the user's behavioral patterns.

[0396] 4. Server: Periodically synchronizes updated profiles and learning results to the device. During synchronization, it checks the consistency and integrity of the data.

[0397] 5. Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[0398] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively. The present invention can provide personalized services based on the user's individual information and emotional information, enriching daily life, and safely protecting the user's personal information with advanced security functions.

[0399] The processing flow will be explained below.

[0400] Initial Setup Phase

[0401] Step 1:

[0402] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0403] Step 2:

[0404] Terminal: Temporarily stores personal information entered by the user, checks the integrity and format of the entered data, and sends it to the server.

[0405] Step 3:

[0406] Server: Receives data sent from the device and stores it in a database. Based on the stored data, it generates an initial profile.

[0407] Step 4:

[0408] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[0409] Step 5:

[0410] Terminal: The initial profile and learning model received from the server are saved locally and the user is notified that the setup is complete.

[0411] Daily Use Phase

[0412] Question and answer processing

[0413] Step 1:

[0414] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[0415] Step 2:

[0416] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[0417] Step 3:

[0418] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[0419] Step 4:

[0420] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[0421] Step 5:

[0422] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[0423] Step 6:

[0424] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[0425] Emotion Recognition Processing

[0426] Step 1:

[0427] User: Speaks, shows facial expressions and actions.

[0428] Step 2:

[0429] On the device: The emotion recognition engine detects the user's tone of voice, facial expressions, movements, and language patterns to analyze emotions in real time.

[0430] Step 3:

[0431] Terminal: Temporarily stores the analyzed emotion data and sends it to the server.

[0432] Step 4:

[0433] Server: Receives emotion data and uses it in the profile generation means and response generation means.

[0434] Step 5:

[0435] Server: Based on the emotional data, it generates a personalized response that is more suited to the user's current situation.

[0436] Step 6:

[0437] The device provides the user with a voice response based on the generated emotion. For example, if the device detects that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[0438] Advanced Security Settings

[0439] Step 1:

[0440] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[0441] Step 2:

[0442] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[0443] Step 3:

[0444] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[0445] Step 4:

[0446] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[0447] Step 5:

[0448] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[0449] Step 6:

[0450] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[0451] Continuous learning phase

[0452] Step 1:

[0453] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[0454] Step 2:

[0455] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[0456] Step 3:

[0457] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[0458] Step 4:

[0459] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[0460] Step 5:

[0461] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[0462] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[0463] Example 2

[0464] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0465] In modern society, people often have many questions and concerns in their daily lives, requiring efficient and personalized responses. Furthermore, it is necessary to provide services that provide higher levels of satisfaction by providing responses that are tailored to the user's emotions and circumstances. However, current systems perform tasks such as user profile creation and updating, voice recognition, emotion recognition, and advanced security authentication separately, resulting in a lack of an integrated and efficient solution.

[0466] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0467] In this invention, the server includes input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns, profile generation means for processing information acquired from the input means, response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means, voice output means for returning the responses generated by the response generation means by voice, learning means for recording and learning the user's behavioral history, update means for updating the profile generation means based on information acquired by the learning means, voice recognition means for analyzing the user's voice and converting it into text data, transmission means for transmitting the text data converted by the voice recognition means to the response generation means, an emotion recognition engine for analyzing the user's emotions, means for personalizing responses based on the emotion data acquired by the emotion recognition engine, and authentication means for performing high-security authentication. This enables personalized responses that reflect the user's situation and emotions.

[0468] 1. "Input means" refers to a device or function that allows a user to input personal information such as personality, hobbies, preferences, and daily behavior patterns into the system.

[0469] 2. "Profile generation means" refers to a device or function that generates a user profile based on information acquired from an input means.

[0470] 3. "Response generation means" refers to a device or function that generates the most appropriate response to a user's question or concern based on the generated user profile.

[0471] 4. "Audio output means" refers to a device or function that provides the response generated by the response generation means to the user by voice.

[0472] 5. "Learning means" refers to a device or function that records the user's behavior history and uses it to improve the system's performance.

[0473] 6. "Update means" refers to a device or function that automatically updates the profile generation means based on information acquired by the learning means.

[0474] 7. "Speech recognition means" means a device or function that captures a user's voice and converts it into text data.

[0475] 8. "Transmitting means" means a device or function that transmits text data converted by the speech recognition means to the response generating means.

[0476] 9. "Emotion recognition engine" means a device or function that analyzes a user's emotions and acquires that data.

[0477] 10. "Personalization means" refers to a device or function that changes responses to suit the user's current situation based on emotional data obtained by an emotion recognition engine.

[0478] 11. "Authentication Means" means a device or function that verifies a user's identity and performs high-security authentication.

[0479] 12. A “generative AI model” is an artificial intelligence-based algorithm that generates appropriate responses to user intent or questions.

[0480] 13. “Prompt sentence” refers to the text data input into a generative AI model, which is used to specifically indicate the user’s question or intention.

[0481] The present invention is a system designed to provide users with personalized responses to questions and concerns they have in their daily lives. This system is composed of multiple components that generate an individual profile based on the user's input information and emotional information, and return appropriate responses. Specific embodiments for implementing the present invention will be described below.

[0482] Hardware and software used

[0483] 1. Terminal: A user device such as a smartphone. This terminal includes a voice input module, a voice recognition engine (e.g., Google Speech-to-Text API), and a voice synthesis engine (e.g., Google Text-to-Speech API).

[0484] 2. Server: A server running a database system (e.g., MySQL or PostgreSQL), a natural language processing engine (e.g., GPT-3), or an emotion recognition engine (e.g., Microsoft Azure Emotion API).

[0485] 3. Generative AI model: An artificial intelligence-based algorithm for analyzing user intent and generating appropriate responses.

[0486] Detailed Description of the Invention

[0487] Initial Setup

[0488] 1. The user starts up a new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0489] 2. The terminal receives the data entered by the user, temporarily stores it in a storage area, verifies the integrity of the data, and then sends it to the server.

[0490] 3. The server stores the received data in a database and generates a user profile, which is then used to create an initial learning model.

[0491] 4. The device receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[0492] daily use

[0493] 1. The user speaks to their smartphone to ask a question or ask for advice, such as, "What do you recommend for dinner tonight?"

[0494] 2. The device captures the voice using the voice input module and converts it into text data using the voice recognition engine.

[0495] 3. The terminal temporarily stores the converted text data, checks the data for consistency, and then sends it to the server.

[0496] 4. The server analyzes the received text data using a natural language processing engine to extract the user's intent. It then generates the optimal response data by referencing the user profile and past data.

[0497] 5. The server sends the generated response data to the terminal.

[0498] 6. The device converts the response data into an audio file using a speech synthesis engine and provides the user with a spoken response. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[0499] emotion recognition

[0500] 1. The user speaks, uses facial expressions, and makes movements.

[0501] 2. The device uses an emotion recognition engine to analyze the user's emotions in real time.

[0502] 3. The device temporarily stores the analyzed emotion data and sends it to the server.

[0503] 4. The server receives the emotion data and uses it in the profile generation means and response generation means.

[0504] 5. The server generates a personalized response based on the emotion data that better suits the user's current situation.

[0505] 6. The device then provides a voice response to the user based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[0506] Security Settings

[0507] 1. The user selects security settings from the device settings menu and chooses an authentication option such as corneal identification, fingerprint authentication, or DNA authentication.

[0508] 2. The device captures and encrypts the biometric data.

[0509] 3. The device sends the encrypted data to the server, which verifies the data's integrity and completeness.

[0510] 4. The server deserializes the encrypted data and stores it in a secure database. It builds an authentication algorithm and associates it with the user profile.

[0511] 5. The server sends the authentication system to the terminal, which installs and configures the system within the terminal.

[0512] 6. The terminal has a security authentication system installed, and grants the user access whenever authentication is successful.

[0513] Continuous learning

[0514] 1. Users use smartphones on a daily basis and use various apps and functions.

[0515] 2. The device automatically collects the user's behavioral history and usage data, temporarily stores it, and sends it to the server.

[0516] 3. The server analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[0517] 4. The server periodically synchronizes the updated profile and learning results to the device, verifying the consistency and completeness of the data.

[0518] 5. The device locally reflects the received update data, providing the user with more personalized services.

[0519] Examples of specific examples and prompts

[0520] Specific examples

[0521] A user asks, "What's the best way to relax after work?"

[0522] The device uses a voice recognition engine to convert the question into text and send it to the server.

[0523] The server analyzes the text, understands the user's intent, and generates optimal advice on how to relax.

[0524] The server sends the generated advice to the terminal, which converts it into an audio file and responds to the user by saying, "Why don't you take a nice, long bath?"

[0525] Prompt Sentence Examples

[0526] "Generate the best response when a user asks, 'What's the best way to relax after work?'"

[0527] "What relaxation methods can we suggest to users when they are feeling stressed?"

[0528] The above is an embodiment of the present invention. This system provides personalized responses based on the user's individual information and emotional information, enriching daily life and ensuring safety.

[0529] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0530] Initial Setup Phase

[0531] Step 1:

[0532] A user starts up their smartphone and enters personal information such as their name, date of birth, personality, blood type, hobbies, etc. The data entered is the user's basic information, such as "Taro Tanaka, May 4, 1990, blood type O, watching movies."

[0533] Input: Personal information (name, date of birth, personality, blood type, hobbies and preferences)

[0534] Output: Personal information data entered

[0535] Step 2:

[0536] The terminal receives the data entered by the user and temporarily stores it in a storage area. At this time, the terminal checks the integrity of the data and monitors whether there are any missing data.

[0537] Input: Personal data entered by the user

[0538] Output: Temporarily saved data

[0539] Step 3:

[0540] The data that has been confirmed by the device is sent to the server using the HTTP protocol, and the data is encrypted.

[0541] Input: Temporarily saved personal information data

[0542] Output: Data sent to the server

[0543] Step 4:

[0544] The server receives the data and stores it in a database (e.g., MySQL or PostgreSQL). It generates a user profile based on the initial data and creates an initial learning model (e.g., TensorFlow or PyTorch).

[0545] Input: Received personal information data

[0546] Output: Generated user profile and initial learning model

[0547] Step 5:

[0548] The server sends the generated initial learning model to the terminal via the endpoint.

[0549] Input: Initial training model

[0550] Output: The trained model sent to the device.

[0551] Step 6:

[0552] The device saves the initial learning model locally and displays a message to the user indicating that setup is complete.

[0553] Input: The received training model

[0554] Output: Locally saved training model and a message saying the setup is complete

[0555] Daily Use Phase

[0556] Step 1:

[0557] Users can ask questions or ask their smartphones by voice, for example, "What do you recommend for dinner tonight?"

[0558] Input: Voice question

[0559] Output: Audio data

[0560] Step 2:

[0561] The device captures audio using the built-in microphone and converts it into text data using a speech recognition engine (e.g., Google Speech-to-Text API).

[0562] Input: Audio data

[0563] Output: Text data

[0564] Step 3:

[0565] The terminal stores the converted text data in a temporary storage area, checks the data for consistency, and then transmits it to the server.

[0566] Input: Text data

[0567] Output: Text data to send to the server

[0568] Step 4:

[0569] The server receives the text data, analyzes it with a natural language processing engine (e.g., GPT-3), extracts the user's intent, and generates the optimal response by referring to the user's profile and past data.

[0570] Input: Received text data

[0571] Output: The generated response data

[0572] Step 5:

[0573] The server transmits the generated response data in text format to the terminal.

[0574] Input: Generated response data

[0575] Output: Response data sent to the terminal

[0576] Step 6:

[0577] The device converts the received response data into an audio file using a speech synthesis engine (for example, Google Text-to-Speech API) and provides a voice response to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[0578] Input: Received response data

[0579] Output: A spoken response provided to the user

[0580] Processing using an emotion recognition engine

[0581] Step 1:

[0582] The user speaks, makes facial expressions, and makes movements.

[0583] Input: voice, facial expressions, movements

[0584] Output: Expressed emotion data

[0585] Step 2:

[0586] The device analyzes the user's emotions in real time using an emotion recognition engine (for example, Microsoft Azure Emotion API).

[0587] Input: Expressed emotion data

[0588] Output: Parsed emotion data

[0589] Step 3:

[0590] The device temporarily stores the analyzed emotion data and transmits it to the server.

[0591] Input: Parsed emotion data

[0592] Output: Emotion data sent to the server

[0593] Step 4:

[0594] The server receives the emotion data and uses it in the profile generation means and response generation means.

[0595] Input: Received emotion data

[0596] Output: Emotion data used in the profile generation and response generation methods

[0597] Step 5:

[0598] The server generates a personalized response based on the emotional data that better suits the user's current situation.

[0599] Input: Emotion data

[0600] Output: Personalized response data

[0601] Step 6:

[0602] The device will provide a voice response to the user based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[0603] Input: Emotion-based response data

[0604] Output: Spoken suggestions

[0605] Advanced Security Settings

[0606] Step 1:

[0607] Users select security settings from the device's settings menu and choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[0608] Input: Select user authentication settings

[0609] Output: The configured authentication options

[0610] Step 2:

[0611] The terminal acquires and encrypts the biometric authentication data according to the selected authentication method.

[0612] Input: Authentication method and biometric data

[0613] Output: Encrypted biometric data

[0614] Step 3:

[0615] The device sends the encrypted data to the server, which verifies the data's integrity and completeness.

[0616] Input: Encrypted biometric data

[0617] Output: Encrypted data sent to the server

[0618] Step 4:

[0619] The server deserializes the encrypted data and stores it in a secure database, building an authentication algorithm and associating it with the user profile.

[0620] Input: Encrypted biometric data

[0621] Output: User profile associated with constructed authentication algorithm

[0622] Step 5:

[0623] The server sends the authentication system to the terminal, which then installs and configures the system within the terminal.

[0624] Input:AuthenticationSystem

[0625] output: Authentication system configured on the device

[0626] Step 6:

[0627] The terminal installs an authentication system and grants the user access whenever authentication is successful.

[0628] Input:AuthenticationSystem

[0629] Output: Access permission granted upon successful authentication

[0630] Continuous learning phase

[0631] Step 1:

[0632] Users use smartphones on a daily basis and utilize various applications and functions.

[0633] Input: Everyday smartphone use

[0634] Output: Application and feature usage data

[0635] Step 2:

[0636] The device automatically collects the user's behavioral history and usage data, temporarily stores it, and sends it to the server.

[0637] Input: Application and feature usage data

[0638] Output: Behavioral history data sent to the server

[0639] Step 3:

[0640] The server analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[0641] Input: Received behavioral history data

[0642] Output: Updated user profile

[0643] Step 4:

[0644] The server periodically synchronizes updated profiles and learning results with the device and checks the consistency and completeness of the data.

[0645] Input: Updated user profile

[0646] Output: Data synced to the device

[0647] Step 5:

[0648] The terminal locally reflects the updated data received from the server, thereby providing the user with a more personalized service.

[0649] Input: Update data synced to the device

[0650] Output: Personalized service

[0651] The above is a description of the specific operations in the processing steps of the present invention. Through this processing, the system provides personalized responses based on the user's individual information and emotional information, realizing a system that can enrich daily life and ensure safety.

[0652] (Application example 2)

[0653] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0654] Current worker support systems in factories provide standard responses without considering the worker's emotions, resulting in insufficient stress management and optimization of work efficiency. Furthermore, when workers ask questions or ask for advice via voice, appropriate responses that take their emotions into consideration are not provided, which can lead to lower worker satisfaction and impact productivity. There is a growing need for a system that can solve these issues and provide responses that take the worker's emotional state into consideration.

[0655] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for inputting the user's personality, hobbies, preferences, and daily behavioral patterns, a profile generation means for processing information acquired from the input means, and a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means. This makes it possible to provide accurate and personalized responses while taking into consideration the user's emotional state.

[0656] The "input means" is a device or interface for inputting information such as the user's personality, hobbies, preferences, and daily behavior patterns.

[0657] The "profile generating means" is a device or program for processing information acquired from the input means and generating an individual profile for the user.

[0658] The "answer generation means" is a device or program for generating appropriate answers to the user's questions and concerns based on the generated user profile.

[0659] The "audio output means" refers to a device or technology for returning the response generated by the response generation means by voice.

[0660] A "learning means" is a device or algorithm that records a user's behavioral history and uses it to train a profile or system.

[0661] The "update means" is a device or technology for updating the profile generation means based on the information acquired by the learning means.

[0662] An "authentication means" is a device or system for performing high-security authentication.

[0663] "Voice input means" refers to a device or technology for capturing spoken voice and converting it into text data using a voice recognition engine.

[0664] "Emotion recognition means" refers to a device or technology for analyzing the user's emotions by analyzing the tone of voice and facial expressions.

[0665] System configuration

[0666] This invention is a "smart industrial assistant" system that supports workers in factories. The system consists of the following main components:

[0667] 1. Input means: An interface for inputting information such as the worker's personality, hobbies and preferences, and daily behavioral patterns, and is typically a smartphone or tablet.

[0668] 2. Profile generation means: A program for processing information obtained from the input means and generating a profile for each worker.

[0669] 3. Response generation means: A program that generates appropriate responses to the worker's questions and concerns based on the generated profile.

[0670] 4. Audio output means: This is a technology for outputting the response generated by the response generation means by voice, and a speaker is used.

[0671] 5. Learning method: An algorithm that records the worker's behavioral history and uses it to train the profile and system.

[0672] 6. Update method: This is a technology that updates the profile generation method based on the information obtained by the learning method.

[0673] 7. Authentication method: A device or system for performing high-security authentication, and options include corneal identification, fingerprint authentication, and DNA authentication.

[0674] 8. Voice input method: This technology captures the voice spoken by the worker and converts it into text data using a voice recognition engine.

[0675] 9. Emotion recognition: This technology analyzes the emotions of workers by analyzing their tone of voice and facial expressions.

[0676] Program processing explanation

[0677] The server uses the Python language, the speech_recognition library for speech recognition, and the pyttsx3 library for speech synthesis. This allows the server to recognize the worker's voice, convert it into text data, and generate an appropriate response based on the worker's profile and emotional information. The generated response is output as audio through the speaker.

[0678] Hardware and software used

[0679] Hardware:

[0680] Microphone: Used to capture audio input.

[0681] Speaker: Used to output responses aloud.

[0682] Robot terminal: Used as an installation platform.

[0683] software:

[0684] Python: Used to implement the program.

[0685] speech_recognition: A library for speech recognition.

[0686] pyttsx3: A text-to-speech engine.

[0687] requests: Used to communicate with the server (calling a dummy server).

[0688] Specific examples

[0689] As an example, if a worker asks, "What are today's inspection procedures?", the voice input means captures the voice and the voice recognition engine converts it into text data. If the emotion recognition means analyzes the worker's tone of voice and facial expression and recognizes that the worker is feeling stressed, the response generation means generates an appropriate response such as "Are you tired? I recommend you take a break," and outputs the response aloud from a speaker via the voice output means.

[0690] Prompt Sentence Examples

[0691] The prompt text is assumed to be as follows:

[0692] "It recognizes the user's emotions and generates responses suggesting a break if they are tired."

[0693] In this way, integrating emotion recognition with personalized responses can improve worker satisfaction and productivity.

[0694] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0695] Step 1:

[0696] Users use their smartphones or tablets to enter information about their personality, hobbies, preferences, and daily behavior patterns.

[0697] Input: User's individual information (e.g., personality, hobbies, preferences, behavioral patterns)

[0698] Specific Actions: The user enters information into the interface and clicks the submit button.

[0699] Output: The input data is temporarily stored.

[0700] Step 2:

[0701] The terminal acquires the input information and transmits it to the profile generating means.

[0702] Input: Saved user personal information

[0703] Specific operation: The device establishes a network connection to send information to the server. When the information transmission is complete, a success message is displayed.

[0704] Output: The server receives the information.

[0705] Step 3:

[0706] The server generates a user profile based on the received information.

[0707] Input: User's personal information sent to the server

[0708] Specific operation: The server executes the profile generation algorithm and stores the user profile in the database. Once the profile generation is complete, it sends a confirmation message to the terminal.

[0709] Output: Generated user profile

[0710] Step 4:

[0711] While working, users can talk to their smartphone or tablet to ask questions or ask for advice, for example, "Tell me about today's inspection procedure."

[0712] Input: User voice input

[0713] Specific operation: When the user speaks, the device's microphone captures the sound.

[0714] Output: Captured audio data

[0715] Step 5:

[0716] The device converts the captured voice data into text data using a voice recognition engine.

[0717] Input: Audio data

[0718] Specific operation: The device uses the speech_recognition library to analyze the voice data and convert it to text.

[0719] Output: Converted text data

[0720] Step 6:

[0721] The terminal transmits the converted text data to the server.

[0722] Input: Text data

[0723] Specific operation: The terminal establishes communication with the server and sends text data over the network.

[0724] Output: The server receives the text data.

[0725] Step 7:

[0726] The server uses the emotion recognition means to analyze the user's emotions.

[0727] Input: Text data and associated audio tone data

[0728] Specific operation: The server uses an emotion recognition engine to analyze the tone of voice and language patterns to determine the user's emotional state.

[0729] Output: Analyzed emotion data (e.g., stress, relaxation, etc.)

[0730] Step 8:

[0731] The server generates an appropriate response based on the profile and emotion data using a response generation means.

[0732] Input: User profile, emotion data, text data

[0733] Specific operation: The server uses the generative AI model to generate a response that matches the user's current state. For example, if the user says "I'm tired," it suggests taking a break.

[0734] Output: The generated response data

[0735] Step 9:

[0736] The server transmits the generated response data to the terminal.

[0737] Input: Response data

[0738] Specific operation: The server establishes a network connection to send data to the terminal and sends response data.

[0739] Output: The terminal receives the response data.

[0740] Step 10:

[0741] The response data received by the terminal is converted into voice using a voice synthesis engine, and a voice response is given to the user.

[0742] Input: Response data

[0743] Specific behavior: The device uses the pyttsx3 library to convert the text data into speech data and responds aloud through the speaker, for example, "Are you tired? I recommend you take a break."

[0744] Output: The user receives the appropriate response via voice.

[0745] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0746] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0747] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0748] [Second embodiment]

[0749] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0750] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0751] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0752] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0753] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0754] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0755] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0756] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0757] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0758] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0759] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0760] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0761] The system of the present invention is designed to provide users with personalized responses to their questions and concerns in everyday life. Specific embodiments for carrying out the present invention will be described below.

[0762] System configuration

[0763] This system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[0764] 1. Input Method

[0765] 2. Profile Generation Method

[0766] 3. Response Generation Method

[0767] 4. Audio output means

[0768] 5. Learning Methods

[0769] 6. Update method

[0770] 7. Authentication Methods

[0771] 8. Voice Recognition Methods

[0772] 9. Means of Transmission

[0773] Program processing flow

[0774] Initial Setup Phase

[0775] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0776] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[0777] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[0778] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[0779] Daily Use Phase

[0780] 1. User: Speaks to their smartphone. Example: "What do you recommend for dinner tonight?"

[0781] 2. Terminal: The voice input module recognizes the voice and converts it into text data, which is then sent to the server.

[0782] 3. Server: The received text data is analyzed using a natural language processing engine to extract the intent. The server then references the user profile and past data to generate the optimal response.

[0783] 4. Server: The generated response is sent to the terminal in text format.

[0784] 5. Terminal: The text data is converted into speech using a speech synthesis engine, and a response is provided to the user. Example: "I think curry would be good for dinner tonight. The ingredients needed are..."

[0785] Specific examples

[0786] Example 1: Menu suggestion

[0787] 1. User: "What should I make today?"

[0788] 2. Device: Converts speech into text and sends it to the server.

[0789] 3. Server: Analyzes the text data and checks the user's past meal history, preferences, and refrigerator inventory information. Then, it generates an appropriate menu.

[0790] 4. Server: Sends the generated menu suggestions in text format to the terminal.

[0791] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Today, I recommend teriyaki chicken."

[0792] Example 2: Advice

[0793] 1. User: "I've been having trouble at work lately. What should I do?"

[0794] 2. Device: Converts speech into text and sends it to the server.

[0795] 3. Server: Analyzes the text data and generates advice based on the user's work situation and past consultation history.

[0796] 4. Server: Sends the generated advice in text format to the terminal.

[0797] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Try to complete your daily tasks little by little, at your own pace, without rushing."

[0798] Advanced Security Settings

[0799] 1. User: Open the security settings from the device settings screen. Select from the setting options of Corneal Identification, Fingerprint Identification, and DNA Identification.

[0800] 2. Device: According to the selected authentication method, the necessary biometric data is collected. The data is stored in a temporary storage area and encrypted for security purposes.

[0801] 3. Server: Receives encrypted biometric data from the device and stores it in a secure database. Builds an authentication system and applies it to each user.

[0802] 4. Terminal: Installs and activates the authentication system received from the server locally, granting the user access whenever authentication is successful.

[0803] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching daily life. In addition, the system's advanced security functions can safely protect the user's personal information.

[0804] The processing flow will be explained below.

[0805] Initial Setup Phase

[0806] Step 1:

[0807] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0808] Step 2:

[0809] Terminal: Temporarily stores personal information entered by the user. Checks the completeness and format of the input data and determines whether the next step can be processed.

[0810] Step 3:

[0811] Terminal: The personal information entered is sent to the server. When sending, the data is encrypted to ensure security.

[0812] Step 4:

[0813] Server: Deserializes the personal information received from the device and stores it in a database. Based on the stored data, an initial profile is generated.

[0814] Step 5:

[0815] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[0816] Step 6:

[0817] Terminal: Saves the initial profile and learning model received from the server locally. Notifies the user that the setup is complete.

[0818] Daily Use Phase

[0819] Step 1:

[0820] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[0821] Step 2:

[0822] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[0823] Step 3:

[0824] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[0825] Step 4:

[0826] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[0827] Step 5:

[0828] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[0829] Step 6:

[0830] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[0831] Advanced Security Settings

[0832] Step 1:

[0833] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[0834] Step 2:

[0835] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[0836] Step 3:

[0837] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[0838] Step 4:

[0839] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[0840] Step 5:

[0841] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[0842] Step 6:

[0843] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[0844] Continuous learning phase

[0845] Step 1:

[0846] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[0847] Step 2:

[0848] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[0849] Step 3:

[0850] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[0851] Step 4:

[0852] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[0853] Step 5:

[0854] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[0855] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[0856] Example 1

[0857] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0858] Current systems that provide personalized responses lack the ability to quickly and accurately process user input and provide appropriate responses to individual questions and concerns in everyday life. Furthermore, continuous learning and profile updates based on user behavioral history are insufficient, resulting in issues with the accuracy and relevance of responses. Furthermore, in terms of security, measures to safely protect personal information are insufficient, and the authentication process is cumbersome, resulting in a lack of user convenience.

[0859] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0860] In this invention, the server includes: input means for inputting a user's personal information, personality, hobbies, preferences, and behavioral patterns; profile generation means for processing the information acquired from the input means; response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; voice output means for returning the responses generated by the response generation means by voice; learning means for recording and learning the user's behavioral history; update means for updating the profile generation means based on the information acquired by the learning means; voice recognition means for analyzing the user's voice and converting it into text data; transmission means for transmitting the text data converted by the voice recognition means to the response generation means; authentication means for performing high-security authentication; and initial learning model generation means for generating an individual learning model using the user profile and storing it on the server. This makes it possible to comprehensively process a user's diverse information to provide highly accurate personalized responses, and to improve the accuracy of responses through continuous learning, thereby realizing a safe and convenient authentication process.

[0861] The "input means" is a means for inputting data such as the user's personal information, personality, hobbies, preferences, and behavioral patterns.

[0862] The "profile generating means" is a means for generating a user profile based on information acquired from the input means.

[0863] The "answer generation means" is a means for generating an appropriate answer to a user's question or concern based on the generated user profile.

[0864] The "audio output means" is a means for returning the response generated by the response generating means to the user by voice.

[0865] The "learning means" is a means for recording the user's behavior history and for the system to continuously learn based on that data.

[0866] The "update means" is a means for updating the profile generation means based on new information acquired by the learning means.

[0867] The "voice recognition means" is a means having the function of analyzing the user's voice and converting it into text data.

[0868] The "transmitting means" is a means for transmitting the text data converted by the speech recognition means to the response generating means.

[0869] "Authentication means" refers to a means for performing an advanced authentication process to ensure user security.

[0870] The "initial learning model generating means" is a means for generating an individual learning model based on a user profile and storing it on a server.

[0871] The "voice synthesis means" is a means for converting the text data generated by the response generation means into voice.

[0872] The system of the present invention is designed to provide users with personalized responses to their everyday questions and concerns. The components and operation of this system will now be described in detail.

[0873] System configuration

[0874] The system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system consists of the following main components:

[0875] 1. Input Method

[0876] A user starts up a new smartphone and inputs personal information such as name, date of birth, personality, blood type, hobbies, and preferences. This information is provided to the system via input means.

[0877] 2. Profile Generation Method

[0878] The terminal stores the user information acquired from the input means in a temporary storage area. After verifying the integrity of the data, it sends it to the server. The server updates the database based on the received data and creates a user profile. Database software (e.g., MySQL or PostgreSQL) is used to generate the profile.

[0879] 3. Response Generation Method

[0880] Once the profile is created, the system generates responses to the user's questions and concerns. This response generation process uses a natural language processing engine (for example, IBM Watson's NLP capabilities) to generate the best possible response based on the user profile and past data.

[0881] 4. Audio output means

[0882] The generated response is sent in text format to the user device, which then uses a speech synthesis engine (e.g., Microsoft Azure's Text-to-Speech service) to convert the text into speech and provide the response to the user.

[0883] 5. Learning Methods

[0884] The system records the user's behavior history and continuously learns from it. This learning process is implemented using a machine learning framework (e.g., TensorFlow).

[0885] 6. Update method

[0886] The information obtained by the learning means is reflected in the profile generation means as appropriate, and the profile is updated, thereby improving the accuracy of the system over time.

[0887] 7. Voice Recognition Methods

[0888] When a user speaks a question, the device's voice input module recognizes the speech and converts it into text data. This process uses a speech recognition API (for example, Google's Speech-to-Text API).

[0889] 8. Means of Transmission

[0890] The text data converted by the speech recognition means is sent to the server via the transmission means.

[0891] 9. Authentication Methods

[0892] The system utilizes biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide a high level of security, ensuring that users' personal information is kept safe.

[0893] 10. Initial learning model generation method

[0894] The server generates an initial learning model based on the user profile and stores it on the server. This initial learning model is generated using a machine learning framework such as TensorFlow.

[0895] Specific examples

[0896] Example 1: Menu suggestion

[0897] When a user speaks to their smartphone, "What should I make today?", the voice input module recognizes the speech and converts it into text data. This text data is sent to the server and analyzed by a natural language processing engine. An appropriate menu is generated by referring to the user's past meal history, preferences, and refrigerator inventory information. The generated menu is sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and suggested to the user.

[0898] Example 2: Advice

[0899] When a user asks, "Work hasn't been going well lately. What should I do?", the voice input module converts the voice into text data. This data is sent to the server, where a natural language processing engine analyzes the user's work situation and past consultation history. The most appropriate advice is generated and sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and conveyed to the user.

[0900] AI chatbot-generated prompts

[0901] Example prompt 1: "Generate a response when a user asks verbally, 'What's for dinner tonight?'"

[0902] Example prompt 2: "Generate advice for when a user says, 'I've been having trouble at work lately. What should I do?'"

[0903] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching the user's daily life and protecting it safely.

[0904] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0905] Step 1:

[0906] The user starts up the new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[0907] Input: Personal information that users enter into input forms.

[0908] Output: The personal information data entered.

[0909] Step 2:

[0910] The device receives user input and stores the entered data in a temporary storage area (the smartphone's RAM). At this point, the data is verified for integrity and sent to the server. The data is encrypted.

[0911] Input: Personal information data provided by the user.

[0912] Data processing: Data validation and encryption (e.g., AES encryption).

[0913] Output: The encrypted personal information data is sent to the server.

[0914] Step 3:

[0915] The server receives the data sent from the device and stores it in a database (MySQL, PostgreSQL, etc.). It generates a user profile based on the received initial data and creates an initial learning model.

[0916] Input: Encrypted personal information data.

[0917] Data processing: Decrypting data and saving it to the database, generating user profiles (executing SQL queries).

[0918] Output: An initial learning model and a user profile are generated and stored in a database.

[0919] Step 4:

[0920] The device receives the initial learning model from the server and saves it locally (in the smartphone's internal storage). A message is displayed to the user indicating that the setup is complete.

[0921] Input: The initial training model sent from the server.

[0922] Data processing: Saving the initial training model.

[0923] Output: The initial training model is saved locally and a message appears saying that the setup is complete.

[0924] Step 5:

[0925] The user speaks to the smartphone using voice. For example, "What do you recommend for dinner tonight?"

[0926] Input: User's voice input.

[0927] Step 6:

[0928] The device's voice input module recognizes speech and converts it into text data. This process uses a speech recognition API (such as Google's Speech-to-Text API). The converted text data is then sent to the server.

[0929] Input: Audio data.

[0930] Data processing: Recognizing voice data and converting it to text.

[0931] Output: Text data is sent to the server.

[0932] Step 7:

[0933] The server analyzes the received text data using a natural language processing engine (such as IBM Watson's NLP function) to extract the intent, and generates the optimal response by referencing the user profile and past data.

[0934] Input: Text data.

[0935] Data processing: Natural language processing, intent extraction, and response generation.

[0936] Output: The generated response text data.

[0937] Step 8:

[0938] The server sends the generated response in text format to the terminal.

[0939] Input: The generated response text data.

[0940] Output: The response text data is sent to the terminal.

[0941] Step 9:

[0942] The device uses a speech synthesis engine (such as Microsoft Azure's Text-to-Speech service) to convert the text data into speech and provide a response to the user. For example: "I think curry would be good for dinner tonight. The ingredients you need are..."

[0943] Input: Response text data.

[0944] Data processing: Converting text data into audio.

[0945] Output: A spoken response.

[0946] Step 10:

[0947] The learning method records the user's behavior history and the system continuously learns based on that data. This is implemented using a machine learning framework (such as TensorFlow).

[0948] Input: User behavior history data.

[0949] Data processing: Preprocessing data and training machine learning models.

[0950] Output: The trained model and updated model parameters.

[0951] Step 11:

[0952] The server updates the profile generation means based on the information obtained by the learning means, thereby improving the profile and increasing the accuracy of responses.

[0953] Input: trained model parameters.

[0954] Data processing: Profile update work.

[0955] Output: The updated user profile.

[0956] Step 12:

[0957] The authentication method uses biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide high-security authentication.

[0958] Input: Biometric data.

[0959] Data Processing: Biometric data collection and authentication.

[0960] Output: Authentication result.

[0961] Step 13:

[0962] The initial learning model generating means generates an individual learning model based on the user profile and stores it in the server.

[0963] Input: User profile data.

[0964] Data processing: generating and saving learning models.

[0965] Output: Initial training model.

[0966] (Application example 1)

[0967] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0968] Customer service in traditional brick-and-mortar stores relies on the knowledge and experience of store staff, making it difficult to provide consistent, high-quality service to customers. In particular, responses to customer questions and concerns are not personalized, which leads to lower customer satisfaction. It is also difficult to accurately grasp daily customer behavior patterns and preferences, making it difficult to recommend optimal products. Furthermore, there are also issues with in-store security authentication, and personal information may not be adequately protected.

[0969] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0970] In this invention, the server includes: an input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns; a profile generation means for processing the information acquired from the input means; a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; a voice output means for returning the responses generated by the response generation means by voice; a learning means for recording and learning the user's behavioral history; an update means for updating the profile generation means based on the information acquired by the learning means; an authentication means for performing high-security authentication; a voice recognition means for analyzing the voice input and converting it into text data; a natural language processing means for analyzing the text data using a natural language processing engine and generating an optimal response by referring to the user's profile and past data; a voice synthesis means for converting the text data generated by the natural language processing means into voice; a robot that operates as a customer service assistant to support operations in a physical store; and a profile generation means for suggesting products in the physical store. This enables personalized responses to customer questions and concerns, thereby contributing not only to improved customer satisfaction but also to improved security.

[0971] "Input means" refers to a device or function for inputting the user's personality, hobbies, preferences, and daily behavior patterns.

[0972] The "profile generating means" refers to a device or function for processing information acquired from the input means and generating a user profile.

[0973] The "response generating means" refers to a device or function for responding to a user's questions or concerns based on the profile generated by the profile generating means.

[0974] The "audio output means" refers to a device or function for returning the response generated by the response generation means by voice.

[0975] The "learning means" refers to a device or function for recording and learning from the user's behavior history.

[0976] The "update means" refers to a device or function for updating the profile generation means based on the information acquired by the learning means.

[0977] "Authentication means" means a device or function for performing high-security authentication.

[0978] "Speech recognition means" refers to a device or function for analyzing voice input and converting it into text data.

[0979] "Natural language processing means" refers to a device or function that uses a natural language processing engine to analyze text data and generate optimal responses by referring to the user's profile and past data.

[0980] The "speech synthesis means" refers to a device or function for converting text data generated by the natural language processing means into speech.

[0981] A "customer service assistant" is a robot that assists with work in physical stores.

[0982] The "profile generation means for making product suggestions" refers to a device or function that generates a profile for making product suggestions in a physical store.

[0983] The system of the present invention aims to provide personalized responses to questions and concerns that users have about their daily lives in brick-and-mortar stores. Specific embodiments for carrying out the invention will now be described.

[0984] System configuration

[0985] The system consists of the following major components:

[0986] 1. Input Method

[0987] 2. Profile Generation Method

[0988] 3. Response Generation Method

[0989] 4. Audio output means

[0990] 5. Learning Methods

[0991] 6. Update method

[0992] 7. Authentication Methods

[0993] 8. Voice Recognition Methods

[0994] 9. Natural Language Processing Tools

[0995] 10. Speech synthesis means

[0996] 11. Customer Service Assistant

[0997] 12. Product proposal profile generation means

[0998] Details of each means are as follows.

[0999] Program Generation and Hardware / Software Usage

[1000] 1. Initial Setup Phase

[1001] User: Launches the application at the store entrance and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[1002] Terminal: The entered information is temporarily stored, encrypted, and sent to the server.

[1003] Server: Generates a user profile based on the received data and creates a personalized product recommendation model.

[1004] Terminal: Saves the initial model received from the server locally and displays a notification that the setup is complete.

[1005] 2. Operational Phase

[1006] User: Asks questions by voice in-store. Example: "What wines do you recommend these days?"

[1007] Terminal: The speech is converted into text using a speech recognition tool (Google Speech-to-Text API) and sent to the server.

[1008] Server: Analyzes the text data using a natural language processing engine (spaCy or BERT) and generates an appropriate response by referencing the user's profile and past data.

[1009] Server: Sends the generated response in text format to the terminal.

[1010] Terminal: Converts text to speech using a speech synthesis engine (Amazon Polly) and provides responses to the user.

[1011] Specific examples

[1012] Example 1

[1013] User: "What wines have you recommended recently?"

[1014] Device: Converts speech to text and sends it to the server.

[1015] Server: Generates a response based on the user's profile and past data, and sends text data to the device saying, "This time, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[1016] Terminal: A speech synthesis engine is used to convert text into speech and provide guidance to the user.

[1017] Prompt Sentence Examples

[1018] Please enter your name, gender, age, preferences, purchase history, and product interest information.

[1019] What wines do you recommend these days?

[1020] "Today, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[1021] Hardware and software used

[1022] Hardware: Smartphones, smart glasses, in-store robots

[1023] software:

[1024] Speech recognition engine: Google Speech-to-Text API

[1025] Natural language processing engine: spaCy, BERT-based model

[1026] Speech synthesis engine: Amazon Polly

[1027] Database: MySQL, Firebase

[1028] Communication protocol: HTTPS, WebSocket

[1029] In this way, the system of the present invention can provide personalized responses to users' questions and concerns, thereby improving customer satisfaction in brick-and-mortar stores. Furthermore, advanced security features can protect users' personal information.

[1030] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1031] Step 1:

[1032] User: Launches the application at the entrance of a physical store and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[1033] Input: Personal information

[1034] Output: Temporarily saved personal information

[1035] Specific operation: By entering information into the input form and pressing the "Submit" button, the data is stored in a temporary storage area on the device.

[1036] Step 2:

[1037] Terminal: The personal information entered is temporarily stored, encrypted for security purposes, and then sent to the server.

[1038] Input:Temporarily saved personal information

[1039] Output: Encrypted personal information data

[1040] Specific operation: Personal information is encrypted inside the device and sent to the server using the HTTPS communication protocol.

[1041] Step 3:

[1042] Server: Decrypts the received encrypted data and stores it in a database. Generates a user profile and creates a personalized product recommendation model.

[1043] Input: Encrypted personal information data

[1044] Output: User profile and proposed model

[1045] Specific operations: Decrypting encrypted data, running the user profile generation algorithm, and saving the generated profile and proposed model to a database.

[1046] Step 4:

[1047] Terminal: The terminal locally stores the user profile and proposed model received from the server and displays a notification to the user that the setup is complete.

[1048] Input: User profile and proposed model data

[1049] Output: Locally saved profile and proposed model, notification that setup is complete

[1050] Specific behavior: Save received data locally, display completion notification

[1051] Step 5:

[1052] User: Asks questions by dictating as they walk through the store. Example: "What wines do you recommend these days?"

[1053] Input: Voice data (question)

[1054] Output: Start speech recognition via speech input interface

[1055] Specific actions: Press the voice input button on your smartphone or smart glasses and dictate your question.

[1056] Step 6:

[1057] Terminal: Using a voice recognition tool (Google Speech-to-Text API), the voice data is converted into text data and sent to the server.

[1058] Input: Audio data

[1059] Output: Text data

[1060] Specific operations: Calling the speech recognition API, generating converted text data, and sending the text data to the server.

[1061] Step 7:

[1062] Server: Analyzes the text data using natural language processing tools (spaCy or BERT engine) and generates an appropriate response by referencing the user's profile and past data.

[1063] Input: Text data, user profile

[1064] Output: Response text data

[1065] What it does: Runs a natural language analysis engine, consults a user profile database, and generates the best possible response.

[1066] Step 8:

[1067] Server: Sends the generated response in text format to the terminal.

[1068] Input: Response text data

[1069] Output: Text data sent to the terminal

[1070] Specific operation: Processing to send text data to the terminal.

[1071] Step 9:

[1072] Terminal: The response text data is converted into speech using a speech synthesis engine (Amazon Polly) and a response is provided to the user.

[1073] Input: Response text data

[1074] Output: Audio data

[1075] Specific operations: Calling the speech synthesis API, generating converted speech data, and outputting speech from the speaker.

[1076] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1077] The system of the present invention is designed in combination with an emotion engine to provide personalized responses to users' questions and concerns in their daily lives. Specific embodiments for implementing the present invention will be described below.

[1078] System configuration

[1079] This system generates an individual profile based on the user's input and emotional information, and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[1080] 1. Input Method

[1081] 2. Profile Generation Method

[1082] 3. Response Generation Method

[1083] 4. Audio output means

[1084] 5. Learning Methods

[1085] 6. Update method

[1086] 7. Authentication Methods

[1087] 8. Voice Recognition Methods

[1088] 9. Means of Transmission

[1089] 10. Emotion Recognition Engine

[1090] Program processing flow

[1091] Initial Setup Phase

[1092] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1093] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[1094] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[1095] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[1096] Daily Use Phase

[1097] 1. User: Talks to the smartphone and asks questions or asks for advice. For example, "What do you recommend for dinner tonight?"

[1098] 2. Terminal: The voice input module captures the spoken voice and converts it into text data using the voice recognition engine.

[1099] 3. Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[1100] 4. Server: The received text data is analyzed using a natural language processing engine to extract the user's intent. The server then generates the optimal response data by referencing the user profile and past data.

[1101] 5. Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[1102] 6. Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it may reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1103] Processing using an emotion recognition engine

[1104] 1. User: Speaks, makes facial expressions and shows movements.

[1105] 2. On the device: The emotion recognition engine detects the user's tone of voice, facial expressions, movements, and language patterns to analyze emotions in real time.

[1106] 3. Terminal: Temporarily stores the analyzed emotion data and sends it to the server.

[1107] 4. Server: Receives emotion data and uses it in the profile generation means and response generation means.

[1108] 5. Server: Based on the emotional data, generate a personalized response that is more suited to the user's current situation.

[1109] 6. Device: Provides the user with a voice response based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1110] Advanced Security Settings

[1111] 1. User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[1112] 2. Terminal: According to the selected authentication method, the necessary biometric authentication data is acquired, temporarily stored, and encrypted.

[1113] 3. Terminal: Sends encrypted authentication data to the server, which checks the data for consistency and integrity upon transmission.

[1114] 4. Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[1115] 5. Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[1116] 6. Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[1117] Continuous learning phase

[1118] 1. User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[1119] 2. Device: Automatically collects user behavior history and usage data, temporarily stores the collected data, and sends it to the server as needed.

[1120] 3. Server: Analyzes the received behavioral history data and updates the user profile. Machine learning algorithms are used to learn the user's behavioral patterns.

[1121] 4. Server: Periodically synchronizes updated profiles and learning results to the device. During synchronization, it checks the consistency and integrity of the data.

[1122] 5. Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[1123] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively. The present invention can provide personalized services based on the user's individual information and emotional information, enriching daily life, and safely protecting the user's personal information with advanced security functions.

[1124] The processing flow will be explained below.

[1125] Initial Setup Phase

[1126] Step 1:

[1127] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1128] Step 2:

[1129] Terminal: Temporarily stores personal information entered by the user, checks the integrity and format of the entered data, and sends it to the server.

[1130] Step 3:

[1131] Server: Receives data sent from the device and stores it in a database. Based on the stored data, it generates an initial profile.

[1132] Step 4:

[1133] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[1134] Step 5:

[1135] Terminal: The initial profile and learning model received from the server are saved locally and the user is notified that the setup is complete.

[1136] Daily Use Phase

[1137] Question and answer processing

[1138] Step 1:

[1139] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[1140] Step 2:

[1141] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[1142] Step 3:

[1143] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[1144] Step 4:

[1145] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[1146] Step 5:

[1147] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[1148] Step 6:

[1149] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1150] Emotion Recognition Processing

[1151] Step 1:

[1152] User: Speaks, shows facial expressions and actions.

[1153] Step 2:

[1154] On the device: The emotion recognition engine detects the user's tone of voice, facial expressions, movements, and language patterns to analyze emotions in real time.

[1155] Step 3:

[1156] Terminal: Temporarily stores the analyzed emotion data and sends it to the server.

[1157] Step 4:

[1158] Server: Receives emotion data and uses it in the profile generation means and response generation means.

[1159] Step 5:

[1160] Server: Based on the emotional data, it generates a personalized response that is more suited to the user's current situation.

[1161] Step 6:

[1162] The device provides the user with a voice response based on the generated emotion. For example, if the device detects that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1163] Advanced Security Settings

[1164] Step 1:

[1165] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[1166] Step 2:

[1167] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[1168] Step 3:

[1169] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[1170] Step 4:

[1171] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[1172] Step 5:

[1173] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[1174] Step 6:

[1175] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[1176] Continuous learning phase

[1177] Step 1:

[1178] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[1179] Step 2:

[1180] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[1181] Step 3:

[1182] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[1183] Step 4:

[1184] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[1185] Step 5:

[1186] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[1187] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[1188] Example 2

[1189] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1190] In modern society, people often have many questions and concerns in their daily lives, requiring efficient and personalized responses. Furthermore, it is necessary to provide services that provide higher levels of satisfaction by providing responses that are tailored to the user's emotions and circumstances. However, current systems perform tasks such as user profile creation and updating, voice recognition, emotion recognition, and advanced security authentication separately, resulting in a lack of an integrated and efficient solution.

[1191] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1192] In this invention, the server includes input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns, profile generation means for processing information acquired from the input means, response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means, voice output means for returning the responses generated by the response generation means by voice, learning means for recording and learning the user's behavioral history, update means for updating the profile generation means based on information acquired by the learning means, voice recognition means for analyzing the user's voice and converting it into text data, transmission means for transmitting the text data converted by the voice recognition means to the response generation means, an emotion recognition engine for analyzing the user's emotions, means for personalizing responses based on the emotion data acquired by the emotion recognition engine, and authentication means for performing high-security authentication. This enables personalized responses that reflect the user's situation and emotions.

[1193] 1. "Input means" refers to a device or function that allows a user to input personal information such as personality, hobbies, preferences, and daily behavior patterns into the system.

[1194] 2. "Profile generation means" refers to a device or function that generates a user profile based on information acquired from an input means.

[1195] 3. "Response generation means" refers to a device or function that generates the most appropriate response to a user's question or concern based on the generated user profile.

[1196] 4. "Audio output means" refers to a device or function that provides the response generated by the response generation means to the user by voice.

[1197] 5. "Learning means" refers to a device or function that records the user's behavior history and uses it to improve the system's performance.

[1198] 6. "Update means" refers to a device or function that automatically updates the profile generation means based on information acquired by the learning means.

[1199] 7. "Speech recognition means" means a device or function that captures a user's voice and converts it into text data.

[1200] 8. "Transmitting means" means a device or function that transmits text data converted by the speech recognition means to the response generating means.

[1201] 9. "Emotion recognition engine" means a device or function that analyzes a user's emotions and acquires that data.

[1202] 10. "Personalization means" refers to a device or function that changes responses to suit the user's current situation based on emotional data obtained by an emotion recognition engine.

[1203] 11. "Authentication Means" means a device or function that verifies a user's identity and performs high-security authentication.

[1204] 12. A “generative AI model” is an artificial intelligence-based algorithm that generates appropriate responses to user intent or questions.

[1205] 13. “Prompt sentence” refers to the text data input into a generative AI model, which is used to specifically indicate the user’s question or intention.

[1206] The present invention is a system designed to provide users with personalized responses to questions and concerns they have in their daily lives. This system is composed of multiple components that generate an individual profile based on the user's input information and emotional information, and return appropriate responses. Specific embodiments for implementing the present invention will be described below.

[1207] Hardware and software used

[1208] 1. Terminal: A user device such as a smartphone. This terminal includes a voice input module, a voice recognition engine (e.g., Google Speech-to-Text API), and a voice synthesis engine (e.g., Google Text-to-Speech API).

[1209] 2. Server: A server running a database system (e.g., MySQL or PostgreSQL), a natural language processing engine (e.g., GPT-3), or an emotion recognition engine (e.g., Microsoft Azure Emotion API).

[1210] 3. Generative AI model: An artificial intelligence-based algorithm for analyzing user intent and generating appropriate responses.

[1211] Detailed Description of the Invention

[1212] Initial Setup

[1213] 1. The user starts up a new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1214] 2. The terminal receives the data entered by the user, temporarily stores it in a storage area, verifies the integrity of the data, and then sends it to the server.

[1215] 3. The server stores the received data in a database and generates a user profile, which is then used to create an initial learning model.

[1216] 4. The device receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[1217] daily use

[1218] 1. The user speaks to their smartphone to ask a question or ask for advice, such as, "What do you recommend for dinner tonight?"

[1219] 2. The device captures the voice using the voice input module and converts it into text data using the voice recognition engine.

[1220] 3. The terminal temporarily stores the converted text data, checks the data for consistency, and then sends it to the server.

[1221] 4. The server analyzes the received text data using a natural language processing engine to extract the user's intent. It then generates the optimal response data by referencing the user profile and past data.

[1222] 5. The server sends the generated response data to the terminal.

[1223] 6. The device converts the response data into an audio file using a speech synthesis engine and provides the user with a spoken response. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1224] emotion recognition

[1225] 1. The user speaks, uses facial expressions, and makes movements.

[1226] 2. The device uses an emotion recognition engine to analyze the user's emotions in real time.

[1227] 3. The device temporarily stores the analyzed emotion data and sends it to the server.

[1228] 4. The server receives the emotion data and uses it in the profile generation means and response generation means.

[1229] 5. The server generates a personalized response based on the emotion data that better suits the user's current situation.

[1230] 6. The device then provides a voice response to the user based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1231] Security Settings

[1232] 1. The user selects security settings from the device settings menu and chooses an authentication option such as corneal identification, fingerprint authentication, or DNA authentication.

[1233] 2. The device captures and encrypts the biometric data.

[1234] 3. The device sends the encrypted data to the server, which verifies the data's integrity and completeness.

[1235] 4. The server deserializes the encrypted data and stores it in a secure database. It builds an authentication algorithm and associates it with the user profile.

[1236] 5. The server sends the authentication system to the terminal, which installs and configures the system within the terminal.

[1237] 6. The terminal has a security authentication system installed, and grants the user access whenever authentication is successful.

[1238] Continuous learning

[1239] 1. Users use smartphones on a daily basis and use various apps and functions.

[1240] 2. The device automatically collects the user's behavioral history and usage data, temporarily stores it, and sends it to the server.

[1241] 3. The server analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[1242] 4. The server periodically synchronizes the updated profile and learning results to the device, verifying the consistency and completeness of the data.

[1243] 5. The device locally reflects the received update data, providing the user with more personalized services.

[1244] Examples of specific examples and prompts

[1245] Specific examples

[1246] A user asks, "What's the best way to relax after work?"

[1247] The device uses a voice recognition engine to convert the question into text and send it to the server.

[1248] The server analyzes the text, understands the user's intent, and generates optimal advice on how to relax.

[1249] The server sends the generated advice to the terminal, which converts it into an audio file and responds to the user by saying, "Why don't you take a nice, long bath?"

[1250] Prompt Sentence Examples

[1251] "Generate the best response when a user asks, 'What's the best way to relax after work?'"

[1252] "What relaxation methods can we suggest to users when they are feeling stressed?"

[1253] The above is an embodiment of the present invention. This system provides personalized responses based on the user's individual information and emotional information, enriching daily life and ensuring safety.

[1254] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1255] Initial Setup Phase

[1256] Step 1:

[1257] A user starts up their smartphone and enters personal information such as their name, date of birth, personality, blood type, hobbies, etc. The data entered is the user's basic information, such as "Taro Tanaka, May 4, 1990, blood type O, watching movies."

[1258] Input: Personal information (name, date of birth, personality, blood type, hobbies and preferences)

[1259] Output: Personal information data entered

[1260] Step 2:

[1261] The terminal receives the data entered by the user and temporarily stores it in a storage area. At this time, the terminal checks the integrity of the data and monitors whether there are any missing data.

[1262] Input: Personal data entered by the user

[1263] Output: Temporarily saved data

[1264] Step 3:

[1265] The data that has been confirmed by the device is sent to the server using the HTTP protocol, and the data is encrypted.

[1266] Input: Temporarily saved personal information data

[1267] Output: Data sent to the server

[1268] Step 4:

[1269] The server receives the data and stores it in a database (e.g., MySQL or PostgreSQL). It generates a user profile based on the initial data and creates an initial learning model (e.g., TensorFlow or PyTorch).

[1270] Input: Received personal information data

[1271] Output: Generated user profile and initial learning model

[1272] Step 5:

[1273] The server sends the generated initial learning model to the terminal via the endpoint.

[1274] Input: Initial training model

[1275] Output: The trained model sent to the device.

[1276] Step 6:

[1277] The device saves the initial learning model locally and displays a message to the user indicating that setup is complete.

[1278] Input: The received training model

[1279] Output: Locally saved training model and a message saying the setup is complete

[1280] Daily Use Phase

[1281] Step 1:

[1282] Users can ask questions or ask their smartphones by voice, for example, "What do you recommend for dinner tonight?"

[1283] Input: Voice question

[1284] Output: Audio data

[1285] Step 2:

[1286] The device captures audio using the built-in microphone and converts it into text data using a speech recognition engine (e.g., Google Speech-to-Text API).

[1287] Input: Audio data

[1288] Output: Text data

[1289] Step 3:

[1290] The terminal stores the converted text data in a temporary storage area, checks the data for consistency, and then transmits it to the server.

[1291] Input: Text data

[1292] Output: Text data to send to the server

[1293] Step 4:

[1294] The server receives the text data, analyzes it with a natural language processing engine (e.g., GPT-3), extracts the user's intent, and generates the optimal response by referring to the user's profile and past data.

[1295] Input: Received text data

[1296] Output: The generated response data

[1297] Step 5:

[1298] The server transmits the generated response data in text format to the terminal.

[1299] Input: Generated response data

[1300] Output: Response data sent to the terminal

[1301] Step 6:

[1302] The device converts the received response data into an audio file using a speech synthesis engine (for example, Google Text-to-Speech API) and provides a voice response to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1303] Input: Received response data

[1304] Output: A spoken response provided to the user

[1305] Processing using an emotion recognition engine

[1306] Step 1:

[1307] The user speaks, makes facial expressions, and makes movements.

[1308] Input: voice, facial expressions, movements

[1309] Output: Expressed emotion data

[1310] Step 2:

[1311] The device analyzes the user's emotions in real time using an emotion recognition engine (for example, Microsoft Azure Emotion API).

[1312] Input: Expressed emotion data

[1313] Output: Parsed emotion data

[1314] Step 3:

[1315] The device temporarily stores the analyzed emotion data and transmits it to the server.

[1316] Input: Parsed emotion data

[1317] Output: Emotion data sent to the server

[1318] Step 4:

[1319] The server receives the emotion data and uses it in the profile generation means and response generation means.

[1320] Input: Received emotion data

[1321] Output: Emotion data used in the profile generation and response generation methods

[1322] Step 5:

[1323] The server generates a personalized response based on the emotional data that better suits the user's current situation.

[1324] Input: Emotion data

[1325] Output: Personalized response data

[1326] Step 6:

[1327] The device will provide a voice response to the user based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1328] Input: Emotion-based response data

[1329] Output: Spoken suggestions

[1330] Advanced Security Settings

[1331] Step 1:

[1332] Users select security settings from the device's settings menu and choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[1333] Input: Select user authentication settings

[1334] Output: The configured authentication options

[1335] Step 2:

[1336] The terminal acquires and encrypts the biometric authentication data according to the selected authentication method.

[1337] Input: Authentication method and biometric data

[1338] Output: Encrypted biometric data

[1339] Step 3:

[1340] The device sends the encrypted data to the server, which verifies the data's integrity and completeness.

[1341] Input: Encrypted biometric data

[1342] Output: Encrypted data sent to the server

[1343] Step 4:

[1344] The server deserializes the encrypted data and stores it in a secure database, building an authentication algorithm and associating it with the user profile.

[1345] Input: Encrypted biometric data

[1346] Output: User profile associated with constructed authentication algorithm

[1347] Step 5:

[1348] The server sends the authentication system to the terminal, which then installs and configures the system within the terminal.

[1349] Input:AuthenticationSystem

[1350] output: Authentication system configured on the device

[1351] Step 6:

[1352] The terminal installs an authentication system and grants the user access whenever authentication is successful.

[1353] Input:AuthenticationSystem

[1354] Output: Access permission granted upon successful authentication

[1355] Continuous learning phase

[1356] Step 1:

[1357] Users use smartphones on a daily basis and utilize various applications and functions.

[1358] Input: Everyday smartphone use

[1359] Output: Application and feature usage data

[1360] Step 2:

[1361] The device automatically collects the user's behavioral history and usage data, temporarily stores it, and sends it to the server.

[1362] Input: Application and feature usage data

[1363] Output: Behavioral history data sent to the server

[1364] Step 3:

[1365] The server analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[1366] Input: Received behavioral history data

[1367] Output: Updated user profile

[1368] Step 4:

[1369] The server periodically synchronizes updated profiles and learning results with the device and checks the consistency and completeness of the data.

[1370] Input: Updated user profile

[1371] Output: Data synced to the device

[1372] Step 5:

[1373] The terminal locally reflects the updated data received from the server, thereby providing the user with a more personalized service.

[1374] Input: Update data synced to the device

[1375] Output: Personalized service

[1376] The above is a description of the specific operations in the processing steps of the present invention. Through this processing, the system provides personalized responses based on the user's individual information and emotional information, realizing a system that can enrich daily life and ensure safety.

[1377] (Application example 2)

[1378] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1379] Current worker support systems in factories provide standard responses without considering the worker's emotions, resulting in insufficient stress management and optimization of work efficiency. Furthermore, when workers ask questions or ask for advice via voice, appropriate responses that take their emotions into consideration are not provided, which can lead to lower worker satisfaction and impact productivity. There is a growing need for a system that can solve these issues and provide responses that take the worker's emotional state into consideration.

[1380] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for inputting the user's personality, hobbies, preferences, and daily behavioral patterns, a profile generation means for processing information acquired from the input means, and a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means. This makes it possible to provide accurate and personalized responses while taking into consideration the user's emotional state.

[1381] The "input means" is a device or interface for inputting information such as the user's personality, hobbies, preferences, and daily behavior patterns.

[1382] The "profile generating means" is a device or program for processing information acquired from the input means and generating an individual profile for the user.

[1383] The "answer generation means" is a device or program for generating appropriate answers to the user's questions and concerns based on the generated user profile.

[1384] The "audio output means" refers to a device or technology for returning the response generated by the response generation means by voice.

[1385] A "learning means" is a device or algorithm that records a user's behavioral history and uses it to train a profile or system.

[1386] The "update means" is a device or technology for updating the profile generation means based on the information acquired by the learning means.

[1387] An "authentication means" is a device or system for performing high-security authentication.

[1388] "Voice input means" refers to a device or technology for capturing spoken voice and converting it into text data using a voice recognition engine.

[1389] "Emotion recognition means" refers to a device or technology for analyzing the user's emotions by analyzing the tone of voice and facial expressions.

[1390] System configuration

[1391] This invention is a "smart industrial assistant" system that supports workers in factories. The system consists of the following main components:

[1392] 1. Input means: An interface for inputting information such as the worker's personality, hobbies and preferences, and daily behavioral patterns, and is typically a smartphone or tablet.

[1393] 2. Profile generation means: A program for processing information obtained from the input means and generating a profile for each worker.

[1394] 3. Response generation means: A program that generates appropriate responses to the worker's questions and concerns based on the generated profile.

[1395] 4. Audio output means: This is a technology for outputting the response generated by the response generation means by voice, and a speaker is used.

[1396] 5. Learning method: An algorithm that records the worker's behavioral history and uses it to train the profile and system.

[1397] 6. Update method: This is a technology that updates the profile generation method based on the information obtained by the learning method.

[1398] 7. Authentication method: A device or system for performing high-security authentication, and options include corneal identification, fingerprint authentication, and DNA authentication.

[1399] 8. Voice input method: This technology captures the voice spoken by the worker and converts it into text data using a voice recognition engine.

[1400] 9. Emotion recognition: This technology analyzes the emotions of workers by analyzing their tone of voice and facial expressions.

[1401] Program processing explanation

[1402] The server uses the Python language, the speech_recognition library for speech recognition, and the pyttsx3 library for speech synthesis. This allows the server to recognize the worker's voice, convert it into text data, and generate an appropriate response based on the worker's profile and emotional information. The generated response is output as audio through the speaker.

[1403] Hardware and software used

[1404] Hardware:

[1405] Microphone: Used to capture audio input.

[1406] Speaker: Used to output responses aloud.

[1407] Robot terminal: Used as an installation platform.

[1408] software:

[1409] Python: Used to implement the program.

[1410] speech_recognition: A library for speech recognition.

[1411] pyttsx3: A text-to-speech engine.

[1412] requests: Used to communicate with the server (calling a dummy server).

[1413] Specific examples

[1414] As an example, if a worker asks, "What are today's inspection procedures?", the voice input means captures the voice and the voice recognition engine converts it into text data. If the emotion recognition means analyzes the worker's tone of voice and facial expression and recognizes that the worker is feeling stressed, the response generation means generates an appropriate response such as "Are you tired? I recommend you take a break," and outputs the response aloud from a speaker via the voice output means.

[1415] Prompt Sentence Examples

[1416] The prompt text is assumed to be as follows:

[1417] "It recognizes the user's emotions and generates responses suggesting a break if they are tired."

[1418] In this way, integrating emotion recognition with personalized responses can improve worker satisfaction and productivity.

[1419] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1420] Step 1:

[1421] Users use their smartphones or tablets to enter information about their personality, hobbies, preferences, and daily behavior patterns.

[1422] Input: User's individual information (e.g., personality, hobbies, preferences, behavioral patterns)

[1423] Specific Actions: The user enters information into the interface and clicks the submit button.

[1424] Output: The input data is temporarily stored.

[1425] Step 2:

[1426] The terminal acquires the input information and transmits it to the profile generating means.

[1427] Input: Saved user personal information

[1428] Specific operation: The device establishes a network connection to send information to the server. When the information transmission is complete, a success message is displayed.

[1429] Output: The server receives the information.

[1430] Step 3:

[1431] The server generates a user profile based on the received information.

[1432] Input: User's personal information sent to the server

[1433] Specific operation: The server executes the profile generation algorithm and stores the user profile in the database. Once the profile generation is complete, it sends a confirmation message to the terminal.

[1434] Output: Generated user profile

[1435] Step 4:

[1436] While working, users can talk to their smartphone or tablet to ask questions or ask for advice, for example, "Tell me about today's inspection procedure."

[1437] Input: User voice input

[1438] Specific operation: When the user speaks, the device's microphone captures the sound.

[1439] Output: Captured audio data

[1440] Step 5:

[1441] The device converts the captured voice data into text data using a voice recognition engine.

[1442] Input: Audio data

[1443] Specific operation: The device uses the speech_recognition library to analyze the voice data and convert it to text.

[1444] Output: Converted text data

[1445] Step 6:

[1446] The terminal transmits the converted text data to the server.

[1447] Input: Text data

[1448] Specific operation: The terminal establishes communication with the server and sends text data over the network.

[1449] Output: The server receives the text data.

[1450] Step 7:

[1451] The server uses the emotion recognition means to analyze the user's emotions.

[1452] Input: Text data and associated audio tone data

[1453] Specific operation: The server uses an emotion recognition engine to analyze the tone of voice and language patterns to determine the user's emotional state.

[1454] Output: Analyzed emotion data (e.g., stress, relaxation, etc.)

[1455] Step 8:

[1456] The server generates an appropriate response based on the profile and emotion data using a response generation means.

[1457] Input: User profile, emotion data, text data

[1458] Specific operation: The server uses the generative AI model to generate a response that matches the user's current state. For example, if the user says "I'm tired," it suggests taking a break.

[1459] Output: The generated response data

[1460] Step 9:

[1461] The server transmits the generated response data to the terminal.

[1462] Input: Response data

[1463] Specific operation: The server establishes a network connection to send data to the terminal and sends response data.

[1464] Output: The terminal receives the response data.

[1465] Step 10:

[1466] The response data received by the terminal is converted into voice using a voice synthesis engine, and a voice response is given to the user.

[1467] Input: Response data

[1468] Specific behavior: The device uses the pyttsx3 library to convert the text data into speech data and responds aloud through the speaker, for example, "Are you tired? I recommend you take a break."

[1469] Output: The user receives the appropriate response via voice.

[1470] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1471] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1472] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1473] [Third embodiment]

[1474] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1475] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1476] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1477] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1478] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1479] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1480] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1481] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1482] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1483] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1484] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1485] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1486] The system of the present invention is designed to provide users with personalized responses to their questions and concerns in everyday life. Specific embodiments for carrying out the present invention will be described below.

[1487] System configuration

[1488] This system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[1489] 1. Input Method

[1490] 2. Profile Generation Method

[1491] 3. Response Generation Method

[1492] 4. Audio output means

[1493] 5. Learning Methods

[1494] 6. Update method

[1495] 7. Authentication Methods

[1496] 8. Voice Recognition Methods

[1497] 9. Means of Transmission

[1498] Program processing flow

[1499] Initial Setup Phase

[1500] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1501] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[1502] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[1503] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[1504] Daily Use Phase

[1505] 1. User: Speaks to their smartphone. Example: "What do you recommend for dinner tonight?"

[1506] 2. Terminal: The voice input module recognizes the voice and converts it into text data, which is then sent to the server.

[1507] 3. Server: The received text data is analyzed using a natural language processing engine to extract the intent. The server then references the user profile and past data to generate the optimal response.

[1508] 4. Server: The generated response is sent to the terminal in text format.

[1509] 5. Terminal: The text data is converted into speech using a speech synthesis engine, and a response is provided to the user. Example: "I think curry would be good for dinner tonight. The ingredients needed are..."

[1510] Specific examples

[1511] Example 1: Menu suggestion

[1512] 1. User: "What should I make today?"

[1513] 2. Device: Converts speech into text and sends it to the server.

[1514] 3. Server: Analyzes the text data and checks the user's past meal history, preferences, and refrigerator inventory information. Then, it generates an appropriate menu.

[1515] 4. Server: Sends the generated menu suggestions in text format to the terminal.

[1516] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Today, I recommend teriyaki chicken."

[1517] Example 2: Advice

[1518] 1. User: "I've been having trouble at work lately. What should I do?"

[1519] 2. Device: Converts speech into text and sends it to the server.

[1520] 3. Server: Analyzes the text data and generates advice based on the user's work situation and past consultation history.

[1521] 4. Server: Sends the generated advice in text format to the terminal.

[1522] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Try to complete your daily tasks little by little, at your own pace, without rushing."

[1523] Advanced Security Settings

[1524] 1. User: Open the security settings from the device settings screen. Select from the setting options of Corneal Identification, Fingerprint Identification, and DNA Identification.

[1525] 2. Device: According to the selected authentication method, the necessary biometric data is collected. The data is stored in a temporary storage area and encrypted for security purposes.

[1526] 3. Server: Receives encrypted biometric data from the device and stores it in a secure database. Builds an authentication system and applies it to each user.

[1527] 4. Terminal: Installs and activates the authentication system received from the server locally, granting the user access whenever authentication is successful.

[1528] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching daily life. In addition, the system's advanced security functions can safely protect the user's personal information.

[1529] The processing flow will be explained below.

[1530] Initial Setup Phase

[1531] Step 1:

[1532] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1533] Step 2:

[1534] Terminal: Temporarily stores personal information entered by the user. Checks the completeness and format of the input data and determines whether the next step can be processed.

[1535] Step 3:

[1536] Terminal: The personal information entered is sent to the server. When sending, the data is encrypted to ensure security.

[1537] Step 4:

[1538] Server: Deserializes the personal information received from the device and stores it in a database. Based on the stored data, an initial profile is generated.

[1539] Step 5:

[1540] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[1541] Step 6:

[1542] Terminal: Saves the initial profile and learning model received from the server locally. Notifies the user that the setup is complete.

[1543] Daily Use Phase

[1544] Step 1:

[1545] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[1546] Step 2:

[1547] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[1548] Step 3:

[1549] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[1550] Step 4:

[1551] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[1552] Step 5:

[1553] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[1554] Step 6:

[1555] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1556] Advanced Security Settings

[1557] Step 1:

[1558] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[1559] Step 2:

[1560] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[1561] Step 3:

[1562] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[1563] Step 4:

[1564] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[1565] Step 5:

[1566] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[1567] Step 6:

[1568] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[1569] Continuous learning phase

[1570] Step 1:

[1571] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[1572] Step 2:

[1573] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[1574] Step 3:

[1575] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[1576] Step 4:

[1577] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[1578] Step 5:

[1579] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[1580] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[1581] Example 1

[1582] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1583] Current systems that provide personalized responses lack the ability to quickly and accurately process user input and provide appropriate responses to individual questions and concerns in everyday life. Furthermore, continuous learning and profile updates based on user behavioral history are insufficient, resulting in issues with the accuracy and relevance of responses. Furthermore, in terms of security, measures to safely protect personal information are insufficient, and the authentication process is cumbersome, resulting in a lack of user convenience.

[1584] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1585] In this invention, the server includes: input means for inputting a user's personal information, personality, hobbies, preferences, and behavioral patterns; profile generation means for processing the information acquired from the input means; response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; voice output means for returning the responses generated by the response generation means by voice; learning means for recording and learning the user's behavioral history; update means for updating the profile generation means based on the information acquired by the learning means; voice recognition means for analyzing the user's voice and converting it into text data; transmission means for transmitting the text data converted by the voice recognition means to the response generation means; authentication means for performing high-security authentication; and initial learning model generation means for generating an individual learning model using the user profile and storing it on the server. This makes it possible to comprehensively process a user's diverse information to provide highly accurate personalized responses, and to improve the accuracy of responses through continuous learning, thereby realizing a safe and convenient authentication process.

[1586] The "input means" is a means for inputting data such as the user's personal information, personality, hobbies, preferences, and behavioral patterns.

[1587] The "profile generating means" is a means for generating a user profile based on information acquired from the input means.

[1588] The "answer generation means" is a means for generating an appropriate answer to a user's question or concern based on the generated user profile.

[1589] The "audio output means" is a means for returning the response generated by the response generating means to the user by voice.

[1590] The "learning means" is a means for recording the user's behavior history and for the system to continuously learn based on that data.

[1591] The "update means" is a means for updating the profile generation means based on new information acquired by the learning means.

[1592] The "voice recognition means" is a means having the function of analyzing the user's voice and converting it into text data.

[1593] The "transmitting means" is a means for transmitting the text data converted by the speech recognition means to the response generating means.

[1594] "Authentication means" refers to a means for performing an advanced authentication process to ensure user security.

[1595] The "initial learning model generating means" is a means for generating an individual learning model based on a user profile and storing it on a server.

[1596] The "voice synthesis means" is a means for converting the text data generated by the response generation means into voice.

[1597] The system of the present invention is designed to provide users with personalized responses to their everyday questions and concerns. The components and operation of this system will now be described in detail.

[1598] System configuration

[1599] The system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system consists of the following main components:

[1600] 1. Input Method

[1601] A user starts up a new smartphone and inputs personal information such as name, date of birth, personality, blood type, hobbies, and preferences. This information is provided to the system via input means.

[1602] 2. Profile Generation Method

[1603] The terminal stores the user information acquired from the input means in a temporary storage area. After verifying the integrity of the data, it sends it to the server. The server updates the database based on the received data and creates a user profile. Database software (e.g., MySQL or PostgreSQL) is used to generate the profile.

[1604] 3. Response Generation Method

[1605] Once the profile is created, the system generates responses to the user's questions and concerns. This response generation process uses a natural language processing engine (for example, IBM Watson's NLP capabilities) to generate the best possible response based on the user profile and past data.

[1606] 4. Audio output means

[1607] The generated response is sent in text format to the user device, which then uses a speech synthesis engine (e.g., Microsoft Azure's Text-to-Speech service) to convert the text into speech and provide the response to the user.

[1608] 5. Learning Methods

[1609] The system records the user's behavior history and continuously learns from it. This learning process is implemented using a machine learning framework (e.g., TensorFlow).

[1610] 6. Update method

[1611] The information obtained by the learning means is reflected in the profile generation means as appropriate, and the profile is updated, thereby improving the accuracy of the system over time.

[1612] 7. Voice Recognition Methods

[1613] When a user speaks a question, the device's voice input module recognizes the speech and converts it into text data. This process uses a speech recognition API (for example, Google's Speech-to-Text API).

[1614] 8. Means of Transmission

[1615] The text data converted by the speech recognition means is sent to the server via the transmission means.

[1616] 9. Authentication Methods

[1617] The system utilizes biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide a high level of security, ensuring that users' personal information is kept safe.

[1618] 10. Initial learning model generation method

[1619] The server generates an initial learning model based on the user profile and stores it on the server. This initial learning model is generated using a machine learning framework such as TensorFlow.

[1620] Specific examples

[1621] Example 1: Menu suggestion

[1622] When a user speaks to their smartphone, "What should I make today?", the voice input module recognizes the speech and converts it into text data. This text data is sent to the server and analyzed by a natural language processing engine. An appropriate menu is generated by referring to the user's past meal history, preferences, and refrigerator inventory information. The generated menu is sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and suggested to the user.

[1623] Example 2: Advice

[1624] When a user asks, "Work hasn't been going well lately. What should I do?", the voice input module converts the voice into text data. This data is sent to the server, where a natural language processing engine analyzes the user's work situation and past consultation history. The most appropriate advice is generated and sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and conveyed to the user.

[1625] AI chatbot-generated prompts

[1626] Example prompt 1: "Generate a response when a user asks verbally, 'What's for dinner tonight?'"

[1627] Example prompt 2: "Generate advice for when a user says, 'I've been having trouble at work lately. What should I do?'"

[1628] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching the user's daily life and protecting it safely.

[1629] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1630] Step 1:

[1631] The user starts up the new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1632] Input: Personal information that users enter into input forms.

[1633] Output: The personal information data entered.

[1634] Step 2:

[1635] The device receives user input and stores the entered data in a temporary storage area (the smartphone's RAM). At this point, the data is verified for integrity and sent to the server. The data is encrypted.

[1636] Input: Personal information data provided by the user.

[1637] Data processing: Data validation and encryption (e.g., AES encryption).

[1638] Output: The encrypted personal information data is sent to the server.

[1639] Step 3:

[1640] The server receives the data sent from the device and stores it in a database (MySQL, PostgreSQL, etc.). It generates a user profile based on the received initial data and creates an initial learning model.

[1641] Input: Encrypted personal information data.

[1642] Data processing: Decrypting data and saving it to the database, generating user profiles (executing SQL queries).

[1643] Output: An initial learning model and a user profile are generated and stored in a database.

[1644] Step 4:

[1645] The device receives the initial learning model from the server and saves it locally (in the smartphone's internal storage). A message is displayed to the user indicating that the setup is complete.

[1646] Input: The initial training model sent from the server.

[1647] Data processing: Saving the initial training model.

[1648] Output: The initial training model is saved locally and a message appears saying that the setup is complete.

[1649] Step 5:

[1650] The user speaks to the smartphone using voice. For example, "What do you recommend for dinner tonight?"

[1651] Input: User's voice input.

[1652] Step 6:

[1653] The device's voice input module recognizes speech and converts it into text data. This process uses a speech recognition API (such as Google's Speech-to-Text API). The converted text data is then sent to the server.

[1654] Input: Audio data.

[1655] Data processing: Recognizing voice data and converting it to text.

[1656] Output: Text data is sent to the server.

[1657] Step 7:

[1658] The server analyzes the received text data using a natural language processing engine (such as IBM Watson's NLP function) to extract the intent, and generates the optimal response by referencing the user profile and past data.

[1659] Input: Text data.

[1660] Data processing: Natural language processing, intent extraction, and response generation.

[1661] Output: The generated response text data.

[1662] Step 8:

[1663] The server sends the generated response in text format to the terminal.

[1664] Input: The generated response text data.

[1665] Output: The response text data is sent to the terminal.

[1666] Step 9:

[1667] The device uses a speech synthesis engine (such as Microsoft Azure's Text-to-Speech service) to convert the text data into speech and provide a response to the user. For example: "I think curry would be good for dinner tonight. The ingredients you need are..."

[1668] Input: Response text data.

[1669] Data processing: Converting text data into audio.

[1670] Output: A spoken response.

[1671] Step 10:

[1672] The learning method records the user's behavior history and the system continuously learns based on that data. This is implemented using a machine learning framework (such as TensorFlow).

[1673] Input: User behavior history data.

[1674] Data processing: Preprocessing data and training machine learning models.

[1675] Output: The trained model and updated model parameters.

[1676] Step 11:

[1677] The server updates the profile generation means based on the information obtained by the learning means, thereby improving the profile and increasing the accuracy of responses.

[1678] Input: trained model parameters.

[1679] Data processing: Profile update work.

[1680] Output: The updated user profile.

[1681] Step 12:

[1682] The authentication method uses biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide high-security authentication.

[1683] Input: Biometric data.

[1684] Data Processing: Biometric data collection and authentication.

[1685] Output: Authentication result.

[1686] Step 13:

[1687] The initial learning model generating means generates an individual learning model based on the user profile and stores it in the server.

[1688] Input: User profile data.

[1689] Data processing: generating and saving learning models.

[1690] Output: Initial training model.

[1691] (Application example 1)

[1692] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1693] Customer service in traditional brick-and-mortar stores relies on the knowledge and experience of store staff, making it difficult to provide consistent, high-quality service to customers. In particular, responses to customer questions and concerns are not personalized, which leads to lower customer satisfaction. It is also difficult to accurately grasp daily customer behavior patterns and preferences, making it difficult to recommend optimal products. Furthermore, there are also issues with in-store security authentication, and personal information may not be adequately protected.

[1694] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1695] In this invention, the server includes: an input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns; a profile generation means for processing the information acquired from the input means; a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; a voice output means for returning the responses generated by the response generation means by voice; a learning means for recording and learning the user's behavioral history; an update means for updating the profile generation means based on the information acquired by the learning means; an authentication means for performing high-security authentication; a voice recognition means for analyzing the voice input and converting it into text data; a natural language processing means for analyzing the text data using a natural language processing engine and generating an optimal response by referring to the user's profile and past data; a voice synthesis means for converting the text data generated by the natural language processing means into voice; a robot that operates as a customer service assistant to support operations in a physical store; and a profile generation means for suggesting products in the physical store. This enables personalized responses to customer questions and concerns, thereby contributing not only to improved customer satisfaction but also to improved security.

[1696] "Input means" refers to a device or function for inputting the user's personality, hobbies, preferences, and daily behavior patterns.

[1697] The "profile generating means" refers to a device or function for processing information acquired from the input means and generating a user profile.

[1698] The "response generating means" refers to a device or function for responding to a user's questions or concerns based on the profile generated by the profile generating means.

[1699] The "audio output means" refers to a device or function for returning the response generated by the response generation means by voice.

[1700] The "learning means" refers to a device or function for recording and learning from the user's behavior history.

[1701] The "update means" refers to a device or function for updating the profile generation means based on the information acquired by the learning means.

[1702] "Authentication means" means a device or function for performing high-security authentication.

[1703] "Speech recognition means" refers to a device or function for analyzing voice input and converting it into text data.

[1704] "Natural language processing means" refers to a device or function that uses a natural language processing engine to analyze text data and generate optimal responses by referring to the user's profile and past data.

[1705] The "speech synthesis means" refers to a device or function for converting text data generated by the natural language processing means into speech.

[1706] A "customer service assistant" is a robot that assists with work in physical stores.

[1707] The "profile generation means for making product suggestions" refers to a device or function that generates a profile for making product suggestions in a physical store.

[1708] The system of the present invention aims to provide personalized responses to questions and concerns that users have about their daily lives in brick-and-mortar stores. Specific embodiments for carrying out the invention will now be described.

[1709] System configuration

[1710] The system consists of the following major components:

[1711] 1. Input Method

[1712] 2. Profile Generation Method

[1713] 3. Response Generation Method

[1714] 4. Audio output means

[1715] 5. Learning Methods

[1716] 6. Update method

[1717] 7. Authentication Methods

[1718] 8. Voice Recognition Methods

[1719] 9. Natural Language Processing Tools

[1720] 10. Speech synthesis means

[1721] 11. Customer Service Assistant

[1722] 12. Product proposal profile generation means

[1723] Details of each means are as follows.

[1724] Program Generation and Hardware / Software Usage

[1725] 1. Initial Setup Phase

[1726] User: Launches the application at the store entrance and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[1727] Terminal: The entered information is temporarily stored, encrypted, and sent to the server.

[1728] Server: Generates a user profile based on the received data and creates a personalized product recommendation model.

[1729] Terminal: Saves the initial model received from the server locally and displays a notification that the setup is complete.

[1730] 2. Operational Phase

[1731] User: Asks questions by voice in-store. Example: "What wines do you recommend these days?"

[1732] Terminal: The speech is converted into text using a speech recognition tool (Google Speech-to-Text API) and sent to the server.

[1733] Server: Analyzes the text data using a natural language processing engine (spaCy or BERT) and generates an appropriate response by referencing the user's profile and past data.

[1734] Server: Sends the generated response in text format to the terminal.

[1735] Terminal: Converts text to speech using a speech synthesis engine (Amazon Polly) and provides responses to the user.

[1736] Specific examples

[1737] Example 1

[1738] User: "What wines have you recommended recently?"

[1739] Device: Converts speech to text and sends it to the server.

[1740] Server: Generates a response based on the user's profile and past data, and sends text data to the device saying, "This time, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[1741] Terminal: A speech synthesis engine is used to convert text into speech and provide guidance to the user.

[1742] Prompt Sentence Examples

[1743] Please enter your name, gender, age, preferences, purchase history, and product interest information.

[1744] What wines do you recommend these days?

[1745] "Today, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[1746] Hardware and software used

[1747] Hardware: Smartphones, smart glasses, in-store robots

[1748] software:

[1749] Speech recognition engine: Google Speech-to-Text API

[1750] Natural language processing engine: spaCy, BERT-based model

[1751] Speech synthesis engine: Amazon Polly

[1752] Database: MySQL, Firebase

[1753] Communication protocol: HTTPS, WebSocket

[1754] In this way, the system of the present invention can provide personalized responses to users' questions and concerns, thereby improving customer satisfaction in brick-and-mortar stores. Furthermore, advanced security features can protect users' personal information.

[1755] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1756] Step 1:

[1757] User: Launches the application at the entrance of a physical store and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[1758] Input: Personal information

[1759] Output: Temporarily saved personal information

[1760] Specific operation: By entering information into the input form and pressing the "Submit" button, the data is stored in a temporary storage area on the device.

[1761] Step 2:

[1762] Terminal: The personal information entered is temporarily stored, encrypted for security purposes, and then sent to the server.

[1763] Input:Temporarily saved personal information

[1764] Output: Encrypted personal information data

[1765] Specific operation: Personal information is encrypted inside the device and sent to the server using the HTTPS communication protocol.

[1766] Step 3:

[1767] Server: Decrypts the received encrypted data and stores it in a database. Generates a user profile and creates a personalized product recommendation model.

[1768] Input: Encrypted personal information data

[1769] Output: User profile and proposed model

[1770] Specific operations: Decrypting encrypted data, running the user profile generation algorithm, and saving the generated profile and proposed model to a database.

[1771] Step 4:

[1772] Terminal: The terminal locally stores the user profile and proposed model received from the server and displays a notification to the user that the setup is complete.

[1773] Input: User profile and proposed model data

[1774] Output: Locally saved profile and proposed model, notification that setup is complete

[1775] Specific behavior: Save received data locally, display completion notification

[1776] Step 5:

[1777] User: Asks questions by dictating as they walk through the store. Example: "What wines do you recommend these days?"

[1778] Input: Voice data (question)

[1779] Output: Start speech recognition via speech input interface

[1780] Specific actions: Press the voice input button on your smartphone or smart glasses and dictate your question.

[1781] Step 6:

[1782] Terminal: Using a voice recognition tool (Google Speech-to-Text API), the voice data is converted into text data and sent to the server.

[1783] Input: Audio data

[1784] Output: Text data

[1785] Specific operations: Calling the speech recognition API, generating converted text data, and sending the text data to the server.

[1786] Step 7:

[1787] Server: Analyzes the text data using natural language processing tools (spaCy or BERT engine) and generates an appropriate response by referencing the user's profile and past data.

[1788] Input: Text data, user profile

[1789] Output: Response text data

[1790] What it does: Runs a natural language analysis engine, consults a user profile database, and generates the best possible response.

[1791] Step 8:

[1792] Server: Sends the generated response in text format to the terminal.

[1793] Input: Response text data

[1794] Output: Text data sent to the terminal

[1795] Specific operation: Processing to send text data to the terminal.

[1796] Step 9:

[1797] Terminal: The response text data is converted into speech using a speech synthesis engine (Amazon Polly) and a response is provided to the user.

[1798] Input: Response text data

[1799] Output: Audio data

[1800] Specific operations: Calling the speech synthesis API, generating converted speech data, and outputting speech from the speaker.

[1801] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1802] The system of the present invention is designed in combination with an emotion engine to provide personalized responses to users' questions and concerns in their daily lives. Specific embodiments for implementing the present invention will be described below.

[1803] System configuration

[1804] This system generates an individual profile based on the user's input and emotional information, and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[1805] 1. Input Method

[1806] 2. Profile Generation Method

[1807] 3. Response Generation Method

[1808] 4. Audio output means

[1809] 5. Learning Methods

[1810] 6. Update method

[1811] 7. Authentication Methods

[1812] 8. Voice Recognition Methods

[1813] 9. Means of Transmission

[1814] 10. Emotion Recognition Engine

[1815] Program processing flow

[1816] Initial Setup Phase

[1817] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1818] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[1819] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[1820] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[1821] Daily Use Phase

[1822] 1. User: Talks to the smartphone and asks questions or asks for advice. For example, "What do you recommend for dinner tonight?"

[1823] 2. Terminal: The voice input module captures the spoken voice and converts it into text data using the voice recognition engine.

[1824] 3. Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[1825] 4. Server: The received text data is analyzed using a natural language processing engine to extract the user's intent. The server then generates the optimal response data by referencing the user profile and past data.

[1826] 5. Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[1827] 6. Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it may reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1828] Processing using an emotion recognition engine

[1829] 1. User: Speaks, makes facial expressions and shows movements.

[1830] 2. On the device: The emotion recognition engine detects the user's tone of voice, facial expressions, movements, and language patterns to analyze emotions in real time.

[1831] 3. Terminal: Temporarily stores the analyzed emotion data and sends it to the server.

[1832] 4. Server: Receives emotion data and uses it in the profile generation means and response generation means.

[1833] 5. Server: Based on the emotional data, generate a personalized response that is more suited to the user's current situation.

[1834] 6. Device: Provides the user with a voice response based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1835] Advanced Security Settings

[1836] 1. User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[1837] 2. Terminal: According to the selected authentication method, the necessary biometric authentication data is acquired, temporarily stored, and encrypted.

[1838] 3. Terminal: Sends encrypted authentication data to the server, which checks the data for consistency and integrity upon transmission.

[1839] 4. Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[1840] 5. Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[1841] 6. Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[1842] Continuous learning phase

[1843] 1. User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[1844] 2. Device: Automatically collects user behavior history and usage data, temporarily stores the collected data, and sends it to the server as needed.

[1845] 3. Server: Analyzes the received behavioral history data and updates the user profile. Machine learning algorithms are used to learn the user's behavioral patterns.

[1846] 4. Server: Periodically synchronizes updated profiles and learning results to the device. During synchronization, it checks the consistency and integrity of the data.

[1847] 5. Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[1848] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively. The present invention can provide personalized services based on the user's individual information and emotional information, enriching daily life, and safely protecting the user's personal information with advanced security functions.

[1849] The processing flow will be explained below.

[1850] Initial Setup Phase

[1851] Step 1:

[1852] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1853] Step 2:

[1854] Terminal: Temporarily stores personal information entered by the user, checks the integrity and format of the entered data, and sends it to the server.

[1855] Step 3:

[1856] Server: Receives data sent from the device and stores it in a database. Based on the stored data, it generates an initial profile.

[1857] Step 4:

[1858] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[1859] Step 5:

[1860] Terminal: The initial profile and learning model received from the server are saved locally and the user is notified that the setup is complete.

[1861] Daily Use Phase

[1862] Question and answer processing

[1863] Step 1:

[1864] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[1865] Step 2:

[1866] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[1867] Step 3:

[1868] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[1869] Step 4:

[1870] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[1871] Step 5:

[1872] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[1873] Step 6:

[1874] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1875] Emotion Recognition Processing

[1876] Step 1:

[1877] User: Speaks, shows facial expressions and actions.

[1878] Step 2:

[1879] On the device: The emotion recognition engine detects the user's tone of voice, facial expressions, movements, and language patterns to analyze emotions in real time.

[1880] Step 3:

[1881] Terminal: Temporarily stores the analyzed emotion data and sends it to the server.

[1882] Step 4:

[1883] Server: Receives emotion data and uses it in the profile generation means and response generation means.

[1884] Step 5:

[1885] Server: Based on the emotional data, it generates a personalized response that is more suited to the user's current situation.

[1886] Step 6:

[1887] The device provides the user with a voice response based on the generated emotion. For example, if the device detects that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1888] Advanced Security Settings

[1889] Step 1:

[1890] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[1891] Step 2:

[1892] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[1893] Step 3:

[1894] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[1895] Step 4:

[1896] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[1897] Step 5:

[1898] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[1899] Step 6:

[1900] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[1901] Continuous learning phase

[1902] Step 1:

[1903] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[1904] Step 2:

[1905] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[1906] Step 3:

[1907] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[1908] Step 4:

[1909] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[1910] Step 5:

[1911] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[1912] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[1913] Example 2

[1914] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1915] In modern society, people often have many questions and concerns in their daily lives, requiring efficient and personalized responses. Furthermore, it is necessary to provide services that provide higher levels of satisfaction by providing responses that are tailored to the user's emotions and circumstances. However, current systems perform tasks such as user profile creation and updating, voice recognition, emotion recognition, and advanced security authentication separately, resulting in a lack of an integrated and efficient solution.

[1916] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1917] In this invention, the server includes input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns, profile generation means for processing information acquired from the input means, response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means, voice output means for returning the responses generated by the response generation means by voice, learning means for recording and learning the user's behavioral history, update means for updating the profile generation means based on information acquired by the learning means, voice recognition means for analyzing the user's voice and converting it into text data, transmission means for transmitting the text data converted by the voice recognition means to the response generation means, an emotion recognition engine for analyzing the user's emotions, means for personalizing responses based on the emotion data acquired by the emotion recognition engine, and authentication means for performing high-security authentication. This enables personalized responses that reflect the user's situation and emotions.

[1918] 1. "Input means" refers to a device or function that allows a user to input personal information such as personality, hobbies, preferences, and daily behavior patterns into the system.

[1919] 2. "Profile generation means" refers to a device or function that generates a user profile based on information acquired from an input means.

[1920] 3. "Response generation means" refers to a device or function that generates the most appropriate response to a user's question or concern based on the generated user profile.

[1921] 4. "Audio output means" refers to a device or function that provides the response generated by the response generation means to the user by voice.

[1922] 5. "Learning means" refers to a device or function that records the user's behavior history and uses it to improve the system's performance.

[1923] 6. "Update means" refers to a device or function that automatically updates the profile generation means based on information acquired by the learning means.

[1924] 7. "Speech recognition means" means a device or function that captures a user's voice and converts it into text data.

[1925] 8. "Transmitting means" means a device or function that transmits text data converted by the speech recognition means to the response generating means.

[1926] 9. "Emotion recognition engine" means a device or function that analyzes a user's emotions and acquires that data.

[1927] 10. "Personalization means" refers to a device or function that changes responses to suit the user's current situation based on emotional data obtained by an emotion recognition engine.

[1928] 11. "Authentication Means" means a device or function that verifies a user's identity and performs high-security authentication.

[1929] 12. A “generative AI model” is an artificial intelligence-based algorithm that generates appropriate responses to user intent or questions.

[1930] 13. “Prompt sentence” refers to the text data input into a generative AI model, which is used to specifically indicate the user’s question or intention.

[1931] The present invention is a system designed to provide users with personalized responses to questions and concerns they have in their daily lives. This system is composed of multiple components that generate an individual profile based on the user's input information and emotional information, and return appropriate responses. Specific embodiments for implementing the present invention will be described below.

[1932] Hardware and software used

[1933] 1. Terminal: A user device such as a smartphone. This terminal includes a voice input module, a voice recognition engine (e.g., Google Speech-to-Text API), and a voice synthesis engine (e.g., Google Text-to-Speech API).

[1934] 2. Server: A server running a database system (e.g., MySQL or PostgreSQL), a natural language processing engine (e.g., GPT-3), or an emotion recognition engine (e.g., Microsoft Azure Emotion API).

[1935] 3. Generative AI model: An artificial intelligence-based algorithm for analyzing user intent and generating appropriate responses.

[1936] Detailed Description of the Invention

[1937] Initial Setup

[1938] 1. The user starts up a new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[1939] 2. The terminal receives the data entered by the user, temporarily stores it in a storage area, verifies the integrity of the data, and then sends it to the server.

[1940] 3. The server stores the received data in a database and generates a user profile, which is then used to create an initial learning model.

[1941] 4. The device receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[1942] daily use

[1943] 1. The user speaks to their smartphone to ask a question or ask for advice, such as, "What do you recommend for dinner tonight?"

[1944] 2. The device captures the voice using the voice input module and converts it into text data using the voice recognition engine.

[1945] 3. The terminal temporarily stores the converted text data, checks the data for consistency, and then sends it to the server.

[1946] 4. The server analyzes the received text data using a natural language processing engine to extract the user's intent. It then generates the optimal response data by referencing the user profile and past data.

[1947] 5. The server sends the generated response data to the terminal.

[1948] 6. The device converts the response data into an audio file using a speech synthesis engine and provides the user with a spoken response. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[1949] emotion recognition

[1950] 1. The user speaks, uses facial expressions, and makes movements.

[1951] 2. The device uses an emotion recognition engine to analyze the user's emotions in real time.

[1952] 3. The device temporarily stores the analyzed emotion data and sends it to the server.

[1953] 4. The server receives the emotion data and uses it in the profile generation means and response generation means.

[1954] 5. The server generates a personalized response based on the emotion data that better suits the user's current situation.

[1955] 6. The device then provides a voice response to the user based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[1956] Security Settings

[1957] 1. The user selects security settings from the device settings menu and chooses an authentication option such as corneal identification, fingerprint authentication, or DNA authentication.

[1958] 2. The device captures and encrypts the biometric data.

[1959] 3. The device sends the encrypted data to the server, which verifies the data's integrity and completeness.

[1960] 4. The server deserializes the encrypted data and stores it in a secure database. It builds an authentication algorithm and associates it with the user profile.

[1961] 5. The server sends the authentication system to the terminal, which installs and configures the system within the terminal.

[1962] 6. The terminal has a security authentication system installed, and grants the user access whenever authentication is successful.

[1963] Continuous learning

[1964] 1. Users use smartphones on a daily basis and use various apps and functions.

[1965] 2. The device automatically collects the user's behavioral history and usage data, temporarily stores it, and sends it to the server.

[1966] 3. The server analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[1967] 4. The server periodically synchronizes the updated profile and learning results to the device, verifying the consistency and completeness of the data.

[1968] 5. The device locally reflects the received update data, providing the user with more personalized services.

[1969] Examples of specific examples and prompts

[1970] Specific examples

[1971] A user asks, "What's the best way to relax after work?"

[1972] The device uses a voice recognition engine to convert the question into text and send it to the server.

[1973] The server analyzes the text, understands the user's intent, and generates optimal advice on how to relax.

[1974] The server sends the generated advice to the terminal, which converts it into an audio file and responds to the user by saying, "Why don't you take a nice, long bath?"

[1975] Prompt Sentence Examples

[1976] "Generate the best response when a user asks, 'What's the best way to relax after work?'"

[1977] "What relaxation methods can we suggest to users when they are feeling stressed?"

[1978] The above is an embodiment of the present invention. This system provides personalized responses based on the user's individual information and emotional information, enriching daily life and ensuring safety.

[1979] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1980] Initial Setup Phase

[1981] Step 1:

[1982] A user starts up their smartphone and enters personal information such as their name, date of birth, personality, blood type, hobbies, etc. The data entered is the user's basic information, such as "Taro Tanaka, May 4, 1990, blood type O, watching movies."

[1983] Input: Personal information (name, date of birth, personality, blood type, hobbies and preferences)

[1984] Output: Personal information data entered

[1985] Step 2:

[1986] The terminal receives the data entered by the user and temporarily stores it in a storage area. At this time, the terminal checks the integrity of the data and monitors whether there are any missing data.

[1987] Input: Personal data entered by the user

[1988] Output: Temporarily saved data

[1989] Step 3:

[1990] The data that has been confirmed by the device is sent to the server using the HTTP protocol, and the data is encrypted.

[1991] Input: Temporarily saved personal information data

[1992] Output: Data sent to the server

[1993] Step 4:

[1994] The server receives the data and stores it in a database (e.g., MySQL or PostgreSQL). It generates a user profile based on the initial data and creates an initial learning model (e.g., TensorFlow or PyTorch).

[1995] Input: Received personal information data

[1996] Output: Generated user profile and initial learning model

[1997] Step 5:

[1998] The server sends the generated initial learning model to the terminal via the endpoint.

[1999] Input: Initial training model

[2000] Output: The trained model sent to the device.

[2001] Step 6:

[2002] The device saves the initial learning model locally and displays a message to the user indicating that setup is complete.

[2003] Input: The received training model

[2004] Output: Locally saved training model and a message saying the setup is complete

[2005] Daily Use Phase

[2006] Step 1:

[2007] Users can ask questions or ask their smartphones by voice, for example, "What do you recommend for dinner tonight?"

[2008] Input: Voice question

[2009] Output: Audio data

[2010] Step 2:

[2011] The device captures audio using the built-in microphone and converts it into text data using a speech recognition engine (e.g., Google Speech-to-Text API).

[2012] Input: Audio data

[2013] Output: Text data

[2014] Step 3:

[2015] The terminal stores the converted text data in a temporary storage area, checks the data for consistency, and then transmits it to the server.

[2016] Input: Text data

[2017] Output: Text data to send to the server

[2018] Step 4:

[2019] The server receives the text data, analyzes it with a natural language processing engine (e.g., GPT-3), extracts the user's intent, and generates the optimal response by referring to the user's profile and past data.

[2020] Input: Received text data

[2021] Output: The generated response data

[2022] Step 5:

[2023] The server transmits the generated response data in text format to the terminal.

[2024] Input: Generated response data

[2025] Output: Response data sent to the terminal

[2026] Step 6:

[2027] The device converts the received response data into an audio file using a speech synthesis engine (for example, Google Text-to-Speech API) and provides a voice response to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[2028] Input: Received response data

[2029] Output: A spoken response provided to the user

[2030] Processing using an emotion recognition engine

[2031] Step 1:

[2032] The user speaks, makes facial expressions, and makes movements.

[2033] Input: voice, facial expressions, movements

[2034] Output: Expressed emotion data

[2035] Step 2:

[2036] The device analyzes the user's emotions in real time using an emotion recognition engine (for example, Microsoft Azure Emotion API).

[2037] Input: Expressed emotion data

[2038] Output: Parsed emotion data

[2039] Step 3:

[2040] The device temporarily stores the analyzed emotion data and transmits it to the server.

[2041] Input: Parsed emotion data

[2042] Output: Emotion data sent to the server

[2043] Step 4:

[2044] The server receives the emotion data and uses it in the profile generation means and response generation means.

[2045] Input: Received emotion data

[2046] Output: Emotion data used in the profile generation and response generation methods

[2047] Step 5:

[2048] The server generates a personalized response based on the emotional data that better suits the user's current situation.

[2049] Input: Emotion data

[2050] Output: Personalized response data

[2051] Step 6:

[2052] The device will provide a voice response to the user based on the generated emotion. For example, if the device recognizes that the user is feeling stressed, it will suggest, "Shall I play some relaxing music?"

[2053] Input: Emotion-based response data

[2054] Output: Spoken suggestions

[2055] Advanced Security Settings

[2056] Step 1:

[2057] Users select security settings from the device's settings menu and choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[2058] Input: Select user authentication settings

[2059] Output: The configured authentication options

[2060] Step 2:

[2061] The terminal acquires and encrypts the biometric authentication data according to the selected authentication method.

[2062] Input: Authentication method and biometric data

[2063] Output: Encrypted biometric data

[2064] Step 3:

[2065] The device sends the encrypted data to the server, which verifies the data's integrity and completeness.

[2066] Input: Encrypted biometric data

[2067] Output: Encrypted data sent to the server

[2068] Step 4:

[2069] The server deserializes the encrypted data and stores it in a secure database, building an authentication algorithm and associating it with the user profile.

[2070] Input: Encrypted biometric data

[2071] Output: User profile associated with constructed authentication algorithm

[2072] Step 5:

[2073] The server sends the authentication system to the terminal, which then installs and configures the system within the terminal.

[2074] Input:AuthenticationSystem

[2075] output: Authentication system configured on the device

[2076] Step 6:

[2077] The terminal installs an authentication system and grants the user access whenever authentication is successful.

[2078] Input:AuthenticationSystem

[2079] Output: Access permission granted upon successful authentication

[2080] Continuous learning phase

[2081] Step 1:

[2082] Users use smartphones on a daily basis and utilize various applications and functions.

[2083] Input: Everyday smartphone use

[2084] Output: Application and feature usage data

[2085] Step 2:

[2086] The device automatically collects the user's behavioral history and usage data, temporarily stores it, and sends it to the server.

[2087] Input: Application and feature usage data

[2088] Output: Behavioral history data sent to the server

[2089] Step 3:

[2090] The server analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[2091] Input: Received behavioral history data

[2092] Output: Updated user profile

[2093] Step 4:

[2094] The server periodically synchronizes updated profiles and learning results with the device and checks the consistency and completeness of the data.

[2095] Input: Updated user profile

[2096] Output: Data synced to the device

[2097] Step 5:

[2098] The terminal locally reflects the updated data received from the server, thereby providing the user with a more personalized service.

[2099] Input: Update data synced to the device

[2100] Output: Personalized service

[2101] The above is a description of the specific operations in the processing steps of the present invention. Through this processing, the system provides personalized responses based on the user's individual information and emotional information, realizing a system that can enrich daily life and ensure safety.

[2102] (Application example 2)

[2103] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2104] Current worker support systems in factories provide standard responses without considering the worker's emotions, resulting in insufficient stress management and optimization of work efficiency. Furthermore, when workers ask questions or ask for advice via voice, appropriate responses that take their emotions into consideration are not provided, which can lead to lower worker satisfaction and impact productivity. There is a growing need for a system that can solve these issues and provide responses that take the worker's emotional state into consideration.

[2105] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an input means for inputting the user's personality, hobbies, preferences, and daily behavioral patterns, a profile generation means for processing information acquired from the input means, and a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means. This makes it possible to provide accurate and personalized responses while taking into consideration the user's emotional state.

[2106] The "input means" is a device or interface for inputting information such as the user's personality, hobbies, preferences, and daily behavior patterns.

[2107] The "profile generating means" is a device or program for processing information acquired from the input means and generating an individual profile for the user.

[2108] The "answer generation means" is a device or program for generating appropriate answers to the user's questions and concerns based on the generated user profile.

[2109] The "audio output means" refers to a device or technology for returning the response generated by the response generation means by voice.

[2110] A "learning means" is a device or algorithm that records a user's behavioral history and uses it to train a profile or system.

[2111] The "update means" is a device or technology for updating the profile generation means based on the information acquired by the learning means.

[2112] An "authentication means" is a device or system for performing high-security authentication.

[2113] "Voice input means" refers to a device or technology for capturing spoken voice and converting it into text data using a voice recognition engine.

[2114] "Emotion recognition means" refers to a device or technology for analyzing the user's emotions by analyzing the tone of voice and facial expressions.

[2115] System configuration

[2116] This invention is a "smart industrial assistant" system that supports workers in factories. The system consists of the following main components:

[2117] 1. Input means: An interface for inputting information such as the worker's personality, hobbies and preferences, and daily behavioral patterns, and is typically a smartphone or tablet.

[2118] 2. Profile generation means: A program for processing information obtained from the input means and generating a profile for each worker.

[2119] 3. Response generation means: A program that generates appropriate responses to the worker's questions and concerns based on the generated profile.

[2120] 4. Audio output means: This is a technology for outputting the response generated by the response generation means by voice, and a speaker is used.

[2121] 5. Learning method: An algorithm that records the worker's behavioral history and uses it to train the profile and system.

[2122] 6. Update method: This is a technology that updates the profile generation method based on the information obtained by the learning method.

[2123] 7. Authentication method: A device or system for performing high-security authentication, and options include corneal identification, fingerprint authentication, and DNA authentication.

[2124] 8. Voice input method: This technology captures the voice spoken by the worker and converts it into text data using a voice recognition engine.

[2125] 9. Emotion recognition: This technology analyzes the emotions of workers by analyzing their tone of voice and facial expressions.

[2126] Program processing explanation

[2127] The server uses the Python language, the speech_recognition library for speech recognition, and the pyttsx3 library for speech synthesis. This allows the server to recognize the worker's voice, convert it into text data, and generate an appropriate response based on the worker's profile and emotional information. The generated response is output as audio through the speaker.

[2128] Hardware and software used

[2129] Hardware:

[2130] Microphone: Used to capture audio input.

[2131] Speaker: Used to output responses aloud.

[2132] Robot terminal: Used as an installation platform.

[2133] software:

[2134] Python: Used to implement the program.

[2135] speech_recognition: A library for speech recognition.

[2136] pyttsx3: A text-to-speech engine.

[2137] requests: Used to communicate with the server (calling a dummy server).

[2138] Specific examples

[2139] As an example, if a worker asks, "What are today's inspection procedures?", the voice input means captures the voice and the voice recognition engine converts it into text data. If the emotion recognition means analyzes the worker's tone of voice and facial expression and recognizes that the worker is feeling stressed, the response generation means generates an appropriate response such as "Are you tired? I recommend you take a break," and outputs the response aloud from a speaker via the voice output means.

[2140] Prompt Sentence Examples

[2141] The prompt text is assumed to be as follows:

[2142] "It recognizes the user's emotions and generates responses suggesting a break if they are tired."

[2143] In this way, integrating emotion recognition with personalized responses can improve worker satisfaction and productivity.

[2144] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2145] Step 1:

[2146] Users use their smartphones or tablets to enter information about their personality, hobbies, preferences, and daily behavior patterns.

[2147] Input: User's individual information (e.g., personality, hobbies, preferences, behavioral patterns)

[2148] Specific Actions: The user enters information into the interface and clicks the submit button.

[2149] Output: The input data is temporarily stored.

[2150] Step 2:

[2151] The terminal acquires the input information and transmits it to the profile generating means.

[2152] Input: Saved user personal information

[2153] Specific operation: The device establishes a network connection to send information to the server. When the information transmission is complete, a success message is displayed.

[2154] Output: The server receives the information.

[2155] Step 3:

[2156] The server generates a user profile based on the received information.

[2157] Input: User's personal information sent to the server

[2158] Specific operation: The server executes the profile generation algorithm and stores the user profile in the database. Once the profile generation is complete, it sends a confirmation message to the terminal.

[2159] Output: Generated user profile

[2160] Step 4:

[2161] While working, users can talk to their smartphone or tablet to ask questions or ask for advice, for example, "Tell me about today's inspection procedure."

[2162] Input: User voice input

[2163] Specific operation: When the user speaks, the device's microphone captures the sound.

[2164] Output: Captured audio data

[2165] Step 5:

[2166] The device converts the captured voice data into text data using a voice recognition engine.

[2167] Input: Audio data

[2168] Specific operation: The device uses the speech_recognition library to analyze the voice data and convert it to text.

[2169] Output: Converted text data

[2170] Step 6:

[2171] The terminal transmits the converted text data to the server.

[2172] Input: Text data

[2173] Specific operation: The terminal establishes communication with the server and sends text data over the network.

[2174] Output: The server receives the text data.

[2175] Step 7:

[2176] The server uses the emotion recognition means to analyze the user's emotions.

[2177] Input: Text data and associated audio tone data

[2178] Specific operation: The server uses an emotion recognition engine to analyze the tone of voice and language patterns to determine the user's emotional state.

[2179] Output: Analyzed emotion data (e.g., stress, relaxation, etc.)

[2180] Step 8:

[2181] The server generates an appropriate response based on the profile and emotion data using a response generation means.

[2182] Input: User profile, emotion data, text data

[2183] Specific operation: The server uses the generative AI model to generate a response that matches the user's current state. For example, if the user says "I'm tired," it suggests taking a break.

[2184] Output: The generated response data

[2185] Step 9:

[2186] The server transmits the generated response data to the terminal.

[2187] Input: Response data

[2188] Specific operation: The server establishes a network connection to send data to the terminal and sends response data.

[2189] Output: The terminal receives the response data.

[2190] Step 10:

[2191] The response data received by the terminal is converted into voice using a voice synthesis engine, and a voice response is given to the user.

[2192] Input: Response data

[2193] Specific behavior: The device uses the pyttsx3 library to convert the text data into speech data and responds aloud through the speaker, for example, "Are you tired? I recommend you take a break."

[2194] Output: The user receives the appropriate response via voice.

[2195] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[2196] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2197] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[2198] [Fourth embodiment]

[2199] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[2200] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2201] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2202] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[2203] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[2204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[2205] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[2206] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[2207] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[2208] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2209] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2210] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[2211] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2212] The system of the present invention is designed to provide users with personalized responses to their questions and concerns in everyday life. Specific embodiments for carrying out the present invention will be described below.

[2213] System configuration

[2214] This system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[2215] 1. Input Method

[2216] 2. Profile Generation Method

[2217] 3. Response Generation Method

[2218] 4. Audio output means

[2219] 5. Learning Methods

[2220] 6. Update method

[2221] 7. Authentication Methods

[2222] 8. Voice Recognition Methods

[2223] 9. Means of Transmission

[2224] Program processing flow

[2225] Initial Setup Phase

[2226] 1. User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[2227] 2. Terminal: Receives user input and stores the input data in a temporary storage area. At this point, the data is verified for integrity and sent to the server.

[2228] 3. Server: Receives data sent from the device and stores it in a database. Generates a user profile based on the initial data and creates an initial learning model.

[2229] 4. Terminal: Receives the initial training model from the server, saves it locally, and displays a message to the user that the setup is complete.

[2230] Daily Use Phase

[2231] 1. User: Speaks to their smartphone. Example: "What do you recommend for dinner tonight?"

[2232] 2. Terminal: The voice input module recognizes the voice and converts it into text data, which is then sent to the server.

[2233] 3. Server: The received text data is analyzed using a natural language processing engine to extract the intent. The server then references the user profile and past data to generate the optimal response.

[2234] 4. Server: The generated response is sent to the terminal in text format.

[2235] 5. Terminal: The text data is converted into speech using a speech synthesis engine, and a response is provided to the user. Example: "I think curry would be good for dinner tonight. The ingredients needed are..."

[2236] Specific examples

[2237] Example 1: Menu suggestion

[2238] 1. User: "What should I make today?"

[2239] 2. Device: Converts speech into text and sends it to the server.

[2240] 3. Server: Analyzes the text data and checks the user's past meal history, preferences, and refrigerator inventory information. Then, it generates an appropriate menu.

[2241] 4. Server: Sends the generated menu suggestions in text format to the terminal.

[2242] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Today, I recommend teriyaki chicken."

[2243] Example 2: Advice

[2244] 1. User: "I've been having trouble at work lately. What should I do?"

[2245] 2. Device: Converts speech into text and sends it to the server.

[2246] 3. Server: Analyzes the text data and generates advice based on the user's work situation and past consultation history.

[2247] 4. Server: Sends the generated advice in text format to the terminal.

[2248] 5. Device: The text data is converted into speech using a speech synthesis engine, and the user is told, "Try to complete your daily tasks little by little, at your own pace, without rushing."

[2249] Advanced Security Settings

[2250] 1. User: Open the security settings from the device settings screen. Select from the setting options of Corneal Identification, Fingerprint Identification, and DNA Identification.

[2251] 2. Device: According to the selected authentication method, the necessary biometric data is collected. The data is stored in a temporary storage area and encrypted for security purposes.

[2252] 3. Server: Receives encrypted biometric data from the device and stores it in a secure database. Builds an authentication system and applies it to each user.

[2253] 4. Terminal: Installs and activates the authentication system received from the server locally, granting the user access whenever authentication is successful.

[2254] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching daily life. In addition, the system's advanced security functions can safely protect the user's personal information.

[2255] The processing flow will be explained below.

[2256] Initial Setup Phase

[2257] Step 1:

[2258] User: Starts up a new smartphone and begins the initial setup. Enters personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[2259] Step 2:

[2260] Terminal: Temporarily stores personal information entered by the user. Checks the completeness and format of the input data and determines whether the next step can be processed.

[2261] Step 3:

[2262] Terminal: The personal information entered is sent to the server. When sending, the data is encrypted to ensure security.

[2263] Step 4:

[2264] Server: Deserializes the personal information received from the device and stores it in a database. Based on the stored data, an initial profile is generated.

[2265] Step 5:

[2266] Server: Sends the generated initial profile along with the learning model to the terminal. Data integrity is checked when sending.

[2267] Step 6:

[2268] Terminal: Saves the initial profile and learning model received from the server locally. Notifies the user that the setup is complete.

[2269] Daily Use Phase

[2270] Step 1:

[2271] User: Talks to the smartphone to ask questions or ask for advice. For example, "What do you recommend for dinner tonight?"

[2272] Step 2:

[2273] Terminal: The voice input module captures the spoken voice and converts it into text data using a voice recognition engine.

[2274] Step 3:

[2275] Terminal: The converted text data is temporarily stored, and after checking the data for consistency, it is sent to the server.

[2276] Step 4:

[2277] Server: Analyzes the received text data using a natural language processing engine to extract the user's intent. Refers to the user profile and past data to generate the optimal response data.

[2278] Step 5:

[2279] Server: The generated response data is sent to the terminal in text format. The integrity of the data is verified when it is sent.

[2280] Step 6:

[2281] Terminal: The response data received from the server is converted into an audio file using a speech synthesis engine, and a voice response is provided to the user. For example, it might reply, "I think curry would be good for dinner tonight. The ingredients you need are..."

[2282] Advanced Security Settings

[2283] Step 1:

[2284] User: Select the security settings item from the device settings menu. Choose from various authentication options (corneal identification, fingerprint authentication, DNA authentication, etc.).

[2285] Step 2:

[2286] Terminal: Acquires the necessary biometric authentication data according to the selected authentication method. The acquired data is temporarily stored and encrypted.

[2287] Step 3:

[2288] Terminal: Sends encrypted authentication data to the server. Checks the integrity and completeness of the data as it is sent.

[2289] Step 4:

[2290] Server: Deserializes the encrypted biometric data and stores it in a secure database. Builds the authentication algorithm and associates it with the user profile.

[2291] Step 5:

[2292] Server: Sends the authentication system to the terminal and installs and configures the system within the terminal.

[2293] Step 6:

[2294] Terminal: A security authentication system is installed and allows the user access whenever authentication is successful.

[2295] Continuous learning phase

[2296] Step 1:

[2297] User: Uses a smartphone on a daily basis and utilizes various applications and functions.

[2298] Step 2:

[2299] Terminal: Automatically collects user behavior history and usage data. Collected data is temporarily stored and sent to a server as needed.

[2300] Step 3:

[2301] Server: Analyzes the received behavioral history data and updates the user profile. It uses machine learning algorithms to learn the user's behavioral patterns.

[2302] Step 4:

[2303] Server: Periodically synchronizes updated profiles and learning results to the device. Checks data consistency and completeness during synchronization.

[2304] Step 5:

[2305] Terminal: The updated data received from the server is reflected locally, providing more personalized services to the user.

[2306] In this way, by clarifying the specific processing and flow at each step, the system can be used effectively.

[2307] Example 1

[2308] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2309] Current systems that provide personalized responses lack the ability to quickly and accurately process user input and provide appropriate responses to individual questions and concerns in everyday life. Furthermore, continuous learning and profile updates based on user behavioral history are insufficient, resulting in issues with the accuracy and relevance of responses. Furthermore, in terms of security, measures to safely protect personal information are insufficient, and the authentication process is cumbersome, resulting in a lack of user convenience.

[2310] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2311] In this invention, the server includes: input means for inputting a user's personal information, personality, hobbies, preferences, and behavioral patterns; profile generation means for processing the information acquired from the input means; response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; voice output means for returning the responses generated by the response generation means by voice; learning means for recording and learning the user's behavioral history; update means for updating the profile generation means based on the information acquired by the learning means; voice recognition means for analyzing the user's voice and converting it into text data; transmission means for transmitting the text data converted by the voice recognition means to the response generation means; authentication means for performing high-security authentication; and initial learning model generation means for generating an individual learning model using the user profile and storing it on the server. This makes it possible to comprehensively process a user's diverse information to provide highly accurate personalized responses, and to improve the accuracy of responses through continuous learning, thereby realizing a safe and convenient authentication process.

[2312] The "input means" is a means for inputting data such as the user's personal information, personality, hobbies, preferences, and behavioral patterns.

[2313] The "profile generating means" is a means for generating a user profile based on information acquired from the input means.

[2314] The "answer generation means" is a means for generating an appropriate answer to a user's question or concern based on the generated user profile.

[2315] The "audio output means" is a means for returning the response generated by the response generating means to the user by voice.

[2316] The "learning means" is a means for recording the user's behavior history and for the system to continuously learn based on that data.

[2317] The "update means" is a means for updating the profile generation means based on new information acquired by the learning means.

[2318] The "voice recognition means" is a means having the function of analyzing the user's voice and converting it into text data.

[2319] The "transmitting means" is a means for transmitting the text data converted by the speech recognition means to the response generating means.

[2320] "Authentication means" refers to a means for performing an advanced authentication process to ensure user security.

[2321] The "initial learning model generating means" is a means for generating an individual learning model based on a user profile and storing it on a server.

[2322] The "voice synthesis means" is a means for converting the text data generated by the response generation means into voice.

[2323] The system of the present invention is designed to provide users with personalized responses to their everyday questions and concerns. The components and operation of this system will now be described in detail.

[2324] System configuration

[2325] The system generates an individual profile based on the user's input information and provides appropriate responses to the user's questions and concerns. The system consists of the following main components:

[2326] 1. Input Method

[2327] A user starts up a new smartphone and inputs personal information such as name, date of birth, personality, blood type, hobbies, and preferences. This information is provided to the system via input means.

[2328] 2. Profile Generation Method

[2329] The terminal stores the user information acquired from the input means in a temporary storage area. After verifying the integrity of the data, it sends it to the server. The server updates the database based on the received data and creates a user profile. Database software (e.g., MySQL or PostgreSQL) is used to generate the profile.

[2330] 3. Response Generation Method

[2331] Once the profile is created, the system generates responses to the user's questions and concerns. This response generation process uses a natural language processing engine (for example, IBM Watson's NLP capabilities) to generate the best possible response based on the user profile and past data.

[2332] 4. Audio output means

[2333] The generated response is sent in text format to the user device, which then uses a speech synthesis engine (e.g., Microsoft Azure's Text-to-Speech service) to convert the text into speech and provide the response to the user.

[2334] 5. Learning Methods

[2335] The system records the user's behavior history and continuously learns from it. This learning process is implemented using a machine learning framework (e.g., TensorFlow).

[2336] 6. Update method

[2337] The information obtained by the learning means is reflected in the profile generation means as appropriate, and the profile is updated, thereby improving the accuracy of the system over time.

[2338] 7. Voice Recognition Methods

[2339] When a user speaks a question, the device's voice input module recognizes the speech and converts it into text data. This process uses a speech recognition API (for example, Google's Speech-to-Text API).

[2340] 8. Means of Transmission

[2341] The text data converted by the speech recognition means is sent to the server via the transmission means.

[2342] 9. Authentication Methods

[2343] The system utilizes biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide a high level of security, ensuring that users' personal information is kept safe.

[2344] 10. Initial learning model generation method

[2345] The server generates an initial learning model based on the user profile and stores it on the server. This initial learning model is generated using a machine learning framework such as TensorFlow.

[2346] Specific examples

[2347] Example 1: Menu suggestion

[2348] When a user speaks to their smartphone, "What should I make today?", the voice input module recognizes the speech and converts it into text data. This text data is sent to the server and analyzed by a natural language processing engine. An appropriate menu is generated by referring to the user's past meal history, preferences, and refrigerator inventory information. The generated menu is sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and suggested to the user.

[2349] Example 2: Advice

[2350] When a user asks, "Work hasn't been going well lately. What should I do?", the voice input module converts the voice into text data. This data is sent to the server, where a natural language processing engine analyzes the user's work situation and past consultation history. The most appropriate advice is generated and sent in text format to the user's device, where it is converted into voice by a speech synthesis engine and conveyed to the user.

[2351] AI chatbot-generated prompts

[2352] Example prompt 1: "Generate a response when a user asks verbally, 'What's for dinner tonight?'"

[2353] Example prompt 2: "Generate advice for when a user says, 'I've been having trouble at work lately. What should I do?'"

[2354] In this way, the system of the present invention can provide personalized services based on the user's individual information, enriching the user's daily life and protecting it safely.

[2355] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2356] Step 1:

[2357] The user starts up the new smartphone and begins the initial setup, entering personal information such as name, date of birth, personality, blood type, hobbies, and preferences.

[2358] Input: Personal information that users enter into input forms.

[2359] Output: The personal information data entered.

[2360] Step 2:

[2361] The device receives user input and stores the entered data in a temporary storage area (the smartphone's RAM). At this point, the data is verified for integrity and sent to the server. The data is encrypted.

[2362] Input: Personal information data provided by the user.

[2363] Data processing: Data validation and encryption (e.g., AES encryption).

[2364] Output: The encrypted personal information data is sent to the server.

[2365] Step 3:

[2366] The server receives the data sent from the device and stores it in a database (MySQL, PostgreSQL, etc.). It generates a user profile based on the received initial data and creates an initial learning model.

[2367] Input: Encrypted personal information data.

[2368] Data processing: Decrypting data and saving it to the database, generating user profiles (executing SQL queries).

[2369] Output: An initial learning model and a user profile are generated and stored in a database.

[2370] Step 4:

[2371] The device receives the initial learning model from the server and saves it locally (in the smartphone's internal storage). A message is displayed to the user indicating that the setup is complete.

[2372] Input: The initial training model sent from the server.

[2373] Data processing: Saving the initial training model.

[2374] Output: The initial training model is saved locally and a message appears saying that the setup is complete.

[2375] Step 5:

[2376] The user speaks to the smartphone using voice. For example, "What do you recommend for dinner tonight?"

[2377] Input: User's voice input.

[2378] Step 6:

[2379] The device's voice input module recognizes speech and converts it into text data. This process uses a speech recognition API (such as Google's Speech-to-Text API). The converted text data is then sent to the server.

[2380] Input: Audio data.

[2381] Data processing: Recognizing voice data and converting it to text.

[2382] Output: Text data is sent to the server.

[2383] Step 7:

[2384] The server analyzes the received text data using a natural language processing engine (such as IBM Watson's NLP function) to extract the intent, and generates the optimal response by referencing the user profile and past data.

[2385] Input: Text data.

[2386] Data processing: Natural language processing, intent extraction, and response generation.

[2387] Output: The generated response text data.

[2388] Step 8:

[2389] The server sends the generated response in text format to the terminal.

[2390] Input: The generated response text data.

[2391] Output: The response text data is sent to the terminal.

[2392] Step 9:

[2393] The device uses a speech synthesis engine (such as Microsoft Azure's Text-to-Speech service) to convert the text data into speech and provide a response to the user. For example: "I think curry would be good for dinner tonight. The ingredients you need are..."

[2394] Input: Response text data.

[2395] Data processing: Converting text data into audio.

[2396] Output: A spoken response.

[2397] Step 10:

[2398] The learning method records the user's behavior history and the system continuously learns based on that data. This is implemented using a machine learning framework (such as TensorFlow).

[2399] Input: User behavior history data.

[2400] Data processing: Preprocessing data and training machine learning models.

[2401] Output: The trained model and updated model parameters.

[2402] Step 11:

[2403] The server updates the profile generation means based on the information obtained by the learning means, thereby improving the profile and increasing the accuracy of responses.

[2404] Input: trained model parameters.

[2405] Data processing: Profile update work.

[2406] Output: The updated user profile.

[2407] Step 12:

[2408] The authentication method uses biometric authentication technologies such as corneal identification, fingerprint authentication, and DNA authentication to provide high-security authentication.

[2409] Input: Biometric data.

[2410] Data Processing: Biometric data collection and authentication.

[2411] Output: Authentication result.

[2412] Step 13:

[2413] The initial learning model generating means generates an individual learning model based on the user profile and stores it in the server.

[2414] Input: User profile data.

[2415] Data processing: generating and saving learning models.

[2416] Output: Initial training model.

[2417] (Application example 1)

[2418] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2419] Customer service in traditional brick-and-mortar stores relies on the knowledge and experience of store staff, making it difficult to provide consistent, high-quality service to customers. In particular, responses to customer questions and concerns are not personalized, which leads to lower customer satisfaction. It is also difficult to accurately grasp daily customer behavior patterns and preferences, making it difficult to recommend optimal products. Furthermore, there are also issues with in-store security authentication, and personal information may not be adequately protected.

[2420] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2421] In this invention, the server includes: an input means for inputting a user's personality, hobbies, preferences, and daily behavioral patterns; a profile generation means for processing the information acquired from the input means; a response generation means for responding to the user's questions and concerns based on the profile generated by the profile generation means; a voice output means for returning the responses generated by the response generation means by voice; a learning means for recording and learning the user's behavioral history; an update means for updating the profile generation means based on the information acquired by the learning means; an authentication means for performing high-security authentication; a voice recognition means for analyzing the voice input and converting it into text data; a natural language processing means for analyzing the text data using a natural language processing engine and generating an optimal response by referring to the user's profile and past data; a voice synthesis means for converting the text data generated by the natural language processing means into voice; a robot that operates as a customer service assistant to support operations in a physical store; and a profile generation means for suggesting products in the physical store. This enables personalized responses to customer questions and concerns, thereby contributing not only to improved customer satisfaction but also to improved security.

[2422] "Input means" refers to a device or function for inputting the user's personality, hobbies, preferences, and daily behavior patterns.

[2423] The "profile generating means" refers to a device or function for processing information acquired from the input means and generating a user profile.

[2424] The "response generating means" refers to a device or function for responding to a user's questions or concerns based on the profile generated by the profile generating means.

[2425] The "audio output means" refers to a device or function for returning the response generated by the response generation means by voice.

[2426] The "learning means" refers to a device or function for recording and learning from the user's behavior history.

[2427] The "update means" refers to a device or function for updating the profile generation means based on the information acquired by the learning means.

[2428] "Authentication means" means a device or function for performing high-security authentication.

[2429] "Speech recognition means" refers to a device or function for analyzing voice input and converting it into text data.

[2430] "Natural language processing means" refers to a device or function that uses a natural language processing engine to analyze text data and generate optimal responses by referring to the user's profile and past data.

[2431] The "speech synthesis means" refers to a device or function for converting text data generated by the natural language processing means into speech.

[2432] A "customer service assistant" is a robot that assists with work in physical stores.

[2433] The "profile generation means for making product suggestions" refers to a device or function that generates a profile for making product suggestions in a physical store.

[2434] The system of the present invention aims to provide personalized responses to questions and concerns that users have about their daily lives in brick-and-mortar stores. Specific embodiments for carrying out the invention will now be described.

[2435] System configuration

[2436] The system consists of the following major components:

[2437] 1. Input Method

[2438] 2. Profile Generation Method

[2439] 3. Response Generation Method

[2440] 4. Audio output means

[2441] 5. Learning Methods

[2442] 6. Update method

[2443] 7. Authentication Methods

[2444] 8. Voice Recognition Methods

[2445] 9. Natural Language Processing Tools

[2446] 10. Speech synthesis means

[2447] 11. Customer Service Assistant

[2448] 12. Product proposal profile generation means

[2449] Details of each means are as follows.

[2450] Program Generation and Hardware / Software Usage

[2451] 1. Initial Setup Phase

[2452] User: Launches the application at the store entrance and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[2453] Terminal: The entered information is temporarily stored, encrypted, and sent to the server.

[2454] Server: Generates a user profile based on the received data and creates a personalized product recommendation model.

[2455] Terminal: Saves the initial model received from the server locally and displays a notification that the setup is complete.

[2456] 2. Operational Phase

[2457] User: Asks questions by voice in-store. Example: "What wines do you recommend these days?"

[2458] Terminal: The speech is converted into text using a speech recognition tool (Google Speech-to-Text API) and sent to the server.

[2459] Server: Analyzes the text data using a natural language processing engine (spaCy or BERT) and generates an appropriate response by referencing the user's profile and past data.

[2460] Server: Sends the generated response in text format to the terminal.

[2461] Terminal: Converts text to speech using a speech synthesis engine (Amazon Polly) and provides responses to the user.

[2462] Specific examples

[2463] Example 1

[2464] User: "What wines have you recommended recently?"

[2465] Device: Converts speech to text and sends it to the server.

[2466] Server: Generates a response based on the user's profile and past data, and sends text data to the device saying, "This time, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[2467] Terminal: A speech synthesis engine is used to convert text into speech and provide guidance to the user.

[2468] Prompt Sentence Examples

[2469] Please enter your name, gender, age, preferences, purchase history, and product interest information.

[2470] What wines do you recommend these days?

[2471] "Today, I recommend the red wine 'Montes Alpha'. It's full-bodied and goes well with meat dishes."

[2472] Hardware and software used

[2473] Hardware: Smartphones, smart glasses, in-store robots

[2474] software:

[2475] Speech recognition engine: Google Speech-to-Text API

[2476] Natural language processing engine: spaCy, BERT-based model

[2477] Speech synthesis engine: Amazon Polly

[2478] Database: MySQL, Firebase

[2479] Communication protocol: HTTPS, WebSocket

[2480] In this way, the system of the present invention can provide personalized responses to users' questions and concerns, thereby improving customer satisfaction in brick-and-mortar stores. Furthermore, advanced security features can protect users' personal information.

[2481] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2482] Step 1:

[2483] User: Launches the application at the entrance of a physical store and enters personal information (name, gender, age, preferences, purchase history, products of interest, etc.).

[2484] Input: Personal information

[2485] Output: Temporarily saved personal information

[2486] Specific operation: By entering information into the input form and pressing the "Submit" button, the data is stored in a temporary storage area on the device.

[2487] Step 2:

[2488] Terminal: The personal information entered is temporarily stored, encrypted for security purposes, and then sent to the server.

[2489] Input:Temporarily saved personal information

[2490] Output: Encrypted personal information data

[2491] Specific operation: Personal information is encrypted inside the device and sent to the server using the HTTPS communication protocol.

[2492] Step 3:

[2493] Server: Decrypts the received encrypted data and stores it in a database. Generates a user profile and creates a personalized product recommendation model.

[2494] Input: Encrypted personal information data

[2495] Output: User profile and proposed model

[2496] Specific operations: Decrypting encrypted data, running the user profile generation algorithm, and saving the generated profile and proposed model to a database.

[2497] Step 4:

[2498] Terminal: The terminal locally stores the user profile and proposed model received from the server and displays a notification to the user that the setup is complete.

[2499] Input: User profile and proposed model data

[2500] Output: Locally saved profile and proposed model, notification that setup is complete

[2501] Specific behavior: Save received data locally, display completion notification

[2502] Step 5:

[2503] User: Asks questions by dictating as they walk through the store. Example: "What wines do you recommend these days?"

[2504] Input: Voice data (question)

[2505] Output: Start speech recognition via speech input interface

[2506] Specific actions: Press the voice input button on your smartphone or smart glasses and dictate your question.

[2507] Step 6:

[2508] Terminal: Using a voice recognition tool (Google Speech-to-Text API), the voice data is converted into text data and sent to the server.

[2509] Input: Audio data

[2510] Output: Text data

[2511] Specific operations: Calling the speech recognition API, generating converted text data, and sending the text data to the server.

[2512] Step 7:

[2513] Server: Analyzes the text data using natural language processing tools (spaCy or BERT engine) and generates an appropriate response by referencing the user's profile and past data.

[2514] Input: Text data, user profile

[2515] Output: Response text data

[2516] What it does: Runs a natural language analysis engine, consults a user profile database, and generates the best possible response.

[2517] Step 8:

[2518] Server: Sends the generated response in text format to the terminal.

[2519] Input: Response text data

[2520] Output: Text data sent to the terminal

[2521] Specific operation: Processing to send text data to the terminal.

[2522] Step 9:

[2523] Terminal: The response text data is converted into speech using a speech synthesis engine (Amazon Polly) and a response is provided to the user.

[2524] Input: Response text data

[2525] Output: Audio data

[2526] Specific operations: Calling the speech synthesis API, generating converted speech data, and outputting speech from the speaker.

[2527] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2528] The system of the present invention is designed in combination with an emotion engine to provide personalized responses to users' questions and concerns in their daily lives. Specific embodiments for implementing the present invention will be described below.

[2529] System configuration

[2530] This system generates an individual profile based on the user's input and emotional information, and provides appropriate responses to the user's questions and concerns. The system is broadly composed of the following main components:

[2531] 1. Input Method

[2532] 2. Profile Generation Method

[2533] 3. Response Generation Method

[2534] 4. Audio output means

[2535] 5. Learning Methods

[2536] 6. Update method

[2537] 7. Authentication Methods

[2538] 8. Voice Recognition Methods

[2539] 9. Means of Transmission

[2540] 10. Emotion Recognition Engine

[2541] Program processing flow

[2542] Initial Setup Phase ...

Claims

1. an input means for inputting the user's personality, hobbies, preferences, and daily behavior patterns; a profile generating means for processing information acquired from the input means; a response generating means for responding to questions and concerns of the user based on the profile generated by the profile generating means; a voice output means for returning the response generated by the response generating means by voice; a learning means for recording and learning a user's behavior history; an update means for updating the profile generation means based on information acquired by the learning means; an authentication method for performing high-security authentication; A system including:

2. 10. The system of claim 1, wherein the system authenticates the user using one of corneal identification, fingerprint authentication, and DNA authentication.

3. A speech recognition means for analyzing a user's speech and converting it into text data; a transmitting means for transmitting the text data converted by the voice recognition means to the response generating means; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A