System

The system addresses the challenge of AI systems failing to understand user preferences by converting voice data to text, analyzing it, and generating personalized advice, thereby enhancing user trust and conversational experiences.

JP2026035233APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Current AI systems struggle to fully understand and remember a user's past conversations, preferences, and values, leading to a lack of personalized advice and ineffective addressing of personal concerns, particularly for elderly individuals living alone.

Method used

A system that acquires voice data, converts it into text, analyzes the text to extract keywords and context, updates a user profile, and generates personalized advice based on past conversation data and preferences, using natural language processing and machine learning algorithms.

Benefits of technology

The system provides personalized and specific advice by remembering user preferences and past conversations, enhancing user trust and improving conversational experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035233000001_ABST
    Figure 2026035233000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for obtaining voice data from a user and converting the voice data into text data; means for analyzing the converted text data and extracting keywords and context; means for updating a profile of the user based on the extracted keywords and context and storing past conversation data; means for generating appropriate advice and conversation for the user based on the updated profile and the past conversation data; and means for providing the generated advice and conversation to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Current AI systems have difficulty fully understanding and remembering a user's past conversations, preferences, and values. As a result, they are unable to provide personalized advice and are limited to general answers. They also have the problem of being unable to effectively address users' personal concerns or specific issues. The lack of a reliable AI partner is a major challenge, particularly for elderly people living alone and those with no one to turn to. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for acquiring voice data from a user and converting it into text data, and a means for analyzing the converted text data and extracting keywords and context. It also includes a means for updating a user's profile based on the extracted keywords and context, storing past conversation data, and generating appropriate advice and conversation for the user based on the updated profile and past conversation data. This system is capable of receiving a new request from the user, analyzing the request, and understanding the intent. Furthermore, by referencing past conversation data and the user's profile, the system can generate and notify optimal advice, thereby generating specific suggestions based on the user's preferences and values, and providing the suggestions if there is a match. In this way, it is possible to build individual trust with users and provide more personalized services.

[0006] "Voice data" refers to voice signal information input by the user.

[0007] "Text data" is information obtained by converting voice data into characters.

[0008] "Analysis" is the process of analyzing the user's input data and extracting meaning, keywords, and context.

[0009] "Keywords" are important words or phrases that appear in a user's speech.

[0010] "Context" is information that indicates the overall meaning and background information of the user's utterance.

[0011] A "profile" is a data set that contains personal information such as a user's preferences, values, and past conversation data.

[0012] "Advice" means advice or suggestions provided based on the user's situation or needs.

[0013] A "conversation" is a series of exchanges of dialogue between a user and a system.

[0014] "System" refers to the entire technology infrastructure that collects and analyzes user data, updates profiles, and generates and provides advice.

[0015] A "request" is a question or request that a user makes to the system.

[0016] "Notification" is the act of the system conveying advice or information to the user.

[0017] "Specific Offers" are advice or information that is individually tailored to the user based on their preferences and values. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes a series of means for collecting user input data, analyzing it, memorizing it, learning from it, and generating optimal advice. The specific operation of this system is shown below.

[0040] System Configuration

[0041] The system mainly consists of the following elements:

[0042] 1. Devices that collect user voice data

[0043] 2. A device with the processing power to convert voice data into text data

[0044] 3. A server that analyzes text data and extracts keywords and context

[0045] 4. A server that updates the user's profile based on the extracted information and stores past conversation data.

[0046] 5. Server and terminal that generates and notifies optimal advice based on the user's profile and past data

[0047] Program processing overview

[0048] 1. Data Collection

[0049] The user speaks to the device, "I had Italian food with my friends yesterday."

[0050] The device records this voice data and converts it into text data using a voice-to-text conversion engine.

[0051] The converted text data becomes "I ate Italian food with my friends yesterday."

[0052] 2. Data transmission and analysis

[0053] The terminal transmits the converted text data to the server.

[0054] The server receives the text data and analyzes it using natural language processing (NLP) techniques.

[0055] Keywords extracted through analysis include "yesterday," "friends," "Italian food," and "ate," and the context of each is also taken into consideration.

[0056] 3. Update your profile

[0057] The server adds new information to the user's profile, such as "I have a penchant for Italian food" or "I like spending time with friends."

[0058] The server uses machine learning algorithms to continuously update and learn from the user's profile.

[0059] 4. Generating and Providing Advice

[0060] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0061] The device records this voice data, converts it into text, and sends it to the server.

[0062] The server analyzes the request and uses historical profile data to generate appropriate advice.

[0063] For example, based on past data, suggestions might be generated such as, "I remember you enjoying Italian food. Why not try that newly opened restaurant?"

[0064] The device notifies the user of the generated advice.

[0065] Specific examples

[0066] Specifically, when a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend," the following process occurs.

[0067] The device records the voice data and converts it into text.

[0068] The server analyzes the text data and references past profile data.

[0069] The server takes into account factors such as "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past" and generates specific suggestions such as "the latest mystery novel."

[0070] The device provides the generated advice to the user, who can refer to it.

[0071] The system gives users a personalized AI partner that understands their preferences and past conversations, allowing them to receive faster, more relevant advice.

[0072] The processing flow will be explained below.

[0073] Step 1:

[0074] The user says to the device, "I had Italian food with my friends yesterday."

[0075] The device captures the user's speech as audio data.

[0076] Step 2:

[0077] The terminal converts the acquired voice data into text data using voice recognition technology.

[0078] The converted text data becomes "I ate Italian food with my friends yesterday."

[0079] Step 3:

[0080] The terminal transmits the converted text data to the server.

[0081] Step 4:

[0082] The server receives the text data and uses natural language processing (NLP) techniques to perform grammatical analysis and keyword extraction.

[0083] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0084] Step 5:

[0085] The server references past conversation data and the user's profile, and updates the profile based on newly extracted keywords and context.

[0086] For example, a user can add information to their profile such as "I like Italian food" or "I like spending time with friends."

[0087] Step 6:

[0088] The server uses machine learning models to learn user preferences and behavioral patterns and update the profile.

[0089] Step 7:

[0090] The user talks to the device, saying, "I'm wondering what to do on my day off."

[0091] The device acquires the user's speech as audio data and converts it into text data.

[0092] Step 8:

[0093] The terminal transmits the converted text data to the server.

[0094] Step 9:

[0095] The server receives the text data and analyzes the request using NLP techniques.

[0096] Keywords such as "holiday" and "not sure what to do" are extracted from the analysis results.

[0097] Step 10:

[0098] The server references past profile data and generates optimal advice based on the user's preferences and past behavior.

[0099] For example, a suggestion might be generated: "I remember enjoying Italian food. Why not try that newly opened restaurant?"

[0100] Step 11:

[0101] The server transmits the generated advice to the terminal.

[0102] Step 12:

[0103] The device will notify the user of the received advice in voice or text format.

[0104] This series of steps allows users to receive personalized advice, resulting in a more specific and reliable conversational experience.

[0105] Example 1

[0106] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0107] It is difficult for current AI systems to effectively memorize and analyze a user's past behavior and preferences and provide personalized advice based on the results. Conventional technologies lack sufficient user profile updates and past data reference for generating advice, resulting in a lack of improvement in the user experience. Therefore, there is a need for a system that can easily build a profile from a user's voice data and quickly provide appropriate advice.

[0108] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0109] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user using a generative AI model based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for receiving a new request from the user and analyzing the request to understand its intent, means for generating optimal advice by referring to past conversation data and the user's profile, means for generating specific suggestions based on the user's preferences and values, means for checking whether the suggested content matches the user's past actions and conversations, and means for providing the suggestion to the user if a match is found. This makes it possible to build a profile from the user's voice data and provide quick and accurate advice.

[0110] "Voice data" refers to digital data that records the voice uttered by the user.

[0111] "Text data" is digital data that has been converted from voice data into character information.

[0112] "Keywords" are important words or phrases extracted from text data.

[0113] "Context" refers to the semantic background, including the context of the keyword and the environment in which it is used.

[0114] A "user profile" is an individual data set that includes a user's preferences, past behavior, and conversational data.

[0115] "Generative AI models" are artificial intelligence algorithms and models used to analyze data and generate advice.

[0116] An "HTTP request" is a request of a communication protocol used to send data from a terminal to a server.

[0117] "Natural language processing" is a general term for technology that analyzes human language and extracts meaning.

[0118] A "machine learning algorithm" is a computational method for learning from data and making predictions or classifications.

[0119] "Advice" is a suggestion or advice provided by the system based on the user's profile and past behavior.

[0120] A "request" is a question or input instruction that a user makes to the system.

[0121] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes various means for collecting user input data, analyzing, memorizing, learning from it, and generating optimal advice. The specific operation of this system is described below.

[0122] System Configuration

[0123] The system mainly consists of the following elements:

[0124] 1. Devices that collect user voice data (smartphones, tablets, etc.)

[0125] 2. A speech-to-text engine (such as Google® Cloud Speech-to-Text) to convert voice data into text data.

[0126] 3. A server that analyzes text data and extracts keywords and context (using SpaCy or NLTK)

[0127] 4. A server that updates user profiles based on the extracted information and stores past conversation data (using Scikit-learn and Tensorflow (registered trademark))

[0128] 5. Server and device that generates and notifies optimal advice based on the user's profile and past data (using generative AI models)

[0129] Data collection

[0130] The user asks questions or makes requests to the system in natural language. For example, they might say, "I had Italian food with my friends yesterday." This voice data is recorded by the device and converted into text data.

[0131] Data analysis and profile updates

[0132] The converted text data is sent to a server and analyzed using natural language processing technology. For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted from the text data. The server then adds this information to the user's profile and stores it as past conversation data.

[0133] Generating and Serving Advice

[0134] When the user makes a new request, such as "I'm not sure what to do on my day off," the device re-records this voice data, converts it to text data, and sends it to the server. The server analyzes the request and generates appropriate advice by referencing past profile data. A generative AI model is used to create appropriate suggestions based on the user's past data. For example, a suggestion might be, "I remember you enjoyed Italian food. Why not try that newly opened restaurant?"

[0135] Examples of concrete examples and prompts

[0136] As a concrete scenario, consider the case where a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend." In this case, the following flow will occur.

[0137] The device records audio data and converts it to text using Google Cloud Speech-to-Text.

[0138] The server analyzes the received text data using SpaCy and extracts keywords such as "friend," "birthday," "present," and "worried."

[0139] The server refers to the user profile and takes into account things like "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past."

[0140] The server uses a generative AI model to generate specific suggestions such as "latest mystery novels."

[0141] The device notifies the user of the generated advice.

[0142] Users can use the suggestions to help them choose a gift.

[0143] Examples of prompts that users can enter include, "I'm not sure what to do on my day off. Could you give me some advice?" or "I'm having trouble deciding on a birthday present for my friend." Based on these prompts, the system can provide users with quick and accurate advice.

[0144] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0145] Program processing flow

[0146] Step 1:

[0147] The user speaks. Specifically, the user says to the terminal, "I ate Italian food with my friends yesterday," and voice data is generated.

[0148] Input: Audio data

[0149] Output: Audio data

[0150] Step 2:

[0151] The device converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text, for example, "I ate Italian food with my friends yesterday."

[0152] Specific operation: The device calls the voice recognition software and saves the voice data as text data.

[0153] Input: Audio data

[0154] Output: Text data

[0155] Step 3:

[0156] The terminal sends the converted text data to the server in JSON format using an HTTP request.

[0157] Specific operation: The terminal creates an HTTP request and sends text data as packets to the server.

[0158] Input: Text data

[0159] Output: HTTP request to the server

[0160] Step 4:

[0161] The server analyzes the received text data and extracts keywords and context from the text data using natural language processing tools (e.g., SpaCy, NLTK). For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted.

[0162] Specific operation: The server calls a natural language processing tool to break down the text data into keywords and context.

[0163] Input: Text data

[0164] Output: Extracted keywords and context data

[0165] Step 5:

[0166] The server updates the user's profile based on the extracted information and remembers past conversation data. Using machine learning algorithms (e.g., Scikit-learn, TensorFlow), the profile is updated and stored in a database. For example, profile information such as "I like Italian food" and "I like spending time with friends" are added.

[0167] What it does: The server uses machine learning algorithms to update the profile and save it in a database.

[0168] Input: Keywords and contextual data

[0169] Output: Updated user profile

[0170] Step 6:

[0171] The user issues a new request, for example, saying to the device, "I'm not sure what to do on my day off," and voice data is generated again.

[0172] Input: New audio data

[0173] Output: New audio data

[0174] Step 7:

[0175] The device converts this new voice data back into text data and sends it to the server. Google Cloud Speech-to-Text converts the voice data into text data, creates an HTTP request, and sends it to the server.

[0176] Specific operation: The device calls the voice recognition software, converts the voice data into text data, creates an HTTP request, and sends it to the server.

[0177] Input: New audio data

[0178] Output: Text data, HTTP request to server

[0179] Step 8:

[0180] The server analyzes the new text data and generates appropriate advice by referencing past profile data. A generative AI model (such as GPT-3®) is used to generate advice based on past profile data, such as "I remember you enjoying Italian food. Why don't you try that newly opened restaurant?"

[0181] How it works: The server calls the generative AI model and generates new advice by referencing past profile data.

[0182] Input: New text data, user profile

[0183] Output: Generated advice

[0184] Step 9:

[0185] The device will notify the user of the generated advice, which will be provided to the user via text message or voice message.

[0186] Specific operation: The device displays or audibly notifies the user of the generated advice via the user interface.

[0187] Input: Generated advice

[0188] Output: User notification

[0189] (Application example 1)

[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0191] In recent years, mail-order sales systems have been required to provide personalized product recommendations to users. However, conventional systems often fail to fully utilize the user's past conversations and preferences, and can only provide general recommendations. Therefore, a system that can provide accurate and individualized product recommendations to users is needed.

[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0193] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for generating a proposal for a specific product based on the generated advice, and means for notifying the user of the proposed product. This makes it possible to make personalized product proposals by utilizing the user's past conversations and preferences.

[0194] "Voice data" refers to data that is a digital recording of voice information uttered by a user.

[0195] "Text data" is voice data expressed as text information.

[0196] "Keywords" are contextually significant words or phrases extracted from text data through analysis.

[0197] "Context" refers to the additional information or meaning that comes from the context in which a keyword is used.

[0198] A "profile" is a collection of personalized data that includes a user's preferences, past conversations, behavioral patterns, etc.

[0199] "Advice" refers to instructions or suggestions for actions or choices that are generated based on a user's profile and past conversation data.

[0200] "Product suggestions" are suggestions for specific products that users are encouraged to purchase, generated based on their profile and past conversation data.

[0201] "Notification" is a means of communication used to communicate system-generated advice and product suggestions to users.

[0202] This invention relates to a system that can memorize a user's past conversations and preferences and provide personalized advice and product suggestions based on them. The system includes a series of means for collecting, analyzing, memorizing, and learning from user input data to generate optimal advice and product suggestions.

[0203] System Configuration

[0204] The system mainly consists of the following elements:

[0205] 1. Speech recognition engine: Captures the user's voice data and converts it into text data. Specifically, this is done by collecting voice data using the microphone on a smartphone or smart glasses, and using voice recognition software such as the Google Speech-to-Text API.

[0206] 2. Text analysis engine: Analyzes text data and extracts keywords and context using natural language processing techniques such as Hugging Face's Transformer model.

[0207] 3. Profile Management Server: Updates the user's profile based on extracted keywords and context and remembers past conversation data.

[0208] 4. Advice Generation Engine: Generates appropriate advice and product suggestions for users based on their updated profile and past conversation data. Uses machine learning algorithms such as TensorFlow.

[0209] 5. Notification system: Provides generated advice and product suggestions to users. Software for notifying users of advice and product suggestions via smartphones or smart glasses.

[0210] Example of operation

[0211] When a user says to their smartphone, "I'm looking for a birthday present for my friend," the system operates as follows:

[0212] 1. Speech capture and text conversion:

[0213] The server converts the user's voice into text using the Google Speech-to-Text API.

[0214] 2. Text Analysis:

[0215] The text analysis engine uses Hugging Face's Transformer model to extract keywords such as "friend," "birthday," and "present" from the text data.

[0216] 3. Profile Update:

[0217] The profile management server updates the user's profile by referencing the extracted keywords and past data.

[0218] 4. Generating advice and suggestions:

[0219] An advice generation engine generates optimal birthday gift suggestions based on the updated profile.

[0220] 5. Notice:

[0221] A notification system sends generated suggestions to the user, such as "What new mystery novels or popular electronic gadgets might this friend like?"

[0222] Examples of AI-generated prompts

[0223] Prompts to suggest birthday gift ideas based on user voice data:

[0224] User dictation: "What would be a good birthday gift for my friend?"

[0225] Converts speech to text.

[0226] The converted text is then analyzed for keywords and context using a natural language processing model.

[0227] Extract important keywords such as "friends," "birthday," and "presents."

[0228] It retrieves related products from the database and notifies the user.

[0229] Example of desired result:

[0230] "How about a new mystery novel or a popular electronic gadget that this friend might like?"

[0231] This allows users to receive personalized advice based on an understanding of their preferences and past conversations, enabling them to choose the best product.

[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0233] Step 1:

[0234] A user speaks to their smartphone, for example, "I'm looking for a birthday present for my friend." The smartphone's microphone captures the voice data, which is sent to the Google Speech-to-Text API. The Google Speech-to-Text API converts the voice data into text data, which becomes "I'm looking for a birthday present for my friend." The input is voice data, and the output is text data.

[0235] Step 2:

[0236] The text data is sent to the server. The server analyzes the text data using Hugging Face's Transformer model. This analysis extracts important keywords in the text, such as "friend," "birthday," and "present," as well as their context. The input is the text data, and the output is the keywords and context.

[0237] Step 3:

[0238] The server updates the user's profile based on the extracted keywords and context. It references the user's past conversation data and adds new information about "friends," "birthdays," and "gifts." This profile is stored in a database. The input is keywords and context, and the output is the updated profile.

[0239] Step 4:

[0240] When the user makes a new request, such as "Tell me what kind of gift would be good," the device receives this voice data and converts it into text data using the Google Speech-to-Text API again. The input is voice data, and the output is text data.

[0241] Step 5:

[0242] The server analyzes the new text data and again uses the Hugging Face Transformer model to understand its intent. It analyzes important keywords and context within the text and extracts intents such as "present" or "birthday." The input is the new text data, and the output is the intended keywords.

[0243] Step 6:

[0244] The server references past conversation data and the user's profile data to generate specific suggestions, such as "new mystery novels" or "popular electronic gadgets," using machine learning algorithms such as TensorFlow. The input is the updated profile and past data, and the output is specific product suggestions.

[0245] Step 7:

[0246] The server notifies the smartphone of the generated product suggestions. The smartphone screen displays a message saying, "How about a new mystery novel or a popular electronic gadget that your friend might like?" The input is the product suggestions, and the output is a notification to the user.

[0247] By going through the above processing steps, the user can receive personalized product suggestions based on past conversations and preferences.

[0248] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0249] This invention relates to a system that acquires a user's voice data and text data, analyzes that data, and provides personalized advice. In particular, by combining it with an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. The specific operating method and program processing of this system are described below.

[0250] System Configuration

[0251] The system consists of the following components:

[0252] 1. Device that collects user voice data

[0253] 2. A device with the processing power to convert voice data into text data

[0254] 3. A server that analyzes text data and extracts keywords and context

[0255] 4. A server that updates user profiles based on the extracted information and stores past conversation data.

[0256] 5. A server that uses an emotion engine to recognize user emotions and update the profile.

[0257] 6. Server and device that generates and notifies optimal advice based on the user's profile, past data, and emotional information

[0258] Program processing overview

[0259] 1. Data Collection

[0260] The user speaks to the device, "I had Italian food with my friends yesterday."

[0261] The device records this audio data and converts it into text data using a speech-to-text engine.

[0262] The converted text data becomes "I ate Italian food with my friends yesterday."

[0263] 2. Data transmission and analysis

[0264] The terminal transmits the converted text data to the server.

[0265] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[0266] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0267] 3. Emotion recognition

[0268] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[0269] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[0270] 4. Update your profile

[0271] The server updates the user's profile based on the new information and the recognized emotion information.

[0272] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[0273] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[0274] 5. Generating and Providing Advice

[0275] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0276] The terminal records this voice data, converts it into text data, and sends it to the server.

[0277] The server analyzes the request and generates the best advice using an emotion engine, taking into account the user's current emotions.

[0278] For example, if the user is in a relaxed state, the generated advice is, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0279] The device notifies the user of the generated advice in voice or text format.

[0280] Specific examples

[0281] For example, if a user asks for advice on what to buy as a birthday present for a friend, the following process will occur:

[0282] The terminal records the voice data and converts it into text data.

[0283] The server analyzes the text data and references past profile data and emotion information.

[0284] The server takes into consideration "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions such as "the latest mystery novel."

[0285] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[0286] The system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[0287] The processing flow will be explained below.

[0288] Step 1:

[0289] The user says to the device, "I had Italian food with my friends yesterday."

[0290] The device captures the user's speech as audio data.

[0291] Step 2:

[0292] The terminal converts the acquired voice data into text data using voice recognition technology.

[0293] The converted text data becomes "I ate Italian food with my friends yesterday."

[0294] Step 3:

[0295] The terminal transmits the converted text data to the server.

[0296] Step 4:

[0297] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[0298] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0299] Step 5:

[0300] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[0301] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[0302] Step 6:

[0303] The server updates the user's profile based on newly extracted keywords, context and recognized emotion information.

[0304] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[0305] Step 7:

[0306] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[0307] Step 8:

[0308] The user talks to the device, saying, "I'm wondering what to do on my day off."

[0309] The device captures the user's speech as audio data and converts it into text data.

[0310] Step 9:

[0311] The terminal transmits the converted text data to the server.

[0312] Step 10:

[0313] The server analyzes the received text data to understand the intent and emotional information of the request.

[0314] For example, the keywords "holiday" and "not sure what to do" and the user's relaxed state are analyzed.

[0315] Step 11:

[0316] The server references past conversation data and profiles to generate optimal advice based on the user's preferences and past behavior.

[0317] For example, if the user is in a relaxed state, the system generates advice such as, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0318] Step 12:

[0319] The server transmits the generated advice to the terminal.

[0320] Step 13:

[0321] The device will notify the user of the received advice in voice or text format.

[0322] This series of steps allows users to receive faster, more personalized advice that adapts to their current emotional state and past behavior.

[0323] Example 2

[0324] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0325] Conventional user assistance systems simply convert user voice data into text data and extract context and keywords, making it difficult to provide personalized advice that takes into account the user's emotional state. Therefore, there is a need for a system that can generate and provide more appropriate advice based on the user's emotions, thereby improving the user experience.

[0326] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0327] In this invention, the server includes means for acquiring a user's voice data and converting it into text data, means for analyzing the text data and extracting keywords and context, and means for analyzing the user's emotional information using an emotion engine and updating the profile, thereby enabling the server to generate and provide appropriate advice and communication that takes into account the user's emotional state.

[0328] "Audio data from a user" refers to data that is a recording of what a user says.

[0329] "Text data" refers to character information obtained by analyzing voice data.

[0330] A "keyword" is a word or phrase that indicates important information from text data.

[0331] "Context" refers to information that indicates the relationship between keywords and the flow of the entire sentence.

[0332] "Profile" means data that records the characteristics and preferences of a user.

[0333] An "emotion engine" is software for analyzing emotional states from voice and text.

[0334] "Content data" refers to data relating to conversations and actions with users that have been recorded in the past.

[0335] "Advice" means any recommendation or instruction provided to a User.

[0336] "Communication" refers to activities that involve dialogue and exchange of information with users.

[0337] A "request" is a question or request made by a user to the system.

[0338] "Intent" refers to the goal or desire the user wants to achieve through their request.

[0339] "Past behavior" refers to the activities and behavioral history that a user has undertaken to date.

[0340] "Values" refer to the fundamental beliefs and ways of thinking that govern a user's thoughts and actions.

[0341] The present invention relates to a system that acquires a user's voice data and text data, analyzes them, and provides personalized advice. In particular, by combining an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. This system is implemented using the following hardware and software.

[0342] System configuration and hardware / software used

[0343] Hardware

[0344] 1. Device: A digital device with a microphone (e.g., smart speaker, smartphone) that captures user voice data.

[0345] 2. Server: A cloud or on-premise server for managing data analysis and profiles.

[0346] software

[0347] 1. Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[0348] 2. Natural Language Processing (NLP) engine: SpaCy is used to perform contextual analysis of text data and extract keywords.

[0349] 3. Emotion Recognition Engine: Analyzes the user's emotional information using IBM Watson (registered trademark) Tone Analyzer.

[0350] 4. Machine learning algorithm: TensorFlow is used to continuously update and learn user profiles and emotion data.

[0351] 5. Advice Generation Engine: Uses OpenAI's GPT-4 model to generate optimal advice.

[0352] 6. Web Framework: Use Django Web Framework for profile management and database operations.

[0353] Specific system operation examples

[0354] Data collection

[0355] The user speaks to the device, "I had Italian food with my friends yesterday."

[0356] The device records this voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data will be "I ate Italian food with my friends yesterday."

[0357] Data analysis and profile updates

[0358] The device sends the converted text data to the server using the REST API.

[0359] The server receives the text data and uses SpaCy to perform grammatical analysis and extract keywords. The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0360] The server then uses IBM Watson Tone Analyzer to analyze emotional information from the voice and text data and recognize emotions such as "happy" or "relaxed."

[0361] The server updates the user's profile with the new information and sentiment, adding traits such as "enjoys Italian food" and "is relaxed when with friends" to the database using the Django Web Framework.

[0362] Advice generation and delivery

[0363] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0364] The device records this voice data, converts it into text data using the Google Cloud Speech-to-Text API, and sends it to the server.

[0365] The server analyzes the request, takes into account the user's current emotions using an emotion engine, and then generates the optimal advice using OpenAI's GPT-4 model. For example, if the user is in a relaxed state, the advice generated might be, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0366] The device notifies the user of the generated advice in voice or text format.

[0367] Specific examples

[0368] As a specific example, consider the case where a user asks for advice, saying, "I'm having trouble deciding on a birthday present for my friend."

[0369] The device records audio data and converts it into text data using the Google Cloud Speech-to-Text API.

[0370] The server analyzes the text data and references past profile data and emotional information using IBM Watson Tone Analyzer.

[0371] Next, the server uses OpenAI's GPT-4 model to consider "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions, such as "the latest mystery novel."

[0372] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[0373] In this way, the system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[0374] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0375] Step 1:

[0376] The user speaks to the device, "I had Italian food with my friends yesterday."

[0377] Input: User's voice data.

[0378] To record this audio data, the device uses a microphone to capture the audio.

[0379] Output: Recorded audio data file.

[0380] Step 2:

[0381] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts it into text data.

[0382] Input: Audio data file.

[0383] The Google Cloud Speech-to-Text API uses speech recognition technology to analyze audio data and generate corresponding text data.

[0384] Output: Text data: "I ate Italian food with my friends yesterday."

[0385] Step 3:

[0386] The device sends the converted text data to the server using the REST API.

[0387] Input: Text data.

[0388] The device packages the text data in JSON format and sends it to the server using an HTTP request.

[0389] Output: The text data received on the server.

[0390] Step 4:

[0391] The server analyzes the received text data using SpaCy, performing grammatical analysis and keyword extraction.

[0392] Input: Text data.

[0393] The SpaCy engine analyzes text data and extracts keywords such as "yesterday," "friends," "Italian food," and "ate." It also performs contextual analysis to understand the relationships between keywords.

[0394] Output: Extracted keywords and context information.

[0395] Step 5:

[0396] The server uses IBM Watson Tone Analyzer to analyze emotional information from the extracted text data.

[0397] Input: Keywords and contextual information.

[0398] IBM Watson Tone Analyzer uses emotion recognition algorithms to identify emotions such as "happy" or "relaxed" from the tone and content of text.

[0399] Output: Recognized emotion information.

[0400] Step 6:

[0401] The server uses the Django Web Framework to update the user's profile based on the new information and the recognized emotion information.

[0402] Input: Recognized emotion information and extracted keywords.

[0403] The profile database stores the new information and adds characteristics to the profile, such as "I enjoy Italian food" and "I feel relaxed when I'm with friends."

[0404] Output: Updated user profile.

[0405] Step 7:

[0406] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0407] Input: The user's new voice data.

[0408] The terminal re-records this voice data, converts it into text data, and sends it to the server.

[0409] Output: The new text data received on the server.

[0410] Step 8:

[0411] The server parses the request, again using an emotion engine to take into account the user's current emotions, and generates the best advice using OpenAI's GPT-4 model.

[0412] Input: New text data and existing user profile.

[0413] OpenAI's GPT-4 model combines past profile information with your current emotional state to generate a recommendation like, "Why not try a new restaurant to rediscover that Italian food you enjoyed before?"

[0414] Output: The generated advice.

[0415] Step 9:

[0416] The terminal provides the generated advice to the user.

[0417] Input: The generated advice.

[0418] Using a text-to-speech engine, the device will verbally inform the user, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0419] Output: The user receives the advice.

[0420] (Application example 2)

[0421] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0422] Modern online shopping sites lack personalized product recommendations that utilize users' voice data and emotional information. In particular, there is a need for a system that can improve the user experience by providing optimal advice and products based on the user's emotional state. The lack of such personalized services is causing users to lose their motivation to purchase and making it difficult to select the right product.

[0423] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data from the user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and product recommendations for the user based on the updated profile, past conversation data, and emotional information, and means for providing the generated advice and product recommendations to the user. This enables personalized product recommendations that take into account the user's emotional state and past behavior.

[0424] "User's voice data" refers to voice information uttered by a user through a voice input device, and is data including voice signals such as conversations and instructions.

[0425] "Text data" is voice data converted into character information, and is data consisting of a sequence of characters.

[0426] "Keywords" are important words or phrases extracted from text data that are meaningful for analyzing context.

[0427] "Context" refers to the background and situation of sentences and phrases present in text data, and is information that helps understand the meaning.

[0428] A "profile" is a user information system constructed by combining a user's attribute information, behavioral history, preferences, emotional state, etc.

[0429] "Emotional information" is data that indicates emotional states such as joy, sadness, anger, and stress, analyzed from a user's statements and actions.

[0430] "Product recommendation" refers to the act and content of suggesting specific products or services to users based on their profile and emotional information.

[0431] "Appropriate advice" means providing useful information and guidance tailored to the user's needs and circumstances.

[0432] "Means for acquiring voice data" refers to a device or method for collecting voice information from a user using a microphone or the like.

[0433] The "means for converting voice data into text data" refers to a device or method for converting voice data into text information using voice recognition technology.

[0434] The "means for analyzing converted text data" refers to a device or method that uses natural language processing technology to analyze text data and extract context and meaning.

[0435] The "means for storing past conversation data" refers to a device or method for storing the history of dialogue with the user in a database or the like.

[0436] The "means for generating advice and product recommendations" is a device or method for generating appropriate advice and product recommendations based on the profile and emotion information.

[0437] A "means for providing generated advice and product recommendations" is a device or method that notifies a user of advice or product recommendations in audio or text form.

[0438] This invention is a system that acquires a user's voice data, converts it into text data, and provides personalized advice and product recommendations that incorporate emotional information. This system improves the user experience and helps them make more appropriate product selections.

[0439] The general flow of the system is as follows: First, the user inputs voice data using a device such as a smartphone. The voice data is acquired in real time and converted into text data using the Google Cloud Speech-to-Text API. This text data is then sent to the server.

[0440] After receiving the converted text data, the server analyzes it using the Google Cloud Natural Language API to extract keywords and context. It then uses IBM Watson Tone Analyzer to analyze emotional information from the text data. The server then updates the user's profile based on the emotional information and the extracted keywords and context.

[0441] When updating profiles, machine learning algorithms such as TensorFlow are used to sequentially learn and update the user's past conversation data and behavioral history, resulting in personalized information that reflects the user's preferences and emotional state in real time.

[0442] Next, when the user makes a new request, the system analyzes the request and references past conversation data and profiles to generate optimal advice and product recommendations, which are then sent to the user via a device such as a smartphone.

[0443] As a concrete example, if a user says, "I've been feeling stressed lately, so I want something that will help me relax," the following processing will occur.

[0444] First, the smartphone's microphone is used to capture voice data, which is then converted into text using the Google Cloud Speech-to-Text API. The converted text data is then sent to a server, where it is analyzed using the Google Natural Language API to extract keywords such as "stress" and "relaxation." Next, IBM Watson Tone Analyzer is used to analyze the emotional data and identify high stress levels.

[0445] Based on this information, the server uses machine learning algorithms such as TensorFlow to update the profile and generate appropriate product recommendations. For example, recommendations such as "Lavender aroma diffuser" or "Soothing music CD" are generated. The generated advice and product recommendations are then displayed on the user's smartphone as a notification saying, "How about a lavender aroma diffuser to relieve stress?"

[0446] The following is a specific example of a prompt sentence to be input to the generative AI model:

[0447] Prompt: "The user says, 'I'm tired and looking for something to relax.' Speech is converted to text and the keywords 'tired' and 'relaxation' are extracted. Sentiment analysis identifies this as a state of high stress. Based on the user's profile, it is determined that they need something to relax, and the system recommends the purchase of a 'lavender aroma diffuser,' which has a relaxing effect."

[0448] The system enables personalized product recommendations that take into account the user's emotional state and past behavior, improving the user experience.

[0449] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0450] Step 1:

[0451] The user speaks into the smartphone's microphone, saying, "I've been feeling stressed lately, so I want something that will help me relax." The input voice data is picked up by the smartphone's microphone. The output is the original voice data.

[0452] Step 2:

[0453] The device sends the acquired voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data and outputs the converted text data.

[0454] Step 3:

[0455] The terminal transmits the converted text data to the server, and the server receives the text data.

[0456] Step 4:

[0457] The server uses the Google Cloud Natural Language API to parse the received text data and extract keywords and context. The input is the received text data, and the output is the extracted keywords (e.g., "stress," "relax") and context information.

[0458] Step 5:

[0459] The server uses IBM Watson Tone Analyzer to analyze emotional information from text data. The input is text data, and the output is emotional data (e.g., high stress).

[0460] Step 6:

[0461] The server updates the user's profile based on the extracted keywords and emotion data. The profile is updated sequentially using machine learning algorithms such as TensorFlow based on the user's past conversation data and behavioral history. The input is keywords, context, emotion data, and past profile data, and the output is the updated profile.

[0462] Step 7:

[0463] When a user inputs a new request, the server analyzes the request. For example, if a user inputs "I want a relaxation item," the input voice data is converted into text data and analyzed. The input is the user's new request, and the output is the analysis result.

[0464] Step 8:

[0465] The server refers to the past conversation data and the updated profile to generate optimal advice and product recommendations. For example, a "lavender aroma diffuser" is recommended. The input is the parsed request, the past conversation data, and the updated profile, and the output is the generated advice and product recommendations.

[0466] Step 9:

[0467] The terminal notifies the user of the generated advice and product recommendation. For example, a notification such as "How about a lavender aroma diffuser to relieve stress?" is displayed. The input is the generated advice and product recommendation, and the output is a notification to the user.

[0468] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0469] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0470] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0471] [Second embodiment]

[0472] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0473] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0474] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0475] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0476] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0477] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0478] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0479] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0480] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0481] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0482] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0483] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0484] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes a series of means for collecting user input data, analyzing it, memorizing it, learning from it, and generating optimal advice. The specific operation of this system is shown below.

[0485] System Configuration

[0486] The system mainly consists of the following elements:

[0487] 1. Devices that collect user voice data

[0488] 2. A device with the processing power to convert voice data into text data

[0489] 3. A server that analyzes text data and extracts keywords and context

[0490] 4. A server that updates the user's profile based on the extracted information and stores past conversation data.

[0491] 5. Server and terminal that generates and notifies optimal advice based on the user's profile and past data

[0492] Program processing overview

[0493] 1. Data Collection

[0494] The user speaks to the device, "I had Italian food with my friends yesterday."

[0495] The device records this voice data and converts it into text data using a voice-to-text conversion engine.

[0496] The converted text data becomes "I ate Italian food with my friends yesterday."

[0497] 2. Data transmission and analysis

[0498] The terminal transmits the converted text data to the server.

[0499] The server receives the text data and analyzes it using natural language processing (NLP) techniques.

[0500] Keywords extracted through analysis include "yesterday," "friends," "Italian food," and "ate," and the context of each is also taken into consideration.

[0501] 3. Update your profile

[0502] The server adds new information to the user's profile, such as "I have a penchant for Italian food" or "I like spending time with friends."

[0503] The server uses machine learning algorithms to continuously update and learn from the user's profile.

[0504] 4. Generating and Providing Advice

[0505] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0506] The device records this voice data, converts it into text, and sends it to the server.

[0507] The server analyzes the request and uses historical profile data to generate appropriate advice.

[0508] For example, based on past data, suggestions might be generated such as, "I remember you enjoying Italian food. Why not try that newly opened restaurant?"

[0509] The device notifies the user of the generated advice.

[0510] Specific examples

[0511] Specifically, when a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend," the following process occurs.

[0512] The device records the voice data and converts it into text.

[0513] The server analyzes the text data and references past profile data.

[0514] The server takes into account factors such as "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past" and generates specific suggestions such as "the latest mystery novel."

[0515] The device provides the generated advice to the user, who can refer to it.

[0516] The system gives users a personalized AI partner that understands their preferences and past conversations, allowing them to receive faster, more relevant advice.

[0517] The processing flow will be explained below.

[0518] Step 1:

[0519] The user says to the device, "I had Italian food with my friends yesterday."

[0520] The device captures the user's speech as audio data.

[0521] Step 2:

[0522] The terminal converts the acquired voice data into text data using voice recognition technology.

[0523] The converted text data becomes "I ate Italian food with my friends yesterday."

[0524] Step 3:

[0525] The terminal transmits the converted text data to the server.

[0526] Step 4:

[0527] The server receives the text data and uses natural language processing (NLP) techniques to perform grammatical analysis and keyword extraction.

[0528] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0529] Step 5:

[0530] The server references past conversation data and the user's profile, and updates the profile based on newly extracted keywords and context.

[0531] For example, a user can add information to their profile such as "I like Italian food" or "I like spending time with friends."

[0532] Step 6:

[0533] The server uses machine learning models to learn user preferences and behavioral patterns and update the profile.

[0534] Step 7:

[0535] The user talks to the device, saying, "I'm wondering what to do on my day off."

[0536] The device acquires the user's speech as audio data and converts it into text data.

[0537] Step 8:

[0538] The terminal transmits the converted text data to the server.

[0539] Step 9:

[0540] The server receives the text data and analyzes the request using NLP techniques.

[0541] Keywords such as "holiday" and "not sure what to do" are extracted from the analysis results.

[0542] Step 10:

[0543] The server references past profile data and generates optimal advice based on the user's preferences and past behavior.

[0544] For example, a suggestion might be generated: "I remember enjoying Italian food. Why not try that newly opened restaurant?"

[0545] Step 11:

[0546] The server transmits the generated advice to the terminal.

[0547] Step 12:

[0548] The device will notify the user of the received advice in voice or text format.

[0549] This series of steps allows users to receive personalized advice, resulting in a more specific and reliable conversational experience.

[0550] Example 1

[0551] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0552] It is difficult for current AI systems to effectively memorize and analyze a user's past behavior and preferences and provide personalized advice based on the results. Conventional technologies lack sufficient user profile updates and past data reference for generating advice, resulting in a lack of improvement in the user experience. Therefore, there is a need for a system that can easily build a profile from a user's voice data and quickly provide appropriate advice.

[0553] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0554] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user using a generative AI model based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for receiving a new request from the user and analyzing the request to understand its intent, means for generating optimal advice by referring to past conversation data and the user's profile, means for generating specific suggestions based on the user's preferences and values, means for checking whether the suggested content matches the user's past actions and conversations, and means for providing the suggestion to the user if a match is found. This makes it possible to build a profile from the user's voice data and provide quick and accurate advice.

[0555] "Voice data" refers to digital data that records the voice uttered by the user.

[0556] "Text data" is digital data that has been converted from voice data into character information.

[0557] "Keywords" are important words or phrases extracted from text data.

[0558] "Context" refers to the semantic background, including the context of the keyword and the environment in which it is used.

[0559] A "user profile" is an individual data set that includes a user's preferences, past behavior, and conversational data.

[0560] "Generative AI models" are artificial intelligence algorithms and models used to analyze data and generate advice.

[0561] An "HTTP request" is a request of a communication protocol used to send data from a terminal to a server.

[0562] "Natural language processing" is a general term for technology that analyzes human language and extracts meaning.

[0563] A "machine learning algorithm" is a computational method for learning from data and making predictions or classifications.

[0564] "Advice" is a suggestion or advice provided by the system based on the user's profile and past behavior.

[0565] A "request" is a question or input instruction that a user makes to the system.

[0566] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes various means for collecting user input data, analyzing, memorizing, learning from it, and generating optimal advice. The specific operation of this system is described below.

[0567] System Configuration

[0568] The system mainly consists of the following elements:

[0569] 1. Devices that collect user voice data (smartphones, tablets, etc.)

[0570] 2. A speech-to-text engine (such as Google Cloud Speech-to-Text) to convert the audio data into text.

[0571] 3. A server that analyzes text data and extracts keywords and context (using SpaCy or NLTK)

[0572] 4. A server that updates user profiles based on the extracted information and stores past conversation data (using Scikit-learn and TensorFlow).

[0573] 5. Server and device that generates and notifies optimal advice based on the user's profile and past data (using generative AI models)

[0574] Data collection

[0575] The user asks questions or makes requests to the system in natural language. For example, they might say, "I had Italian food with my friends yesterday." This voice data is recorded by the device and converted into text data.

[0576] Data analysis and profile updates

[0577] The converted text data is sent to a server and analyzed using natural language processing technology. For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted from the text data. The server then adds this information to the user's profile and stores it as past conversation data.

[0578] Generating and Serving Advice

[0579] When the user makes a new request, such as "I'm not sure what to do on my day off," the device re-records this voice data, converts it to text data, and sends it to the server. The server analyzes the request and generates appropriate advice by referencing past profile data. A generative AI model is used to create appropriate suggestions based on the user's past data. For example, a suggestion might be, "I remember you enjoyed Italian food. Why not try that newly opened restaurant?"

[0580] Examples of concrete examples and prompts

[0581] As a concrete scenario, consider the case where a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend." In this case, the following flow will occur.

[0582] The device records audio data and converts it to text using Google Cloud Speech-to-Text.

[0583] The server analyzes the received text data using SpaCy and extracts keywords such as "friend," "birthday," "present," and "worried."

[0584] The server refers to the user profile and takes into account things like "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past."

[0585] The server uses a generative AI model to generate specific suggestions such as "latest mystery novels."

[0586] The device notifies the user of the generated advice.

[0587] Users can use the suggestions to help them choose a gift.

[0588] Examples of prompts that users can enter include, "I'm not sure what to do on my day off. Could you give me some advice?" or "I'm having trouble deciding on a birthday present for my friend." Based on these prompts, the system can provide users with quick and accurate advice.

[0589] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0590] Program processing flow

[0591] Step 1:

[0592] The user speaks. Specifically, the user says to the terminal, "I ate Italian food with my friends yesterday," and voice data is generated.

[0593] Input: Audio data

[0594] Output: Audio data

[0595] Step 2:

[0596] The device converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text, for example, "I ate Italian food with my friends yesterday."

[0597] Specific operation: The device calls the voice recognition software and saves the voice data as text data.

[0598] Input: Audio data

[0599] Output: Text data

[0600] Step 3:

[0601] The terminal sends the converted text data to the server in JSON format using an HTTP request.

[0602] Specific operation: The terminal creates an HTTP request and sends text data as packets to the server.

[0603] Input: Text data

[0604] Output: HTTP request to the server

[0605] Step 4:

[0606] The server analyzes the received text data and extracts keywords and context from the text data using natural language processing tools (e.g., SpaCy, NLTK). For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted.

[0607] Specific operation: The server calls a natural language processing tool to break down the text data into keywords and context.

[0608] Input: Text data

[0609] Output: Extracted keywords and context data

[0610] Step 5:

[0611] The server updates the user's profile based on the extracted information and remembers past conversation data. Using machine learning algorithms (e.g., Scikit-learn, TensorFlow), the profile is updated and stored in a database. For example, profile information such as "I like Italian food" and "I like spending time with friends" are added.

[0612] What it does: The server uses machine learning algorithms to update the profile and save it in a database.

[0613] Input: Keywords and contextual data

[0614] Output: Updated user profile

[0615] Step 6:

[0616] The user issues a new request, for example, saying to the device, "I'm not sure what to do on my day off," and voice data is generated again.

[0617] Input: New audio data

[0618] Output: New audio data

[0619] Step 7:

[0620] The device converts this new voice data back into text data and sends it to the server. Google Cloud Speech-to-Text converts the voice data into text data, creates an HTTP request, and sends it to the server.

[0621] Specific operation: The device calls the voice recognition software, converts the voice data into text data, creates an HTTP request, and sends it to the server.

[0622] Input: New audio data

[0623] Output: Text data, HTTP request to server

[0624] Step 8:

[0625] The server analyzes the new text data and generates appropriate advice by referencing past profile data. Using a generative AI model (such as GPT-3), it generates advice based on past profile data, such as "I remember you enjoying Italian food. Why don't you try that newly opened restaurant?"

[0626] How it works: The server calls the generative AI model and generates new advice by referencing past profile data.

[0627] Input: New text data, user profile

[0628] Output: Generated advice

[0629] Step 9:

[0630] The device will notify the user of the generated advice, which will be provided to the user via text message or voice message.

[0631] Specific operation: The device displays or audibly notifies the user of the generated advice via the user interface.

[0632] Input: Generated advice

[0633] Output: User notification

[0634] (Application example 1)

[0635] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0636] In recent years, mail-order sales systems have been required to provide personalized product recommendations to users. However, conventional systems often fail to fully utilize the user's past conversations and preferences, and can only provide general recommendations. Therefore, a system that can provide accurate and individualized product recommendations to users is needed.

[0637] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0638] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for generating a proposal for a specific product based on the generated advice, and means for notifying the user of the proposed product. This makes it possible to make personalized product proposals by utilizing the user's past conversations and preferences.

[0639] "Voice data" refers to data that is a digital recording of voice information uttered by a user.

[0640] "Text data" is voice data expressed as text information.

[0641] "Keywords" are contextually significant words or phrases extracted from text data through analysis.

[0642] "Context" refers to the additional information or meaning that comes from the context in which a keyword is used.

[0643] A "profile" is a collection of personalized data that includes a user's preferences, past conversations, behavioral patterns, etc.

[0644] "Advice" refers to instructions or suggestions for actions or choices that are generated based on a user's profile and past conversation data.

[0645] "Product suggestions" are suggestions for specific products that users are encouraged to purchase, generated based on their profile and past conversation data.

[0646] "Notification" is a means of communication used to communicate system-generated advice and product suggestions to users.

[0647] This invention relates to a system that can memorize a user's past conversations and preferences and provide personalized advice and product suggestions based on them. The system includes a series of means for collecting, analyzing, memorizing, and learning from user input data to generate optimal advice and product suggestions.

[0648] System Configuration

[0649] The system mainly consists of the following elements:

[0650] 1. Speech recognition engine: Captures the user's voice data and converts it into text data. Specifically, this is done by collecting voice data using the microphone on a smartphone or smart glasses, and using voice recognition software such as the Google Speech-to-Text API.

[0651] 2. Text analysis engine: Analyzes text data and extracts keywords and context using natural language processing techniques such as Hugging Face's Transformer model.

[0652] 3. Profile Management Server: Updates the user's profile based on extracted keywords and context and remembers past conversation data.

[0653] 4. Advice Generation Engine: Generates appropriate advice and product suggestions for users based on their updated profile and past conversation data. Uses machine learning algorithms such as TensorFlow.

[0654] 5. Notification system: Provides generated advice and product suggestions to users. Software for notifying users of advice and product suggestions via smartphones or smart glasses.

[0655] Example of operation

[0656] When a user says to their smartphone, "I'm looking for a birthday present for my friend," the system operates as follows:

[0657] 1. Speech capture and text conversion:

[0658] The server converts the user's voice into text using the Google Speech-to-Text API.

[0659] 2. Text Analysis:

[0660] The text analysis engine uses Hugging Face's Transformer model to extract keywords such as "friend," "birthday," and "present" from the text data.

[0661] 3. Profile Update:

[0662] The profile management server updates the user's profile by referencing the extracted keywords and past data.

[0663] 4. Generating advice and suggestions:

[0664] An advice generation engine generates optimal birthday gift suggestions based on the updated profile.

[0665] 5. Notice:

[0666] A notification system sends generated suggestions to the user, such as "What new mystery novels or popular electronic gadgets might this friend like?"

[0667] Examples of AI-generated prompts

[0668] Prompts to suggest birthday gift ideas based on user voice data:

[0669] User dictation: "What would be a good birthday gift for my friend?"

[0670] Converts speech to text.

[0671] The converted text is then analyzed for keywords and context using a natural language processing model.

[0672] Extract important keywords such as "friends," "birthday," and "presents."

[0673] It retrieves related products from the database and notifies the user.

[0674] Example of desired result:

[0675] "How about a new mystery novel or a popular electronic gadget that this friend might like?"

[0676] This allows users to receive personalized advice based on an understanding of their preferences and past conversations, enabling them to choose the best product.

[0677] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0678] Step 1:

[0679] A user speaks to their smartphone, for example, "I'm looking for a birthday present for my friend." The smartphone's microphone captures the voice data, which is sent to the Google Speech-to-Text API. The Google Speech-to-Text API converts the voice data into text data, which becomes "I'm looking for a birthday present for my friend." The input is voice data, and the output is text data.

[0680] Step 2:

[0681] The text data is sent to the server. The server analyzes the text data using Hugging Face's Transformer model. This analysis extracts important keywords in the text, such as "friend," "birthday," and "present," as well as their context. The input is the text data, and the output is the keywords and context.

[0682] Step 3:

[0683] The server updates the user's profile based on the extracted keywords and context. It references the user's past conversation data and adds new information about "friends," "birthdays," and "gifts." This profile is stored in a database. The input is keywords and context, and the output is the updated profile.

[0684] Step 4:

[0685] When the user makes a new request, such as "Tell me what kind of gift would be good," the device receives this voice data and converts it into text data using the Google Speech-to-Text API again. The input is voice data, and the output is text data.

[0686] Step 5:

[0687] The server analyzes the new text data and again uses the Hugging Face Transformer model to understand its intent. It analyzes important keywords and context within the text and extracts intents such as "present" or "birthday." The input is the new text data, and the output is the intended keywords.

[0688] Step 6:

[0689] The server references past conversation data and the user's profile data to generate specific suggestions, such as "new mystery novels" or "popular electronic gadgets," using machine learning algorithms such as TensorFlow. The input is the updated profile and past data, and the output is specific product suggestions.

[0690] Step 7:

[0691] The server notifies the smartphone of the generated product suggestions. The smartphone screen displays a message saying, "How about a new mystery novel or a popular electronic gadget that your friend might like?" The input is the product suggestions, and the output is a notification to the user.

[0692] By going through the above processing steps, the user can receive personalized product suggestions based on past conversations and preferences.

[0693] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0694] This invention relates to a system that acquires a user's voice data and text data, analyzes that data, and provides personalized advice. In particular, by combining it with an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. The specific operating method and program processing of this system are described below.

[0695] System Configuration

[0696] The system consists of the following components:

[0697] 1. Device that collects user voice data

[0698] 2. A device with the processing power to convert voice data into text data

[0699] 3. A server that analyzes text data and extracts keywords and context

[0700] 4. A server that updates user profiles based on the extracted information and stores past conversation data.

[0701] 5. A server that uses an emotion engine to recognize user emotions and update the profile.

[0702] 6. Server and device that generates and notifies optimal advice based on the user's profile, past data, and emotional information

[0703] Program processing overview

[0704] 1. Data Collection

[0705] The user speaks to the device, "I had Italian food with my friends yesterday."

[0706] The device records this audio data and converts it into text data using a speech-to-text engine.

[0707] The converted text data becomes "I ate Italian food with my friends yesterday."

[0708] 2. Data transmission and analysis

[0709] The terminal transmits the converted text data to the server.

[0710] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[0711] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0712] 3. Emotion recognition

[0713] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[0714] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[0715] 4. Update your profile

[0716] The server updates the user's profile based on the new information and the recognized emotion information.

[0717] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[0718] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[0719] 5. Generating and Providing Advice

[0720] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0721] The terminal records this voice data, converts it into text data, and sends it to the server.

[0722] The server analyzes the request and generates the best advice using an emotion engine, taking into account the user's current emotions.

[0723] For example, if the user is in a relaxed state, the generated advice is, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0724] The device notifies the user of the generated advice in voice or text format.

[0725] Specific examples

[0726] For example, if a user asks for advice on what to buy as a birthday present for a friend, the following process will occur:

[0727] The terminal records the voice data and converts it into text data.

[0728] The server analyzes the text data and references past profile data and emotion information.

[0729] The server takes into consideration "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions such as "the latest mystery novel."

[0730] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[0731] The system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[0732] The processing flow will be explained below.

[0733] Step 1:

[0734] The user says to the device, "I had Italian food with my friends yesterday."

[0735] The device captures the user's speech as audio data.

[0736] Step 2:

[0737] The terminal converts the acquired voice data into text data using voice recognition technology.

[0738] The converted text data becomes "I ate Italian food with my friends yesterday."

[0739] Step 3:

[0740] The terminal transmits the converted text data to the server.

[0741] Step 4:

[0742] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[0743] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0744] Step 5:

[0745] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[0746] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[0747] Step 6:

[0748] The server updates the user's profile based on newly extracted keywords, context and recognized emotion information.

[0749] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[0750] Step 7:

[0751] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[0752] Step 8:

[0753] The user talks to the device, saying, "I'm wondering what to do on my day off."

[0754] The device captures the user's speech as audio data and converts it into text data.

[0755] Step 9:

[0756] The terminal transmits the converted text data to the server.

[0757] Step 10:

[0758] The server analyzes the received text data to understand the intent and emotional information of the request.

[0759] For example, the keywords "holiday" and "not sure what to do" and the user's relaxed state are analyzed.

[0760] Step 11:

[0761] The server references past conversation data and profiles to generate optimal advice based on the user's preferences and past behavior.

[0762] For example, if the user is in a relaxed state, the system generates advice such as, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0763] Step 12:

[0764] The server transmits the generated advice to the terminal.

[0765] Step 13:

[0766] The device will notify the user of the received advice in voice or text format.

[0767] This series of steps allows users to receive faster, more personalized advice that adapts to their current emotional state and past behavior.

[0768] Example 2

[0769] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0770] Conventional user assistance systems simply convert user voice data into text data and extract context and keywords, making it difficult to provide personalized advice that takes into account the user's emotional state. Therefore, there is a need for a system that can generate and provide more appropriate advice based on the user's emotions, thereby improving the user experience.

[0771] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0772] In this invention, the server includes means for acquiring a user's voice data and converting it into text data, means for analyzing the text data and extracting keywords and context, and means for analyzing the user's emotional information using an emotion engine and updating the profile, thereby enabling the server to generate and provide appropriate advice and communication that takes into account the user's emotional state.

[0773] "Audio data from a user" refers to data that is a recording of what a user says.

[0774] "Text data" refers to character information obtained by analyzing voice data.

[0775] A "keyword" is a word or phrase that indicates important information from text data.

[0776] "Context" refers to information that indicates the relationship between keywords and the flow of the entire sentence.

[0777] "Profile" means data that records the characteristics and preferences of a user.

[0778] An "emotion engine" is software for analyzing emotional states from voice and text.

[0779] "Content data" refers to data relating to conversations and actions with users that have been recorded in the past.

[0780] "Advice" means any recommendation or instruction provided to a User.

[0781] "Communication" refers to activities that involve dialogue and exchange of information with users.

[0782] A "request" is a question or request made by a user to the system.

[0783] "Intent" refers to the goal or desire the user wants to achieve through their request.

[0784] "Past behavior" refers to the activities and behavioral history that a user has undertaken to date.

[0785] "Values" refer to the fundamental beliefs and ways of thinking that govern a user's thoughts and actions.

[0786] The present invention relates to a system that acquires a user's voice data and text data, analyzes them, and provides personalized advice. In particular, by combining an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. This system is implemented using the following hardware and software.

[0787] System configuration and hardware / software used

[0788] Hardware

[0789] 1. Device: A digital device with a microphone (e.g., smart speaker, smartphone) that captures user voice data.

[0790] 2. Server: A cloud or on-premise server for managing data analysis and profiles.

[0791] software

[0792] 1. Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[0793] 2. Natural Language Processing (NLP) engine: SpaCy is used to perform contextual analysis of text data and extract keywords.

[0794] 3. Emotion Recognition Engine: Analyzes user emotional information using IBM Watson Tone Analyzer.

[0795] 4. Machine learning algorithm: TensorFlow is used to continuously update and learn user profiles and emotion data.

[0796] 5. Advice Generation Engine: Uses OpenAI's GPT-4 model to generate optimal advice.

[0797] 6. Web Framework: Use Django Web Framework for profile management and database operations.

[0798] Specific system operation examples

[0799] Data collection

[0800] The user speaks to the device, "I had Italian food with my friends yesterday."

[0801] The device records this voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data will be "I ate Italian food with my friends yesterday."

[0802] Data analysis and profile updates

[0803] The device sends the converted text data to the server using the REST API.

[0804] The server receives the text data and uses SpaCy to perform grammatical analysis and extract keywords. The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0805] The server then uses IBM Watson Tone Analyzer to analyze emotional information from the voice and text data and recognize emotions such as "happy" or "relaxed."

[0806] The server updates the user's profile with the new information and sentiment, adding traits such as "enjoys Italian food" and "is relaxed when with friends" to the database using the Django Web Framework.

[0807] Advice generation and delivery

[0808] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0809] The device records this voice data, converts it into text data using the Google Cloud Speech-to-Text API, and sends it to the server.

[0810] The server analyzes the request, takes into account the user's current emotions using an emotion engine, and then generates the optimal advice using OpenAI's GPT-4 model. For example, if the user is in a relaxed state, the advice generated might be, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0811] The device notifies the user of the generated advice in voice or text format.

[0812] Specific examples

[0813] As a specific example, consider the case where a user asks for advice, saying, "I'm having trouble deciding on a birthday present for my friend."

[0814] The device records audio data and converts it into text data using the Google Cloud Speech-to-Text API.

[0815] The server analyzes the text data and references past profile data and emotional information using IBM Watson Tone Analyzer.

[0816] Next, the server uses OpenAI's GPT-4 model to consider "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions, such as "the latest mystery novel."

[0817] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[0818] In this way, the system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[0819] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0820] Step 1:

[0821] The user speaks to the device, "I had Italian food with my friends yesterday."

[0822] Input: User's voice data.

[0823] To record this audio data, the device uses a microphone to capture the audio.

[0824] Output: Recorded audio data file.

[0825] Step 2:

[0826] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts it into text data.

[0827] Input: Audio data file.

[0828] The Google Cloud Speech-to-Text API uses speech recognition technology to analyze audio data and generate corresponding text data.

[0829] Output: Text data: "I ate Italian food with my friends yesterday."

[0830] Step 3:

[0831] The device sends the converted text data to the server using the REST API.

[0832] Input: Text data.

[0833] The device packages the text data in JSON format and sends it to the server using an HTTP request.

[0834] Output: The text data received on the server.

[0835] Step 4:

[0836] The server analyzes the received text data using SpaCy, performing grammatical analysis and keyword extraction.

[0837] Input: Text data.

[0838] The SpaCy engine analyzes text data and extracts keywords such as "yesterday," "friends," "Italian food," and "ate." It also performs contextual analysis to understand the relationships between keywords.

[0839] Output: Extracted keywords and context information.

[0840] Step 5:

[0841] The server uses IBM Watson Tone Analyzer to analyze emotional information from the extracted text data.

[0842] Input: Keywords and contextual information.

[0843] IBM Watson Tone Analyzer uses emotion recognition algorithms to identify emotions such as "happy" or "relaxed" from the tone and content of text.

[0844] Output: Recognized emotion information.

[0845] Step 6:

[0846] The server uses the Django Web Framework to update the user's profile based on the new information and the recognized emotion information.

[0847] Input: Recognized emotion information and extracted keywords.

[0848] The profile database stores the new information and adds characteristics to the profile, such as "I enjoy Italian food" and "I feel relaxed when I'm with friends."

[0849] Output: Updated user profile.

[0850] Step 7:

[0851] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0852] Input: The user's new voice data.

[0853] The terminal re-records this voice data, converts it into text data, and sends it to the server.

[0854] Output: The new text data received on the server.

[0855] Step 8:

[0856] The server parses the request, again using an emotion engine to take into account the user's current emotions, and generates the best advice using OpenAI's GPT-4 model.

[0857] Input: New text data and existing user profile.

[0858] OpenAI's GPT-4 model combines past profile information with your current emotional state to generate a recommendation like, "Why not try a new restaurant to rediscover that Italian food you enjoyed before?"

[0859] Output: The generated advice.

[0860] Step 9:

[0861] The terminal provides the generated advice to the user.

[0862] Input: The generated advice.

[0863] Using a text-to-speech engine, the device will verbally inform the user, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[0864] Output: The user receives the advice.

[0865] (Application example 2)

[0866] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0867] Modern online shopping sites lack personalized product recommendations that utilize users' voice data and emotional information. In particular, there is a need for a system that can improve the user experience by providing optimal advice and products based on the user's emotional state. The lack of such personalized services is causing users to lose their motivation to purchase and making it difficult to select the right product.

[0868] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data from the user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and product recommendations for the user based on the updated profile, past conversation data, and emotional information, and means for providing the generated advice and product recommendations to the user. This enables personalized product recommendations that take into account the user's emotional state and past behavior.

[0869] "User's voice data" refers to voice information uttered by a user through a voice input device, and is data including voice signals such as conversations and instructions.

[0870] "Text data" is voice data converted into character information, and is data consisting of a sequence of characters.

[0871] "Keywords" are important words or phrases extracted from text data that are meaningful for analyzing context.

[0872] "Context" refers to the background and situation of sentences and phrases present in text data, and is information that helps understand the meaning.

[0873] A "profile" is a user information system constructed by combining a user's attribute information, behavioral history, preferences, emotional state, etc.

[0874] "Emotional information" is data that indicates emotional states such as joy, sadness, anger, and stress, analyzed from a user's statements and actions.

[0875] "Product recommendation" refers to the act and content of suggesting specific products or services to users based on their profile and emotional information.

[0876] "Appropriate advice" means providing useful information and guidance tailored to the user's needs and circumstances.

[0877] "Means for acquiring voice data" refers to a device or method for collecting voice information from a user using a microphone or the like.

[0878] The "means for converting voice data into text data" refers to a device or method for converting voice data into text information using voice recognition technology.

[0879] The "means for analyzing converted text data" refers to a device or method that uses natural language processing technology to analyze text data and extract context and meaning.

[0880] The "means for storing past conversation data" refers to a device or method for storing the history of dialogue with the user in a database or the like.

[0881] The "means for generating advice and product recommendations" is a device or method for generating appropriate advice and product recommendations based on the profile and emotion information.

[0882] A "means for providing generated advice and product recommendations" is a device or method that notifies a user of advice or product recommendations in audio or text form.

[0883] This invention is a system that acquires a user's voice data, converts it into text data, and provides personalized advice and product recommendations that incorporate emotional information. This system improves the user experience and helps them make more appropriate product selections.

[0884] The general flow of the system is as follows: First, the user inputs voice data using a device such as a smartphone. The voice data is acquired in real time and converted into text data using the Google Cloud Speech-to-Text API. This text data is then sent to the server.

[0885] After receiving the converted text data, the server analyzes it using the Google Cloud Natural Language API to extract keywords and context. It then uses IBM Watson Tone Analyzer to analyze emotional information from the text data. The server then updates the user's profile based on the emotional information and the extracted keywords and context.

[0886] When updating profiles, machine learning algorithms such as TensorFlow are used to sequentially learn and update the user's past conversation data and behavioral history, resulting in personalized information that reflects the user's preferences and emotional state in real time.

[0887] Next, when the user makes a new request, the system analyzes the request and references past conversation data and profiles to generate optimal advice and product recommendations, which are then sent to the user via a device such as a smartphone.

[0888] As a concrete example, if a user says, "I've been feeling stressed lately, so I want something that will help me relax," the following processing will occur.

[0889] First, the smartphone's microphone is used to capture voice data, which is then converted into text using the Google Cloud Speech-to-Text API. The converted text data is then sent to a server, where it is analyzed using the Google Natural Language API to extract keywords such as "stress" and "relaxation." Next, IBM Watson Tone Analyzer is used to analyze the emotional data and identify high stress levels.

[0890] Based on this information, the server uses machine learning algorithms such as TensorFlow to update the profile and generate appropriate product recommendations. For example, recommendations such as "Lavender aroma diffuser" or "Soothing music CD" are generated. The generated advice and product recommendations are then displayed on the user's smartphone as a notification saying, "How about a lavender aroma diffuser to relieve stress?"

[0891] The following is a specific example of a prompt sentence to be input to the generative AI model:

[0892] Prompt: "The user says, 'I'm tired and looking for something to relax.' Speech is converted to text and the keywords 'tired' and 'relaxation' are extracted. Sentiment analysis identifies this as a state of high stress. Based on the user's profile, it is determined that they need something to relax, and the system recommends the purchase of a 'lavender aroma diffuser,' which has a relaxing effect."

[0893] The system enables personalized product recommendations that take into account the user's emotional state and past behavior, improving the user experience.

[0894] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0895] Step 1:

[0896] The user speaks into the smartphone's microphone, saying, "I've been feeling stressed lately, so I want something that will help me relax." The input voice data is picked up by the smartphone's microphone. The output is the original voice data.

[0897] Step 2:

[0898] The device sends the acquired voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data and outputs the converted text data.

[0899] Step 3:

[0900] The terminal transmits the converted text data to the server, and the server receives the text data.

[0901] Step 4:

[0902] The server uses the Google Cloud Natural Language API to parse the received text data and extract keywords and context. The input is the received text data, and the output is the extracted keywords (e.g., "stress," "relax") and context information.

[0903] Step 5:

[0904] The server uses IBM Watson Tone Analyzer to analyze emotional information from text data. The input is text data, and the output is emotional data (e.g., high stress).

[0905] Step 6:

[0906] The server updates the user's profile based on the extracted keywords and emotion data. The profile is updated sequentially using machine learning algorithms such as TensorFlow based on the user's past conversation data and behavioral history. The input is keywords, context, emotion data, and past profile data, and the output is the updated profile.

[0907] Step 7:

[0908] When a user inputs a new request, the server analyzes the request. For example, if a user inputs "I want a relaxation item," the input voice data is converted into text data and analyzed. The input is the user's new request, and the output is the analysis result.

[0909] Step 8:

[0910] The server refers to the past conversation data and the updated profile to generate optimal advice and product recommendations. For example, a "lavender aroma diffuser" is recommended. The input is the parsed request, the past conversation data, and the updated profile, and the output is the generated advice and product recommendations.

[0911] Step 9:

[0912] The terminal notifies the user of the generated advice and product recommendation. For example, a notification such as "How about a lavender aroma diffuser to relieve stress?" is displayed. The input is the generated advice and product recommendation, and the output is a notification to the user.

[0913] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0914] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0915] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0916] [Third embodiment]

[0917] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0918] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0919] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0920] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0921] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0922] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0923] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0924] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0925] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0926] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0927] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0928] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0929] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes a series of means for collecting user input data, analyzing it, memorizing it, learning from it, and generating optimal advice. The specific operation of this system is shown below.

[0930] System Configuration

[0931] The system mainly consists of the following elements:

[0932] 1. Devices that collect user voice data

[0933] 2. A device with the processing power to convert voice data into text data

[0934] 3. A server that analyzes text data and extracts keywords and context

[0935] 4. A server that updates the user's profile based on the extracted information and stores past conversation data.

[0936] 5. Server and terminal that generates and notifies optimal advice based on the user's profile and past data

[0937] Program processing overview

[0938] 1. Data Collection

[0939] The user speaks to the device, "I had Italian food with my friends yesterday."

[0940] The device records this voice data and converts it into text data using a voice-to-text conversion engine.

[0941] The converted text data becomes "I ate Italian food with my friends yesterday."

[0942] 2. Data transmission and analysis

[0943] The terminal transmits the converted text data to the server.

[0944] The server receives the text data and analyzes it using natural language processing (NLP) techniques.

[0945] Keywords extracted through analysis include "yesterday," "friends," "Italian food," and "ate," and the context of each is also taken into consideration.

[0946] 3. Update your profile

[0947] The server adds new information to the user's profile, such as "I have a penchant for Italian food" or "I like spending time with friends."

[0948] The server uses machine learning algorithms to continuously update and learn from the user's profile.

[0949] 4. Generating and Providing Advice

[0950] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[0951] The device records this voice data, converts it into text, and sends it to the server.

[0952] The server analyzes the request and uses historical profile data to generate appropriate advice.

[0953] For example, based on past data, suggestions might be generated such as, "I remember you enjoying Italian food. Why not try that newly opened restaurant?"

[0954] The device notifies the user of the generated advice.

[0955] Specific examples

[0956] Specifically, when a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend," the following process occurs.

[0957] The device records the voice data and converts it into text.

[0958] The server analyzes the text data and references past profile data.

[0959] The server takes into account factors such as "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past" and generates specific suggestions such as "the latest mystery novel."

[0960] The device provides the generated advice to the user, who can refer to it.

[0961] The system gives users a personalized AI partner that understands their preferences and past conversations, allowing them to receive faster, more relevant advice.

[0962] The processing flow will be explained below.

[0963] Step 1:

[0964] The user says to the device, "I had Italian food with my friends yesterday."

[0965] The device captures the user's speech as audio data.

[0966] Step 2:

[0967] The terminal converts the acquired voice data into text data using voice recognition technology.

[0968] The converted text data becomes "I ate Italian food with my friends yesterday."

[0969] Step 3:

[0970] The terminal transmits the converted text data to the server.

[0971] Step 4:

[0972] The server receives the text data and uses natural language processing (NLP) techniques to perform grammatical analysis and keyword extraction.

[0973] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[0974] Step 5:

[0975] The server references past conversation data and the user's profile, and updates the profile based on newly extracted keywords and context.

[0976] For example, a user can add information to their profile such as "I like Italian food" or "I like spending time with friends."

[0977] Step 6:

[0978] The server uses machine learning models to learn user preferences and behavioral patterns and update the profile.

[0979] Step 7:

[0980] The user talks to the device, saying, "I'm wondering what to do on my day off."

[0981] The device acquires the user's speech as audio data and converts it into text data.

[0982] Step 8:

[0983] The terminal transmits the converted text data to the server.

[0984] Step 9:

[0985] The server receives the text data and analyzes the request using NLP techniques.

[0986] Keywords such as "holiday" and "not sure what to do" are extracted from the analysis results.

[0987] Step 10:

[0988] The server references past profile data and generates optimal advice based on the user's preferences and past behavior.

[0989] For example, a suggestion might be generated: "I remember enjoying Italian food. Why not try that newly opened restaurant?"

[0990] Step 11:

[0991] The server transmits the generated advice to the terminal.

[0992] Step 12:

[0993] The device will notify the user of the received advice in voice or text format.

[0994] This series of steps allows users to receive personalized advice, resulting in a more specific and reliable conversational experience.

[0995] Example 1

[0996] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0997] It is difficult for current AI systems to effectively memorize and analyze a user's past behavior and preferences and provide personalized advice based on the results. Conventional technologies lack sufficient user profile updates and past data reference for generating advice, resulting in a lack of improvement in the user experience. Therefore, there is a need for a system that can easily build a profile from a user's voice data and quickly provide appropriate advice.

[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0999] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user using a generative AI model based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for receiving a new request from the user and analyzing the request to understand its intent, means for generating optimal advice by referring to past conversation data and the user's profile, means for generating specific suggestions based on the user's preferences and values, means for checking whether the suggested content matches the user's past actions and conversations, and means for providing the suggestion to the user if a match is found. This makes it possible to build a profile from the user's voice data and provide quick and accurate advice.

[1000] "Voice data" refers to digital data that records the voice uttered by the user.

[1001] "Text data" is digital data that has been converted from voice data into character information.

[1002] "Keywords" are important words or phrases extracted from text data.

[1003] "Context" refers to the semantic background, including the context of the keyword and the environment in which it is used.

[1004] A "user profile" is an individual data set that includes a user's preferences, past behavior, and conversational data.

[1005] "Generative AI models" are artificial intelligence algorithms and models used to analyze data and generate advice.

[1006] An "HTTP request" is a request of a communication protocol used to send data from a terminal to a server.

[1007] "Natural language processing" is a general term for technology that analyzes human language and extracts meaning.

[1008] A "machine learning algorithm" is a computational method for learning from data and making predictions or classifications.

[1009] "Advice" is a suggestion or advice provided by the system based on the user's profile and past behavior.

[1010] A "request" is a question or input instruction that a user makes to the system.

[1011] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes various means for collecting user input data, analyzing, memorizing, learning from it, and generating optimal advice. The specific operation of this system is described below.

[1012] System Configuration

[1013] The system mainly consists of the following elements:

[1014] 1. Devices that collect user voice data (smartphones, tablets, etc.)

[1015] 2. A speech-to-text engine (such as Google Cloud Speech-to-Text) to convert the audio data into text.

[1016] 3. A server that analyzes text data and extracts keywords and context (using SpaCy or NLTK)

[1017] 4. A server that updates user profiles based on the extracted information and stores past conversation data (using Scikit-learn and TensorFlow).

[1018] 5. Server and device that generates and notifies optimal advice based on the user's profile and past data (using generative AI models)

[1019] Data collection

[1020] The user asks questions or makes requests to the system in natural language. For example, they might say, "I had Italian food with my friends yesterday." This voice data is recorded by the device and converted into text data.

[1021] Data analysis and profile updates

[1022] The converted text data is sent to a server and analyzed using natural language processing technology. For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted from the text data. The server then adds this information to the user's profile and stores it as past conversation data.

[1023] Generating and Serving Advice

[1024] When the user makes a new request, such as "I'm not sure what to do on my day off," the device re-records this voice data, converts it to text data, and sends it to the server. The server analyzes the request and generates appropriate advice by referencing past profile data. A generative AI model is used to create appropriate suggestions based on the user's past data. For example, a suggestion might be, "I remember you enjoyed Italian food. Why not try that newly opened restaurant?"

[1025] Examples of concrete examples and prompts

[1026] As a concrete scenario, consider the case where a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend." In this case, the following flow will occur.

[1027] The device records audio data and converts it to text using Google Cloud Speech-to-Text.

[1028] The server analyzes the received text data using SpaCy and extracts keywords such as "friend," "birthday," "present," and "worried."

[1029] The server refers to the user profile and takes into account things like "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past."

[1030] The server uses a generative AI model to generate specific suggestions such as "latest mystery novels."

[1031] The device notifies the user of the generated advice.

[1032] Users can use the suggestions to help them choose a gift.

[1033] Examples of prompts that users can enter include, "I'm not sure what to do on my day off. Could you give me some advice?" or "I'm having trouble deciding on a birthday present for my friend." Based on these prompts, the system can provide users with quick and accurate advice.

[1034] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1035] Program processing flow

[1036] Step 1:

[1037] The user speaks. Specifically, the user says to the terminal, "I ate Italian food with my friends yesterday," and voice data is generated.

[1038] Input: Audio data

[1039] Output: Audio data

[1040] Step 2:

[1041] The device converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text, for example, "I ate Italian food with my friends yesterday."

[1042] Specific operation: The device calls the voice recognition software and saves the voice data as text data.

[1043] Input: Audio data

[1044] Output: Text data

[1045] Step 3:

[1046] The terminal sends the converted text data to the server in JSON format using an HTTP request.

[1047] Specific operation: The terminal creates an HTTP request and sends text data as packets to the server.

[1048] Input: Text data

[1049] Output: HTTP request to the server

[1050] Step 4:

[1051] The server analyzes the received text data and extracts keywords and context from the text data using natural language processing tools (e.g., SpaCy, NLTK). For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted.

[1052] Specific operation: The server calls a natural language processing tool to break down the text data into keywords and context.

[1053] Input: Text data

[1054] Output: Extracted keywords and context data

[1055] Step 5:

[1056] The server updates the user's profile based on the extracted information and remembers past conversation data. Using machine learning algorithms (e.g., Scikit-learn, TensorFlow), the profile is updated and stored in a database. For example, profile information such as "I like Italian food" and "I like spending time with friends" are added.

[1057] What it does: The server uses machine learning algorithms to update the profile and save it in a database.

[1058] Input: Keywords and contextual data

[1059] Output: Updated user profile

[1060] Step 6:

[1061] The user issues a new request, for example, saying to the device, "I'm not sure what to do on my day off," and voice data is generated again.

[1062] Input: New audio data

[1063] Output: New audio data

[1064] Step 7:

[1065] The device converts this new voice data back into text data and sends it to the server. Google Cloud Speech-to-Text converts the voice data into text data, creates an HTTP request, and sends it to the server.

[1066] Specific operation: The device calls the voice recognition software, converts the voice data into text data, creates an HTTP request, and sends it to the server.

[1067] Input: New audio data

[1068] Output: Text data, HTTP request to server

[1069] Step 8:

[1070] The server analyzes the new text data and generates appropriate advice by referencing past profile data. Using a generative AI model (such as GPT-3), it generates advice based on past profile data, such as "I remember you enjoying Italian food. Why don't you try that newly opened restaurant?"

[1071] How it works: The server calls the generative AI model and generates new advice by referencing past profile data.

[1072] Input: New text data, user profile

[1073] Output: Generated advice

[1074] Step 9:

[1075] The device will notify the user of the generated advice, which will be provided to the user via text message or voice message.

[1076] Specific operation: The device displays or audibly notifies the user of the generated advice via the user interface.

[1077] Input: Generated advice

[1078] Output: User notification

[1079] (Application example 1)

[1080] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1081] In recent years, mail-order sales systems have been required to provide personalized product recommendations to users. However, conventional systems often fail to fully utilize the user's past conversations and preferences, and can only provide general recommendations. Therefore, a system that can provide accurate and individualized product recommendations to users is needed.

[1082] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1083] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for generating a proposal for a specific product based on the generated advice, and means for notifying the user of the proposed product. This makes it possible to make personalized product proposals by utilizing the user's past conversations and preferences.

[1084] "Voice data" refers to data that is a digital recording of voice information uttered by a user.

[1085] "Text data" is voice data expressed as text information.

[1086] "Keywords" are contextually significant words or phrases extracted from text data through analysis.

[1087] "Context" refers to the additional information or meaning that comes from the context in which a keyword is used.

[1088] A "profile" is a collection of personalized data that includes a user's preferences, past conversations, behavioral patterns, etc.

[1089] "Advice" refers to instructions or suggestions for actions or choices that are generated based on a user's profile and past conversation data.

[1090] "Product suggestions" are suggestions for specific products that users are encouraged to purchase, generated based on their profile and past conversation data.

[1091] "Notification" is a means of communication used to communicate system-generated advice and product suggestions to users.

[1092] This invention relates to a system that can memorize a user's past conversations and preferences and provide personalized advice and product suggestions based on them. The system includes a series of means for collecting, analyzing, memorizing, and learning from user input data to generate optimal advice and product suggestions.

[1093] System Configuration

[1094] The system mainly consists of the following elements:

[1095] 1. Speech recognition engine: Captures the user's voice data and converts it into text data. Specifically, this is done by collecting voice data using the microphone on a smartphone or smart glasses, and using voice recognition software such as the Google Speech-to-Text API.

[1096] 2. Text analysis engine: Analyzes text data and extracts keywords and context using natural language processing techniques such as Hugging Face's Transformer model.

[1097] 3. Profile Management Server: Updates the user's profile based on extracted keywords and context and remembers past conversation data.

[1098] 4. Advice Generation Engine: Generates appropriate advice and product suggestions for users based on their updated profile and past conversation data. Uses machine learning algorithms such as TensorFlow.

[1099] 5. Notification system: Provides generated advice and product suggestions to users. Software for notifying users of advice and product suggestions via smartphones or smart glasses.

[1100] Example of operation

[1101] When a user says to their smartphone, "I'm looking for a birthday present for my friend," the system operates as follows:

[1102] 1. Speech capture and text conversion:

[1103] The server converts the user's voice into text using the Google Speech-to-Text API.

[1104] 2. Text Analysis:

[1105] The text analysis engine uses Hugging Face's Transformer model to extract keywords such as "friend," "birthday," and "present" from the text data.

[1106] 3. Profile Update:

[1107] The profile management server updates the user's profile by referencing the extracted keywords and past data.

[1108] 4. Generating advice and suggestions:

[1109] An advice generation engine generates optimal birthday gift suggestions based on the updated profile.

[1110] 5. Notice:

[1111] A notification system sends generated suggestions to the user, such as "What new mystery novels or popular electronic gadgets might this friend like?"

[1112] Examples of AI-generated prompts

[1113] Prompts to suggest birthday gift ideas based on user voice data:

[1114] User dictation: "What would be a good birthday gift for my friend?"

[1115] Converts speech to text.

[1116] The converted text is then analyzed for keywords and context using a natural language processing model.

[1117] Extract important keywords such as "friends," "birthday," and "presents."

[1118] It retrieves related products from the database and notifies the user.

[1119] Example of desired result:

[1120] "How about a new mystery novel or a popular electronic gadget that this friend might like?"

[1121] This allows users to receive personalized advice based on an understanding of their preferences and past conversations, enabling them to choose the best product.

[1122] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1123] Step 1:

[1124] A user speaks to their smartphone, for example, "I'm looking for a birthday present for my friend." The smartphone's microphone captures the voice data, which is sent to the Google Speech-to-Text API. The Google Speech-to-Text API converts the voice data into text data, which becomes "I'm looking for a birthday present for my friend." The input is voice data, and the output is text data.

[1125] Step 2:

[1126] The text data is sent to the server. The server analyzes the text data using Hugging Face's Transformer model. This analysis extracts important keywords in the text, such as "friend," "birthday," and "present," as well as their context. The input is the text data, and the output is the keywords and context.

[1127] Step 3:

[1128] The server updates the user's profile based on the extracted keywords and context. It references the user's past conversation data and adds new information about "friends," "birthdays," and "gifts." This profile is stored in a database. The input is keywords and context, and the output is the updated profile.

[1129] Step 4:

[1130] When the user makes a new request, such as "Tell me what kind of gift would be good," the device receives this voice data and converts it into text data using the Google Speech-to-Text API again. The input is voice data, and the output is text data.

[1131] Step 5:

[1132] The server analyzes the new text data and again uses the Hugging Face Transformer model to understand its intent. It analyzes important keywords and context within the text and extracts intents such as "present" or "birthday." The input is the new text data, and the output is the intended keywords.

[1133] Step 6:

[1134] The server references past conversation data and the user's profile data to generate specific suggestions, such as "new mystery novels" or "popular electronic gadgets," using machine learning algorithms such as TensorFlow. The input is the updated profile and past data, and the output is specific product suggestions.

[1135] Step 7:

[1136] The server notifies the smartphone of the generated product suggestions. The smartphone screen displays a message saying, "How about a new mystery novel or a popular electronic gadget that your friend might like?" The input is the product suggestions, and the output is a notification to the user.

[1137] By going through the above processing steps, the user can receive personalized product suggestions based on past conversations and preferences.

[1138] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1139] This invention relates to a system that acquires a user's voice data and text data, analyzes that data, and provides personalized advice. In particular, by combining it with an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. The specific operating method and program processing of this system are described below.

[1140] System Configuration

[1141] The system consists of the following components:

[1142] 1. Device that collects user voice data

[1143] 2. A device with the processing power to convert voice data into text data

[1144] 3. A server that analyzes text data and extracts keywords and context

[1145] 4. A server that updates user profiles based on the extracted information and stores past conversation data.

[1146] 5. A server that uses an emotion engine to recognize user emotions and update the profile.

[1147] 6. Server and device that generates and notifies optimal advice based on the user's profile, past data, and emotional information

[1148] Program processing overview

[1149] 1. Data Collection

[1150] The user speaks to the device, "I had Italian food with my friends yesterday."

[1151] The device records this audio data and converts it into text data using a speech-to-text engine.

[1152] The converted text data becomes "I ate Italian food with my friends yesterday."

[1153] 2. Data transmission and analysis

[1154] The terminal transmits the converted text data to the server.

[1155] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[1156] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1157] 3. Emotion recognition

[1158] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[1159] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[1160] 4. Update your profile

[1161] The server updates the user's profile based on the new information and the recognized emotion information.

[1162] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[1163] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[1164] 5. Generating and Providing Advice

[1165] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1166] The terminal records this voice data, converts it into text data, and sends it to the server.

[1167] The server analyzes the request and generates the best advice using an emotion engine, taking into account the user's current emotions.

[1168] For example, if the user is in a relaxed state, the generated advice is, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1169] The device notifies the user of the generated advice in voice or text format.

[1170] Specific examples

[1171] For example, if a user asks for advice on what to buy as a birthday present for a friend, the following process will occur:

[1172] The terminal records the voice data and converts it into text data.

[1173] The server analyzes the text data and references past profile data and emotion information.

[1174] The server takes into consideration "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions such as "the latest mystery novel."

[1175] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[1176] The system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[1177] The processing flow will be explained below.

[1178] Step 1:

[1179] The user says to the device, "I had Italian food with my friends yesterday."

[1180] The device captures the user's speech as audio data.

[1181] Step 2:

[1182] The terminal converts the acquired voice data into text data using voice recognition technology.

[1183] The converted text data becomes "I ate Italian food with my friends yesterday."

[1184] Step 3:

[1185] The terminal transmits the converted text data to the server.

[1186] Step 4:

[1187] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[1188] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1189] Step 5:

[1190] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[1191] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[1192] Step 6:

[1193] The server updates the user's profile based on newly extracted keywords, context and recognized emotion information.

[1194] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[1195] Step 7:

[1196] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[1197] Step 8:

[1198] The user talks to the device, saying, "I'm wondering what to do on my day off."

[1199] The device captures the user's speech as audio data and converts it into text data.

[1200] Step 9:

[1201] The terminal transmits the converted text data to the server.

[1202] Step 10:

[1203] The server analyzes the received text data to understand the intent and emotional information of the request.

[1204] For example, the keywords "holiday" and "not sure what to do" and the user's relaxed state are analyzed.

[1205] Step 11:

[1206] The server references past conversation data and profiles to generate optimal advice based on the user's preferences and past behavior.

[1207] For example, if the user is in a relaxed state, the system generates advice such as, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1208] Step 12:

[1209] The server transmits the generated advice to the terminal.

[1210] Step 13:

[1211] The device will notify the user of the received advice in voice or text format.

[1212] This series of steps allows users to receive faster, more personalized advice that adapts to their current emotional state and past behavior.

[1213] Example 2

[1214] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1215] Conventional user assistance systems simply convert user voice data into text data and extract context and keywords, making it difficult to provide personalized advice that takes into account the user's emotional state. Therefore, there is a need for a system that can generate and provide more appropriate advice based on the user's emotions, thereby improving the user experience.

[1216] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1217] In this invention, the server includes means for acquiring a user's voice data and converting it into text data, means for analyzing the text data and extracting keywords and context, and means for analyzing the user's emotional information using an emotion engine and updating the profile, thereby enabling the server to generate and provide appropriate advice and communication that takes into account the user's emotional state.

[1218] "Audio data from a user" refers to data that is a recording of what a user says.

[1219] "Text data" refers to character information obtained by analyzing voice data.

[1220] A "keyword" is a word or phrase that indicates important information from text data.

[1221] "Context" refers to information that indicates the relationship between keywords and the flow of the entire sentence.

[1222] "Profile" means data that records the characteristics and preferences of a user.

[1223] An "emotion engine" is software for analyzing emotional states from voice and text.

[1224] "Content data" refers to data relating to conversations and actions with users that have been recorded in the past.

[1225] "Advice" means any recommendation or instruction provided to a User.

[1226] "Communication" refers to activities that involve dialogue and exchange of information with users.

[1227] A "request" is a question or request made by a user to the system.

[1228] "Intent" refers to the goal or desire the user wants to achieve through their request.

[1229] "Past behavior" refers to the activities and behavioral history that a user has undertaken to date.

[1230] "Values" refer to the fundamental beliefs and ways of thinking that govern a user's thoughts and actions.

[1231] The present invention relates to a system that acquires a user's voice data and text data, analyzes them, and provides personalized advice. In particular, by combining an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. This system is implemented using the following hardware and software.

[1232] System configuration and hardware / software used

[1233] Hardware

[1234] 1. Device: A digital device with a microphone (e.g., smart speaker, smartphone) that captures user voice data.

[1235] 2. Server: A cloud or on-premise server for managing data analysis and profiles.

[1236] software

[1237] 1. Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[1238] 2. Natural Language Processing (NLP) engine: SpaCy is used to perform contextual analysis of text data and extract keywords.

[1239] 3. Emotion Recognition Engine: Analyzes user emotional information using IBM Watson Tone Analyzer.

[1240] 4. Machine learning algorithm: TensorFlow is used to continuously update and learn user profiles and emotion data.

[1241] 5. Advice Generation Engine: Uses OpenAI's GPT-4 model to generate optimal advice.

[1242] 6. Web Framework: Use Django Web Framework for profile management and database operations.

[1243] Specific system operation examples

[1244] Data collection

[1245] The user speaks to the device, "I had Italian food with my friends yesterday."

[1246] The device records this voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data will be "I ate Italian food with my friends yesterday."

[1247] Data analysis and profile updates

[1248] The device sends the converted text data to the server using the REST API.

[1249] The server receives the text data and uses SpaCy to perform grammatical analysis and extract keywords. The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1250] The server then uses IBM Watson Tone Analyzer to analyze emotional information from the voice and text data and recognize emotions such as "happy" or "relaxed."

[1251] The server updates the user's profile with the new information and sentiment, adding traits such as "enjoys Italian food" and "is relaxed when with friends" to the database using the Django Web Framework.

[1252] Advice generation and delivery

[1253] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1254] The device records this voice data, converts it into text data using the Google Cloud Speech-to-Text API, and sends it to the server.

[1255] The server analyzes the request, takes into account the user's current emotions using an emotion engine, and then generates the optimal advice using OpenAI's GPT-4 model. For example, if the user is in a relaxed state, the advice generated might be, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1256] The device notifies the user of the generated advice in voice or text format.

[1257] Specific examples

[1258] As a specific example, consider the case where a user asks for advice, saying, "I'm having trouble deciding on a birthday present for my friend."

[1259] The device records audio data and converts it into text data using the Google Cloud Speech-to-Text API.

[1260] The server analyzes the text data and references past profile data and emotional information using IBM Watson Tone Analyzer.

[1261] Next, the server uses OpenAI's GPT-4 model to consider "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions, such as "the latest mystery novel."

[1262] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[1263] In this way, the system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[1264] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1265] Step 1:

[1266] The user speaks to the device, "I had Italian food with my friends yesterday."

[1267] Input: User's voice data.

[1268] To record this audio data, the device uses a microphone to capture the audio.

[1269] Output: Recorded audio data file.

[1270] Step 2:

[1271] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts it into text data.

[1272] Input: Audio data file.

[1273] The Google Cloud Speech-to-Text API uses speech recognition technology to analyze audio data and generate corresponding text data.

[1274] Output: Text data: "I ate Italian food with my friends yesterday."

[1275] Step 3:

[1276] The device sends the converted text data to the server using the REST API.

[1277] Input: Text data.

[1278] The device packages the text data in JSON format and sends it to the server using an HTTP request.

[1279] Output: The text data received on the server.

[1280] Step 4:

[1281] The server analyzes the received text data using SpaCy, performing grammatical analysis and keyword extraction.

[1282] Input: Text data.

[1283] The SpaCy engine analyzes text data and extracts keywords such as "yesterday," "friends," "Italian food," and "ate." It also performs contextual analysis to understand the relationships between keywords.

[1284] Output: Extracted keywords and context information.

[1285] Step 5:

[1286] The server uses IBM Watson Tone Analyzer to analyze emotional information from the extracted text data.

[1287] Input: Keywords and contextual information.

[1288] IBM Watson Tone Analyzer uses emotion recognition algorithms to identify emotions such as "happy" or "relaxed" from the tone and content of text.

[1289] Output: Recognized emotion information.

[1290] Step 6:

[1291] The server uses the Django Web Framework to update the user's profile based on the new information and the recognized emotion information.

[1292] Input: Recognized emotion information and extracted keywords.

[1293] The profile database stores the new information and adds characteristics to the profile, such as "I enjoy Italian food" and "I feel relaxed when I'm with friends."

[1294] Output: Updated user profile.

[1295] Step 7:

[1296] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1297] Input: The user's new voice data.

[1298] The terminal re-records this voice data, converts it into text data, and sends it to the server.

[1299] Output: The new text data received on the server.

[1300] Step 8:

[1301] The server parses the request, again using an emotion engine to take into account the user's current emotions, and generates the best advice using OpenAI's GPT-4 model.

[1302] Input: New text data and existing user profile.

[1303] OpenAI's GPT-4 model combines past profile information with your current emotional state to generate a recommendation like, "Why not try a new restaurant to rediscover that Italian food you enjoyed before?"

[1304] Output: The generated advice.

[1305] Step 9:

[1306] The terminal provides the generated advice to the user.

[1307] Input: The generated advice.

[1308] Using a text-to-speech engine, the device will verbally inform the user, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1309] Output: The user receives the advice.

[1310] (Application example 2)

[1311] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1312] Modern online shopping sites lack personalized product recommendations that utilize users' voice data and emotional information. In particular, there is a need for a system that can improve the user experience by providing optimal advice and products based on the user's emotional state. The lack of such personalized services is causing users to lose their motivation to purchase and making it difficult to select the right product.

[1313] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data from the user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and product recommendations for the user based on the updated profile, past conversation data, and emotional information, and means for providing the generated advice and product recommendations to the user. This enables personalized product recommendations that take into account the user's emotional state and past behavior.

[1314] "User's voice data" refers to voice information uttered by a user through a voice input device, and is data including voice signals such as conversations and instructions.

[1315] "Text data" is voice data converted into character information, and is data consisting of a sequence of characters.

[1316] "Keywords" are important words or phrases extracted from text data that are meaningful for analyzing context.

[1317] "Context" refers to the background and situation of sentences and phrases present in text data, and is information that helps understand the meaning.

[1318] A "profile" is a user information system constructed by combining a user's attribute information, behavioral history, preferences, emotional state, etc.

[1319] "Emotional information" is data that indicates emotional states such as joy, sadness, anger, and stress, analyzed from a user's statements and actions.

[1320] "Product recommendation" refers to the act and content of suggesting specific products or services to users based on their profile and emotional information.

[1321] "Appropriate advice" means providing useful information and guidance tailored to the user's needs and circumstances.

[1322] "Means for acquiring voice data" refers to a device or method for collecting voice information from a user using a microphone or the like.

[1323] The "means for converting voice data into text data" refers to a device or method for converting voice data into text information using voice recognition technology.

[1324] The "means for analyzing converted text data" refers to a device or method that uses natural language processing technology to analyze text data and extract context and meaning.

[1325] The "means for storing past conversation data" refers to a device or method for storing the history of dialogue with the user in a database or the like.

[1326] The "means for generating advice and product recommendations" is a device or method for generating appropriate advice and product recommendations based on the profile and emotion information.

[1327] A "means for providing generated advice and product recommendations" is a device or method that notifies a user of advice or product recommendations in audio or text form.

[1328] This invention is a system that acquires a user's voice data, converts it into text data, and provides personalized advice and product recommendations that incorporate emotional information. This system improves the user experience and helps them make more appropriate product selections.

[1329] The general flow of the system is as follows: First, the user inputs voice data using a device such as a smartphone. The voice data is acquired in real time and converted into text data using the Google Cloud Speech-to-Text API. This text data is then sent to the server.

[1330] After receiving the converted text data, the server analyzes it using the Google Cloud Natural Language API to extract keywords and context. It then uses IBM Watson Tone Analyzer to analyze emotional information from the text data. The server then updates the user's profile based on the emotional information and the extracted keywords and context.

[1331] When updating profiles, machine learning algorithms such as TensorFlow are used to sequentially learn and update the user's past conversation data and behavioral history, resulting in personalized information that reflects the user's preferences and emotional state in real time.

[1332] Next, when the user makes a new request, the system analyzes the request and references past conversation data and profiles to generate optimal advice and product recommendations, which are then sent to the user via a device such as a smartphone.

[1333] As a concrete example, if a user says, "I've been feeling stressed lately, so I want something that will help me relax," the following processing will occur.

[1334] First, the smartphone's microphone is used to capture voice data, which is then converted into text using the Google Cloud Speech-to-Text API. The converted text data is then sent to a server, where it is analyzed using the Google Natural Language API to extract keywords such as "stress" and "relaxation." Next, IBM Watson Tone Analyzer is used to analyze the emotional data and identify high stress levels.

[1335] Based on this information, the server uses machine learning algorithms such as TensorFlow to update the profile and generate appropriate product recommendations. For example, recommendations such as "Lavender aroma diffuser" or "Soothing music CD" are generated. The generated advice and product recommendations are then displayed on the user's smartphone as a notification saying, "How about a lavender aroma diffuser to relieve stress?"

[1336] The following is a specific example of a prompt sentence to be input to the generative AI model:

[1337] Prompt: "The user says, 'I'm tired and looking for something to relax.' Speech is converted to text and the keywords 'tired' and 'relaxation' are extracted. Sentiment analysis identifies this as a state of high stress. Based on the user's profile, it is determined that they need something to relax, and the system recommends the purchase of a 'lavender aroma diffuser,' which has a relaxing effect."

[1338] The system enables personalized product recommendations that take into account the user's emotional state and past behavior, improving the user experience.

[1339] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1340] Step 1:

[1341] The user speaks into the smartphone's microphone, saying, "I've been feeling stressed lately, so I want something that will help me relax." The input voice data is picked up by the smartphone's microphone. The output is the original voice data.

[1342] Step 2:

[1343] The device sends the acquired voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data and outputs the converted text data.

[1344] Step 3:

[1345] The terminal transmits the converted text data to the server, and the server receives the text data.

[1346] Step 4:

[1347] The server uses the Google Cloud Natural Language API to parse the received text data and extract keywords and context. The input is the received text data, and the output is the extracted keywords (e.g., "stress," "relax") and context information.

[1348] Step 5:

[1349] The server uses IBM Watson Tone Analyzer to analyze emotional information from text data. The input is text data, and the output is emotional data (e.g., high stress).

[1350] Step 6:

[1351] The server updates the user's profile based on the extracted keywords and emotion data. The profile is updated sequentially using machine learning algorithms such as TensorFlow based on the user's past conversation data and behavioral history. The input is keywords, context, emotion data, and past profile data, and the output is the updated profile.

[1352] Step 7:

[1353] When a user inputs a new request, the server analyzes the request. For example, if a user inputs "I want a relaxation item," the input voice data is converted into text data and analyzed. The input is the user's new request, and the output is the analysis result.

[1354] Step 8:

[1355] The server refers to the past conversation data and the updated profile to generate optimal advice and product recommendations. For example, a "lavender aroma diffuser" is recommended. The input is the parsed request, the past conversation data, and the updated profile, and the output is the generated advice and product recommendations.

[1356] Step 9:

[1357] The terminal notifies the user of the generated advice and product recommendation. For example, a notification such as "How about a lavender aroma diffuser to relieve stress?" is displayed. The input is the generated advice and product recommendation, and the output is a notification to the user.

[1358] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1359] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1360] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1361] [Fourth embodiment]

[1362] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1363] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1364] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1365] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1366] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1367] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1368] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1369] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1370] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1371] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1372] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1373] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1374] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1375] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes a series of means for collecting user input data, analyzing it, memorizing it, learning from it, and generating optimal advice. The specific operation of this system is shown below.

[1376] System Configuration

[1377] The system mainly consists of the following elements:

[1378] 1. Devices that collect user voice data

[1379] 2. A device with the processing power to convert voice data into text data

[1380] 3. A server that analyzes text data and extracts keywords and context

[1381] 4. A server that updates the user's profile based on the extracted information and stores past conversation data.

[1382] 5. Server and terminal that generates and notifies optimal advice based on the user's profile and past data

[1383] Program processing overview

[1384] 1. Data Collection

[1385] The user speaks to the device, "I had Italian food with my friends yesterday."

[1386] The device records this voice data and converts it into text data using a voice-to-text conversion engine.

[1387] The converted text data becomes "I ate Italian food with my friends yesterday."

[1388] 2. Data transmission and analysis

[1389] The terminal transmits the converted text data to the server.

[1390] The server receives the text data and analyzes it using natural language processing (NLP) techniques.

[1391] Keywords extracted through analysis include "yesterday," "friends," "Italian food," and "ate," and the context of each is also taken into consideration.

[1392] 3. Update your profile

[1393] The server adds new information to the user's profile, such as "I have a penchant for Italian food" or "I like spending time with friends."

[1394] The server uses machine learning algorithms to continuously update and learn from the user's profile.

[1395] 4. Generating and Providing Advice

[1396] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1397] The device records this voice data, converts it into text, and sends it to the server.

[1398] The server analyzes the request and uses historical profile data to generate appropriate advice.

[1399] For example, based on past data, suggestions might be generated such as, "I remember you enjoying Italian food. Why not try that newly opened restaurant?"

[1400] The device notifies the user of the generated advice.

[1401] Specific examples

[1402] Specifically, when a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend," the following process occurs.

[1403] The device records the voice data and converts it into text.

[1404] The server analyzes the text data and references past profile data.

[1405] The server takes into account factors such as "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past" and generates specific suggestions such as "the latest mystery novel."

[1406] The device provides the generated advice to the user, who can refer to it.

[1407] The system gives users a personalized AI partner that understands their preferences and past conversations, allowing them to receive faster, more relevant advice.

[1408] The processing flow will be explained below.

[1409] Step 1:

[1410] The user says to the device, "I had Italian food with my friends yesterday."

[1411] The device captures the user's speech as audio data.

[1412] Step 2:

[1413] The terminal converts the acquired voice data into text data using voice recognition technology.

[1414] The converted text data becomes "I ate Italian food with my friends yesterday."

[1415] Step 3:

[1416] The terminal transmits the converted text data to the server.

[1417] Step 4:

[1418] The server receives the text data and uses natural language processing (NLP) techniques to perform grammatical analysis and keyword extraction.

[1419] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1420] Step 5:

[1421] The server references past conversation data and the user's profile, and updates the profile based on newly extracted keywords and context.

[1422] For example, a user can add information to their profile such as "I like Italian food" or "I like spending time with friends."

[1423] Step 6:

[1424] The server uses machine learning models to learn user preferences and behavioral patterns and update the profile.

[1425] Step 7:

[1426] The user talks to the device, saying, "I'm wondering what to do on my day off."

[1427] The device acquires the user's speech as audio data and converts it into text data.

[1428] Step 8:

[1429] The terminal transmits the converted text data to the server.

[1430] Step 9:

[1431] The server receives the text data and analyzes the request using NLP techniques.

[1432] Keywords such as "holiday" and "not sure what to do" are extracted from the analysis results.

[1433] Step 10:

[1434] The server references past profile data and generates optimal advice based on the user's preferences and past behavior.

[1435] For example, a suggestion might be generated: "I remember enjoying Italian food. Why not try that newly opened restaurant?"

[1436] Step 11:

[1437] The server transmits the generated advice to the terminal.

[1438] Step 12:

[1439] The device will notify the user of the received advice in voice or text format.

[1440] This series of steps allows users to receive personalized advice, resulting in a more specific and reliable conversational experience.

[1441] Example 1

[1442] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1443] It is difficult for current AI systems to effectively memorize and analyze a user's past behavior and preferences and provide personalized advice based on the results. Conventional technologies lack sufficient user profile updates and past data reference for generating advice, resulting in a lack of improvement in the user experience. Therefore, there is a need for a system that can easily build a profile from a user's voice data and quickly provide appropriate advice.

[1444] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1445] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user using a generative AI model based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for receiving a new request from the user and analyzing the request to understand its intent, means for generating optimal advice by referring to past conversation data and the user's profile, means for generating specific suggestions based on the user's preferences and values, means for checking whether the suggested content matches the user's past actions and conversations, and means for providing the suggestion to the user if a match is found. This makes it possible to build a profile from the user's voice data and provide quick and accurate advice.

[1446] "Voice data" refers to digital data that records the voice uttered by the user.

[1447] "Text data" is digital data that has been converted from voice data into character information.

[1448] "Keywords" are important words or phrases extracted from text data.

[1449] "Context" refers to the semantic background, including the context of the keyword and the environment in which it is used.

[1450] A "user profile" is an individual data set that includes a user's preferences, past behavior, and conversational data.

[1451] "Generative AI models" are artificial intelligence algorithms and models used to analyze data and generate advice.

[1452] An "HTTP request" is a request of a communication protocol used to send data from a terminal to a server.

[1453] "Natural language processing" is a general term for technology that analyzes human language and extracts meaning.

[1454] A "machine learning algorithm" is a computational method for learning from data and making predictions or classifications.

[1455] "Advice" is a suggestion or advice provided by the system based on the user's profile and past behavior.

[1456] A "request" is a question or input instruction that a user makes to the system.

[1457] This invention relates to an AI system that can memorize a user's past conversations and preferences and provide personalized advice based on them. This system includes various means for collecting user input data, analyzing, memorizing, learning from it, and generating optimal advice. The specific operation of this system is described below.

[1458] System Configuration

[1459] The system mainly consists of the following elements:

[1460] 1. Devices that collect user voice data (smartphones, tablets, etc.)

[1461] 2. A speech-to-text engine (such as Google Cloud Speech-to-Text) to convert the audio data into text.

[1462] 3. A server that analyzes text data and extracts keywords and context (using SpaCy or NLTK)

[1463] 4. A server that updates user profiles based on the extracted information and stores past conversation data (using Scikit-learn and TensorFlow).

[1464] 5. Server and device that generates and notifies optimal advice based on the user's profile and past data (using generative AI models)

[1465] Data collection

[1466] The user asks questions or makes requests to the system in natural language. For example, they might say, "I had Italian food with my friends yesterday." This voice data is recorded by the device and converted into text data.

[1467] Data analysis and profile updates

[1468] The converted text data is sent to a server and analyzed using natural language processing technology. For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted from the text data. The server then adds this information to the user's profile and stores it as past conversation data.

[1469] Generating and Serving Advice

[1470] When the user makes a new request, such as "I'm not sure what to do on my day off," the device re-records this voice data, converts it to text data, and sends it to the server. The server analyzes the request and generates appropriate advice by referencing past profile data. A generative AI model is used to create appropriate suggestions based on the user's past data. For example, a suggestion might be, "I remember you enjoyed Italian food. Why not try that newly opened restaurant?"

[1471] Examples of concrete examples and prompts

[1472] As a concrete scenario, consider the case where a user asks for advice saying, "I'm having trouble deciding on a birthday present for my friend." In this case, the following flow will occur.

[1473] The device records audio data and converts it to text using Google Cloud Speech-to-Text.

[1474] The server analyzes the received text data using SpaCy and extracts keywords such as "friend," "birthday," "present," and "worried."

[1475] The server refers to the user profile and takes into account things like "friends' hobbies" and "tendencies in gift-giving that have been discussed in the past."

[1476] The server uses a generative AI model to generate specific suggestions such as "latest mystery novels."

[1477] The device notifies the user of the generated advice.

[1478] Users can use the suggestions to help them choose a gift.

[1479] Examples of prompts that users can enter include, "I'm not sure what to do on my day off. Could you give me some advice?" or "I'm having trouble deciding on a birthday present for my friend." Based on these prompts, the system can provide users with quick and accurate advice.

[1480] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1481] Program processing flow

[1482] Step 1:

[1483] The user speaks. Specifically, the user says to the terminal, "I ate Italian food with my friends yesterday," and voice data is generated.

[1484] Input: Audio data

[1485] Output: Audio data

[1486] Step 2:

[1487] The device converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text, for example, "I ate Italian food with my friends yesterday."

[1488] Specific operation: The device calls the voice recognition software and saves the voice data as text data.

[1489] Input: Audio data

[1490] Output: Text data

[1491] Step 3:

[1492] The terminal sends the converted text data to the server in JSON format using an HTTP request.

[1493] Specific operation: The terminal creates an HTTP request and sends text data as packets to the server.

[1494] Input: Text data

[1495] Output: HTTP request to the server

[1496] Step 4:

[1497] The server analyzes the received text data and extracts keywords and context from the text data using natural language processing tools (e.g., SpaCy, NLTK). For example, keywords such as "yesterday," "friends," "Italian food," and "ate" are extracted.

[1498] Specific operation: The server calls a natural language processing tool to break down the text data into keywords and context.

[1499] Input: Text data

[1500] Output: Extracted keywords and context data

[1501] Step 5:

[1502] The server updates the user's profile based on the extracted information and remembers past conversation data. Using machine learning algorithms (e.g., Scikit-learn, TensorFlow), the profile is updated and stored in a database. For example, profile information such as "I like Italian food" and "I like spending time with friends" are added.

[1503] What it does: The server uses machine learning algorithms to update the profile and save it in a database.

[1504] Input: Keywords and contextual data

[1505] Output: Updated user profile

[1506] Step 6:

[1507] The user issues a new request, for example, saying to the device, "I'm not sure what to do on my day off," and voice data is generated again.

[1508] Input: New audio data

[1509] Output: New audio data

[1510] Step 7:

[1511] The device converts this new voice data back into text data and sends it to the server. Google Cloud Speech-to-Text converts the voice data into text data, creates an HTTP request, and sends it to the server.

[1512] Specific operation: The device calls the voice recognition software, converts the voice data into text data, creates an HTTP request, and sends it to the server.

[1513] Input: New audio data

[1514] Output: Text data, HTTP request to server

[1515] Step 8:

[1516] The server analyzes the new text data and generates appropriate advice by referencing past profile data. Using a generative AI model (such as GPT-3), it generates advice based on past profile data, such as "I remember you enjoying Italian food. Why don't you try that newly opened restaurant?"

[1517] How it works: The server calls the generative AI model and generates new advice by referencing past profile data.

[1518] Input: New text data, user profile

[1519] Output: Generated advice

[1520] Step 9:

[1521] The device will notify the user of the generated advice, which will be provided to the user via text message or voice message.

[1522] Specific operation: The device displays or audibly notifies the user of the generated advice via the user interface.

[1523] Input: Generated advice

[1524] Output: User notification

[1525] (Application example 1)

[1526] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1527] In recent years, mail-order sales systems have been required to provide personalized product recommendations to users. However, conventional systems often fail to fully utilize the user's past conversations and preferences, and can only provide general recommendations. Therefore, a system that can provide accurate and individualized product recommendations to users is needed.

[1528] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1529] In this invention, the server includes means for acquiring voice data from a user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and conversation for the user based on the updated profile and past conversation data, means for providing the generated advice and conversation to the user, means for generating a proposal for a specific product based on the generated advice, and means for notifying the user of the proposed product. This makes it possible to make personalized product proposals by utilizing the user's past conversations and preferences.

[1530] "Voice data" refers to data that is a digital recording of voice information uttered by a user.

[1531] "Text data" is voice data expressed as text information.

[1532] "Keywords" are contextually significant words or phrases extracted from text data through analysis.

[1533] "Context" refers to the additional information or meaning that comes from the context in which a keyword is used.

[1534] A "profile" is a collection of personalized data that includes a user's preferences, past conversations, behavioral patterns, etc.

[1535] "Advice" refers to instructions or suggestions for actions or choices that are generated based on a user's profile and past conversation data.

[1536] "Product suggestions" are suggestions for specific products that users are encouraged to purchase, generated based on their profile and past conversation data.

[1537] "Notification" is a means of communication used to communicate system-generated advice and product suggestions to users.

[1538] This invention relates to a system that can memorize a user's past conversations and preferences and provide personalized advice and product suggestions based on them. The system includes a series of means for collecting, analyzing, memorizing, and learning from user input data to generate optimal advice and product suggestions.

[1539] System Configuration

[1540] The system mainly consists of the following elements:

[1541] 1. Speech recognition engine: Captures the user's voice data and converts it into text data. Specifically, this is done by collecting voice data using the microphone on a smartphone or smart glasses, and using voice recognition software such as the Google Speech-to-Text API.

[1542] 2. Text analysis engine: Analyzes text data and extracts keywords and context using natural language processing techniques such as Hugging Face's Transformer model.

[1543] 3. Profile Management Server: Updates the user's profile based on extracted keywords and context and remembers past conversation data.

[1544] 4. Advice Generation Engine: Generates appropriate advice and product suggestions for users based on their updated profile and past conversation data. Uses machine learning algorithms such as TensorFlow.

[1545] 5. Notification system: Provides generated advice and product suggestions to users. Software for notifying users of advice and product suggestions via smartphones or smart glasses.

[1546] Example of operation

[1547] When a user says to their smartphone, "I'm looking for a birthday present for my friend," the system operates as follows:

[1548] 1. Speech capture and text conversion:

[1549] The server converts the user's voice into text using the Google Speech-to-Text API.

[1550] 2. Text Analysis:

[1551] The text analysis engine uses Hugging Face's Transformer model to extract keywords such as "friend," "birthday," and "present" from the text data.

[1552] 3. Profile Update:

[1553] The profile management server updates the user's profile by referencing the extracted keywords and past data.

[1554] 4. Generating advice and suggestions:

[1555] An advice generation engine generates optimal birthday gift suggestions based on the updated profile.

[1556] 5. Notice:

[1557] A notification system sends generated suggestions to the user, such as "What new mystery novels or popular electronic gadgets might this friend like?"

[1558] Examples of AI-generated prompts

[1559] Prompts to suggest birthday gift ideas based on user voice data:

[1560] User dictation: "What would be a good birthday gift for my friend?"

[1561] Converts speech to text.

[1562] The converted text is then analyzed for keywords and context using a natural language processing model.

[1563] Extract important keywords such as "friends," "birthday," and "presents."

[1564] It retrieves related products from the database and notifies the user.

[1565] Example of desired result:

[1566] "How about a new mystery novel or a popular electronic gadget that this friend might like?"

[1567] This allows users to receive personalized advice based on an understanding of their preferences and past conversations, enabling them to choose the best product.

[1568] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1569] Step 1:

[1570] A user speaks to their smartphone, for example, "I'm looking for a birthday present for my friend." The smartphone's microphone captures the voice data, which is sent to the Google Speech-to-Text API. The Google Speech-to-Text API converts the voice data into text data, which becomes "I'm looking for a birthday present for my friend." The input is voice data, and the output is text data.

[1571] Step 2:

[1572] The text data is sent to the server. The server analyzes the text data using Hugging Face's Transformer model. This analysis extracts important keywords in the text, such as "friend," "birthday," and "present," as well as their context. The input is the text data, and the output is the keywords and context.

[1573] Step 3:

[1574] The server updates the user's profile based on the extracted keywords and context. It references the user's past conversation data and adds new information about "friends," "birthdays," and "gifts." This profile is stored in a database. The input is keywords and context, and the output is the updated profile.

[1575] Step 4:

[1576] When the user makes a new request, such as "Tell me what kind of gift would be good," the device receives this voice data and converts it into text data using the Google Speech-to-Text API again. The input is voice data, and the output is text data.

[1577] Step 5:

[1578] The server analyzes the new text data and again uses the Hugging Face Transformer model to understand its intent. It analyzes important keywords and context within the text and extracts intents such as "present" or "birthday." The input is the new text data, and the output is the intended keywords.

[1579] Step 6:

[1580] The server references past conversation data and the user's profile data to generate specific suggestions, such as "new mystery novels" or "popular electronic gadgets," using machine learning algorithms such as TensorFlow. The input is the updated profile and past data, and the output is specific product suggestions.

[1581] Step 7:

[1582] The server notifies the smartphone of the generated product suggestions. The smartphone screen displays a message saying, "How about a new mystery novel or a popular electronic gadget that your friend might like?" The input is the product suggestions, and the output is a notification to the user.

[1583] By going through the above processing steps, the user can receive personalized product suggestions based on past conversations and preferences.

[1584] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1585] This invention relates to a system that acquires a user's voice data and text data, analyzes that data, and provides personalized advice. In particular, by combining it with an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. The specific operating method and program processing of this system are described below.

[1586] System Configuration

[1587] The system consists of the following components:

[1588] 1. Device that collects user voice data

[1589] 2. A device with the processing power to convert voice data into text data

[1590] 3. A server that analyzes text data and extracts keywords and context

[1591] 4. A server that updates user profiles based on the extracted information and stores past conversation data.

[1592] 5. A server that uses an emotion engine to recognize user emotions and update the profile.

[1593] 6. Server and device that generates and notifies optimal advice based on the user's profile, past data, and emotional information

[1594] Program processing overview

[1595] 1. Data Collection

[1596] The user speaks to the device, "I had Italian food with my friends yesterday."

[1597] The device records this audio data and converts it into text data using a speech-to-text engine.

[1598] The converted text data becomes "I ate Italian food with my friends yesterday."

[1599] 2. Data transmission and analysis

[1600] The terminal transmits the converted text data to the server.

[1601] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[1602] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1603] 3. Emotion recognition

[1604] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[1605] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[1606] 4. Update your profile

[1607] The server updates the user's profile based on the new information and the recognized emotion information.

[1608] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[1609] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[1610] 5. Generating and Providing Advice

[1611] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1612] The terminal records this voice data, converts it into text data, and sends it to the server.

[1613] The server analyzes the request and generates the best advice using an emotion engine, taking into account the user's current emotions.

[1614] For example, if the user is in a relaxed state, the generated advice is, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1615] The device notifies the user of the generated advice in voice or text format.

[1616] Specific examples

[1617] For example, if a user asks for advice on what to buy as a birthday present for a friend, the following process will occur:

[1618] The terminal records the voice data and converts it into text data.

[1619] The server analyzes the text data and references past profile data and emotion information.

[1620] The server takes into consideration "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions such as "the latest mystery novel."

[1621] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[1622] The system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[1623] The processing flow will be explained below.

[1624] Step 1:

[1625] The user says to the device, "I had Italian food with my friends yesterday."

[1626] The device captures the user's speech as audio data.

[1627] Step 2:

[1628] The terminal converts the acquired voice data into text data using voice recognition technology.

[1629] The converted text data becomes "I ate Italian food with my friends yesterday."

[1630] Step 3:

[1631] The terminal transmits the converted text data to the server.

[1632] Step 4:

[1633] The server receives the text data and performs grammatical analysis and keyword extraction using natural language processing (NLP) techniques.

[1634] The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1635] Step 5:

[1636] The server uses an emotion engine to analyze emotion information from the user's voice data and text data.

[1637] For example, emotions such as "happy" or "relaxed" can be recognized from the tone of voice and choice of text.

[1638] Step 6:

[1639] The server updates the user's profile based on newly extracted keywords, context and recognized emotion information.

[1640] For example, characteristics such as "I enjoy Italian food" and "I feel relaxed when I'm with friends" can be added to a profile.

[1641] Step 7:

[1642] The server uses machine learning algorithms to continuously update and learn user profiles and emotional data.

[1643] Step 8:

[1644] The user talks to the device, saying, "I'm wondering what to do on my day off."

[1645] The device captures the user's speech as audio data and converts it into text data.

[1646] Step 9:

[1647] The terminal transmits the converted text data to the server.

[1648] Step 10:

[1649] The server analyzes the received text data to understand the intent and emotional information of the request.

[1650] For example, the keywords "holiday" and "not sure what to do" and the user's relaxed state are analyzed.

[1651] Step 11:

[1652] The server references past conversation data and profiles to generate optimal advice based on the user's preferences and past behavior.

[1653] For example, if the user is in a relaxed state, the system generates advice such as, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1654] Step 12:

[1655] The server transmits the generated advice to the terminal.

[1656] Step 13:

[1657] The device will notify the user of the received advice in voice or text format.

[1658] This series of steps allows users to receive faster, more personalized advice that adapts to their current emotional state and past behavior.

[1659] Example 2

[1660] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1661] Conventional user assistance systems simply convert user voice data into text data and extract context and keywords, making it difficult to provide personalized advice that takes into account the user's emotional state. Therefore, there is a need for a system that can generate and provide more appropriate advice based on the user's emotions, thereby improving the user experience.

[1662] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1663] In this invention, the server includes means for acquiring a user's voice data and converting it into text data, means for analyzing the text data and extracting keywords and context, and means for analyzing the user's emotional information using an emotion engine and updating the profile, thereby enabling the server to generate and provide appropriate advice and communication that takes into account the user's emotional state.

[1664] "Audio data from a user" refers to data that is a recording of what a user says.

[1665] "Text data" refers to character information obtained by analyzing voice data.

[1666] A "keyword" is a word or phrase that indicates important information from text data.

[1667] "Context" refers to information that indicates the relationship between keywords and the flow of the entire sentence.

[1668] "Profile" means data that records the characteristics and preferences of a user.

[1669] An "emotion engine" is software for analyzing emotional states from voice and text.

[1670] "Content data" refers to data relating to conversations and actions with users that have been recorded in the past.

[1671] "Advice" means any recommendation or instruction provided to a User.

[1672] "Communication" refers to activities that involve dialogue and exchange of information with users.

[1673] A "request" is a question or request made by a user to the system.

[1674] "Intent" refers to the goal or desire the user wants to achieve through their request.

[1675] "Past behavior" refers to the activities and behavioral history that a user has undertaken to date.

[1676] "Values" refer to the fundamental beliefs and ways of thinking that govern a user's thoughts and actions.

[1677] The present invention relates to a system that acquires a user's voice data and text data, analyzes them, and provides personalized advice. In particular, by combining an emotion engine, it is possible to grasp the user's emotional state and provide more appropriate advice and conversation. This system is implemented using the following hardware and software.

[1678] System configuration and hardware / software used

[1679] Hardware

[1680] 1. Device: A digital device with a microphone (e.g., smart speaker, smartphone) that captures user voice data.

[1681] 2. Server: A cloud or on-premise server for managing data analysis and profiles.

[1682] software

[1683] 1. Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[1684] 2. Natural Language Processing (NLP) engine: SpaCy is used to perform contextual analysis of text data and extract keywords.

[1685] 3. Emotion Recognition Engine: Analyzes user emotional information using IBM Watson Tone Analyzer.

[1686] 4. Machine learning algorithm: TensorFlow is used to continuously update and learn user profiles and emotion data.

[1687] 5. Advice Generation Engine: Uses OpenAI's GPT-4 model to generate optimal advice.

[1688] 6. Web Framework: Use Django Web Framework for profile management and database operations.

[1689] Specific system operation examples

[1690] Data collection

[1691] The user speaks to the device, "I had Italian food with my friends yesterday."

[1692] The device records this voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data will be "I ate Italian food with my friends yesterday."

[1693] Data analysis and profile updates

[1694] The device sends the converted text data to the server using the REST API.

[1695] The server receives the text data and uses SpaCy to perform grammatical analysis and extract keywords. The extracted keywords include "yesterday," "friends," "Italian food," and "ate."

[1696] The server then uses IBM Watson Tone Analyzer to analyze emotional information from the voice and text data and recognize emotions such as "happy" or "relaxed."

[1697] The server updates the user's profile with the new information and sentiment, adding traits such as "enjoys Italian food" and "is relaxed when with friends" to the database using the Django Web Framework.

[1698] Advice generation and delivery

[1699] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1700] The device records this voice data, converts it into text data using the Google Cloud Speech-to-Text API, and sends it to the server.

[1701] The server analyzes the request, takes into account the user's current emotions using an emotion engine, and then generates the optimal advice using OpenAI's GPT-4 model. For example, if the user is in a relaxed state, the advice generated might be, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1702] The device notifies the user of the generated advice in voice or text format.

[1703] Specific examples

[1704] As a specific example, consider the case where a user asks for advice, saying, "I'm having trouble deciding on a birthday present for my friend."

[1705] The device records audio data and converts it into text data using the Google Cloud Speech-to-Text API.

[1706] The server analyzes the text data and references past profile data and emotional information using IBM Watson Tone Analyzer.

[1707] Next, the server uses OpenAI's GPT-4 model to consider "friends' hobbies," "past gift trends," and "the user's current state of conversation," and generates specific gift suggestions, such as "the latest mystery novel."

[1708] The device will then notify the user of the generated advice: "Your friend might enjoy the latest mystery novel."

[1709] In this way, the system allows users to receive faster, more personalized advice that adapts to their emotional state and past behavior.

[1710] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1711] Step 1:

[1712] The user speaks to the device, "I had Italian food with my friends yesterday."

[1713] Input: User's voice data.

[1714] To record this audio data, the device uses a microphone to capture the audio.

[1715] Output: Recorded audio data file.

[1716] Step 2:

[1717] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts it into text data.

[1718] Input: Audio data file.

[1719] The Google Cloud Speech-to-Text API uses speech recognition technology to analyze audio data and generate corresponding text data.

[1720] Output: Text data: "I ate Italian food with my friends yesterday."

[1721] Step 3:

[1722] The device sends the converted text data to the server using the REST API.

[1723] Input: Text data.

[1724] The device packages the text data in JSON format and sends it to the server using an HTTP request.

[1725] Output: The text data received on the server.

[1726] Step 4:

[1727] The server analyzes the received text data using SpaCy, performing grammatical analysis and keyword extraction.

[1728] Input: Text data.

[1729] The SpaCy engine analyzes text data and extracts keywords such as "yesterday," "friends," "Italian food," and "ate." It also performs contextual analysis to understand the relationships between keywords.

[1730] Output: Extracted keywords and context information.

[1731] Step 5:

[1732] The server uses IBM Watson Tone Analyzer to analyze emotional information from the extracted text data.

[1733] Input: Keywords and contextual information.

[1734] IBM Watson Tone Analyzer uses emotion recognition algorithms to identify emotions such as "happy" or "relaxed" from the tone and content of text.

[1735] Output: Recognized emotion information.

[1736] Step 6:

[1737] The server uses the Django Web Framework to update the user's profile based on the new information and the recognized emotion information.

[1738] Input: Recognized emotion information and extracted keywords.

[1739] The profile database stores the new information and adds characteristics to the profile, such as "I enjoy Italian food" and "I feel relaxed when I'm with friends."

[1740] Output: Updated user profile.

[1741] Step 7:

[1742] The user inputs a new request into the terminal, saying, "I'm not sure what to do on my day off."

[1743] Input: The user's new voice data.

[1744] The terminal re-records this voice data, converts it into text data, and sends it to the server.

[1745] Output: The new text data received on the server.

[1746] Step 8:

[1747] The server parses the request, again using an emotion engine to take into account the user's current emotions, and generates the best advice using OpenAI's GPT-4 model.

[1748] Input: New text data and existing user profile.

[1749] OpenAI's GPT-4 model combines past profile information with your current emotional state to generate a recommendation like, "Why not try a new restaurant to rediscover that Italian food you enjoyed before?"

[1750] Output: The generated advice.

[1751] Step 9:

[1752] The terminal provides the generated advice to the user.

[1753] Input: The generated advice.

[1754] Using a text-to-speech engine, the device will verbally inform the user, "Why not try a new restaurant to rediscover the Italian food you enjoyed before?"

[1755] Output: The user receives the advice.

[1756] (Application example 2)

[1757] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1758] Modern online shopping sites lack personalized product recommendations that utilize users' voice data and emotional information. In particular, there is a need for a system that can improve the user experience by providing optimal advice and products based on the user's emotional state. The lack of such personalized services is causing users to lose their motivation to purchase and making it difficult to select the right product.

[1759] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring voice data from the user and converting it into text data, means for analyzing the converted text data and extracting keywords and context, means for updating the user's profile based on the extracted keywords and context and storing past conversation data, means for generating appropriate advice and product recommendations for the user based on the updated profile, past conversation data, and emotional information, and means for providing the generated advice and product recommendations to the user. This enables personalized product recommendations that take into account the user's emotional state and past behavior.

[1760] "User's voice data" refers to voice information uttered by a user through a voice input device, and is data including voice signals such as conversations and instructions.

[1761] "Text data" is voice data converted into character information, and is data consisting of a sequence of characters.

[1762] "Keywords" are important words or phrases extracted from text data that are meaningful for analyzing context.

[1763] "Context" refers to the background and situation of sentences and phrases present in text data, and is information that helps understand the meaning.

[1764] A "profile" is a user information system constructed by combining a user's attribute information, behavioral history, preferences, emotional state, etc.

[1765] "Emotional information" is data that indicates emotional states such as joy, sadness, anger, and stress, analyzed from a user's statements and actions.

[1766] "Product recommendation" refers to the act and content of suggesting specific products or services to users based on their profile and emotional information.

[1767] "Appropriate advice" means providing useful information and guidance tailored to the user's needs and circumstances.

[1768] "Means for acquiring voice data" refers to a device or method for collecting voice information from a user using a microphone or the like.

[1769] The "means for converting voice data into text data" refers to a device or method for converting voice data into text information using voice recognition technology.

[1770] The "means for analyzing converted text data" refers to a device or method that uses natural language processing technology to analyze text data and extract context and meaning.

[1771] The "means for storing past conversation data" refers to a device or method for storing the history of dialogue with the user in a database or the like.

[1772] The "means for generating advice and product recommendations" is a device or method for generating appropriate advice and product recommendations based on the profile and emotion information.

[1773] A "means for providing generated advice and product recommendations" is a device or method that notifies a user of advice or product recommendations in audio or text form.

[1774] This invention is a system that acquires a user's voice data, converts it into text data, and provides personalized advice and product recommendations that incorporate emotional information. This system improves the user experience and helps them make more appropriate product selections.

[1775] The general flow of the system is as follows: First, the user inputs voice data using a device such as a smartphone. The voice data is acquired in real time and converted into text data using the Google Cloud Speech-to-Text API. This text data is then sent to the server.

[1776] After receiving the converted text data, the server analyzes it using the Google Cloud Natural Language API to extract keywords and context. It then uses IBM Watson Tone Analyzer to analyze emotional information from the text data. The server then updates the user's profile based on the emotional information and the extracted keywords and context.

[1777] When updating profiles, machine learning algorithms such as TensorFlow are used to sequentially learn and update the user's past conversation data and behavioral history, resulting in personalized information that reflects the user's preferences and emotional state in real time.

[1778] Next, when the user makes a new request, the system analyzes the request and references past conversation data and profiles to generate optimal advice and product recommendations, which are then sent to the user via a device such as a smartphone.

[1779] As a concrete example, if a user says, "I've been feeling stressed lately, so I want something that will help me relax," the following processing will occur.

[1780] First, the smartphone's microphone is used to capture voice data, which is then converted into text using the Google Cloud Speech-to-Text API. The converted text data is then sent to a server, where it is analyzed using the Google Natural Language API to extract keywords such as "stress" and "relaxation." Next, IBM Watson Tone Analyzer is used to analyze the emotional data and identify high stress levels.

[1781] Based on this information, the server uses machine learning algorithms such as TensorFlow to update the profile and generate appropriate product recommendations. For example, recommendations such as "Lavender aroma diffuser" or "Soothing music CD" are generated. The generated advice and product recommendations are then displayed on the user's smartphone as a notification saying, "How about a lavender aroma diffuser to relieve stress?"

[1782] The following is a specific example of a prompt sentence to be input to the generative AI model:

[1783] Prompt: "The user says, 'I'm tired and looking for something to relax.' Speech is converted to text and the keywords 'tired' and 'relaxation' are extracted. Sentiment analysis identifies this as a state of high stress. Based on the user's profile, it is determined that they need something to relax, and the system recommends the purchase of a 'lavender aroma diffuser,' which has a relaxing effect."

[1784] The system enables personalized product recommendations that take into account the user's emotional state and past behavior, improving the user experience.

[1785] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1786] Step 1:

[1787] The user speaks into the smartphone's microphone, saying, "I've been feeling stressed lately, so I want something that will help me relax." The input voice data is picked up by the smartphone's microphone. The output is the original voice data.

[1788] Step 2:

[1789] The device sends the acquired voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data and outputs the converted text data.

[1790] Step 3:

[1791] The terminal transmits the converted text data to the server, and the server receives the text data.

[1792] Step 4:

[1793] The server uses the Google Cloud Natural Language API to parse the received text data and extract keywords and context. The input is the received text data, and the output is the extracted keywords (e.g., "stress," "relax") and context information.

[1794] Step 5:

[1795] The server uses IBM Watson Tone Analyzer to analyze emotional information from text data. The input is text data, and the output is emotional data (e.g., high stress).

[1796] Step 6:

[1797] The server updates the user's profile based on the extracted keywords and emotion data. The profile is updated sequentially using machine learning algorithms such as TensorFlow based on the user's past conversation data and behavioral history. The input is keywords, context, emotion data, and past profile data, and the output is the updated profile.

[1798] Step 7:

[1799] When a user inputs a new request, the server analyzes the request. For example, if a user inputs "I want a relaxation item," the input voice data is converted into text data and analyzed. The input is the user's new request, and the output is the analysis result.

[1800] Step 8:

[1801] The server refers to the past conversation data and the updated profile to generate optimal advice and product recommendations. For example, a "lavender aroma diffuser" is recommended. The input is the parsed request, the past conversation data, and the updated profile, and the output is the generated advice and product recommendations.

[1802] Step 9:

[1803] The terminal notifies the user of the generated advice and product recommendation. For example, a notification such as "How about a lavender aroma diffuser to relieve stress?" is displayed. The input is the generated advice and product recommendation, and the output is a notification to the user.

[1804] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1805] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1806] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1807] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1808] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1809] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1810] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1811] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1812] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1813] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1814] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1815] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1816] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1817] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1818] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1819] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1820] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1821] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1822] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1823] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1824] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1825] The following is further disclosed regarding the above embodiment.

[1826] (Claim 1)

[1827] A means for acquiring voice data from a user and converting it into text data;

[1828] means for analyzing the converted text data and extracting keywords and context;

[1829] means for updating a user's profile and storing past conversation data based on the extracted keywords and context;

[1830] means for generating appropriate advice and conversations for the user based on the updated profile and past conversation data;

[1831] The system includes a means for providing generated advice and conversations to a user.

[1832] (Claim 2)

[1833] A means of receiving new requests from users and parsing those requests to understand their intent;

[1834] A means of generating optimal advice by referencing past conversation data and user profiles;

[1835] 10. The system of claim 1, further comprising means for notifying a user of the generated advice.

[1836] (Claim 3)

[1837] A means of generating specific recommendations based on the user's preferences and values;

[1838] A way to verify whether the suggestions match the user's past behavior and conversations, and

[1839] 10. The system of claim 1, further comprising means for providing a suggestion to the user if a match is found.

[1840] "Example 1"

[1841] (Claim 1)

[1842] A means for acquiring voice data from a user and converting it into text data;

[1843] means for analyzing the converted text data and extracting keywords and context;

[1844] means for updating a user's profile and storing past conversation data based on the extracted keywords and context;

[1845] a means for generating appropriate advice and conversations for the user using a generative AI model based on the updated profile and past conversation data;

[1846] The system includes a means for providing generated advice and conversations to a user.

[1847] (Claim 2)

[1848] A means of receiving new requests from users and parsing those requests to understand their intent;

[1849] A means of generating optimal advice by referencing past conversation data and user profiles;

[1850] 10. The system of claim 1, further comprising means for notifying a user of the generated advice.

[1851] (Claim 3)

[1852] A means of generating specific recommendations based on the user's preferences and values;

[1853] A way to verify whether the suggestions match the user's past behavior and conversations, and

[1854] 10. The system of claim 1, further comprising means for providing a suggestion to the user if a match is found.

[1855] "Application Example 1"

[1856] (Claim 1)

[1857] A means for acquiring voice data from a user and converting it into text data;

[1858] means for analyzing the converted text data and extracting keywords and context;

[1859] means for updating a user's profile and storing past conversation data based on the extracted keywords and context;

[1860] means for generating appropriate advice and conversations for the user based on the updated profile and past conversation data;

[1861] a means for providing the generated advice and conversation to the user;

[1862] means for generating recommendations for specific products based on the generated advice;

[1863] The system includes a means for notifying the user of suggested products.

[1864] (Claim 2)

[1865] A means of receiving new requests from users and parsing those requests to understand their intent;

[1866] A means of generating optimal advice by referencing past conversation data and user profiles;

[1867] 2. The system according to claim 1, further comprising means for generating specific product suggestions based on the generated advice and notifying the user of the suggestions.

[1868] (Claim 3)

[1869] A means of generating specific recommendations based on the user's preferences and values;

[1870] A way to verify whether the suggestions match the user's past behavior and conversations, and

[1871] 10. The system of claim 1, further comprising means for generating and providing product suggestions to the user if a match is found.

[1872] "Example 2: Combining Emotion Engines"

[1873] (Claim 1)

[1874] A means for acquiring voice data from a user and converting it into text data;

[1875] means for analyzing the converted text data and extracting keywords and context;

[1876] means for updating a user's profile and storing historical content data based on the extracted keywords and context;

[1877] a means for analyzing the user's emotional information and updating the profile using an emotion engine;

[1878] means for generating appropriate advice and communications for the user based on the updated profile and historical content data;

[1879] A system including a means for providing generated advice and communications to a user.

[1880] (Claim 2)

[1881] A means for receiving a new request from a user, analyzing the request to understand the intent, and generating optimal advice taking into account emotional information;

[1882] 10. The system of claim 1, further comprising means for notifying a user of the generated advice.

[1883] (Claim 3)

[1884] A means of generating specific recommendations based on the user's preferences and values;

[1885] A way to verify whether the suggestions match the user's past behavior and communications;

[1886] 10. The system of claim 1, further comprising means for providing a suggestion to the user if a match is found.

[1887] "Application example 2 when combining emotion engines"

[1888] (Claim 1)

[1889] A means for acquiring voice data from a user and converting it into text data;

[1890] means for analyzing the converted text data and extracting keywords and context;

[1891] means for updating a user's profile and storing past conversation data based on the extracted keywords and context;

[1892] means for generating appropriate advice and product recommendations for the user based on the updated profile and past conversation data and sentiment information;

[1893] The system includes a means for providing the generated advice and product recommendations to a user.

[1894] (Claim 2)

[1895] A means of receiving new requests from users and parsing those requests to understand their intent;

[1896] A means for generating optimal advice and product recommendations by referencing past conversation data, user profile and sentiment information;

[1897] 10. The system of claim 1, further comprising means for notifying a user of the generated advice and product recommendations.

[1898] (Claim 3)

[1899] means for generating product recommendations based on specific suggestions and emotional states of users based on their preferences and values;

[1900] A means to verify whether the suggestions and product recommendations match the user's past behavior and conversations;

[1901] 10. The system of claim 1, further comprising means for providing the suggestion and product recommendation to the user if a match is found. [Explanation of symbols]

[1902] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring voice data from a user and converting it into text data; means for analyzing the converted text data and extracting keywords and context; means for updating a user's profile and storing past conversation data based on the extracted keywords and context; means for generating appropriate advice and conversations for the user based on the updated profile and past conversation data; The system includes a means for providing generated advice and conversations to a user.

2. A means of receiving new requests from users and parsing those requests to understand their intent; A means of generating optimal advice by referencing past conversation data and user profiles; 10. The system of claim 1, further comprising means for notifying a user of the generated advice.

3. A means of generating specific recommendations based on the user's preferences and values; A way to verify whether the suggestions match the user's past behavior and conversations, and 10. The system of claim 1, further comprising means for providing a suggestion to the user if a match is found.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A