System

The system addresses low adoption and customization challenges in smart speakers by acquiring user data, generating profiles, and personalizing functions, facilitating easy setup and adaptive responses.

JP2026016221APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117311
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The smart speaker market faces challenges such as low adoption rates, cumbersome installation and setup, difficulty in differentiating products, and a lack of customization to meet individual user needs.

Method used

A system that includes means for acquiring user data, analyzing it to generate a profile, customizing smart speaker functions, processing voice commands, enabling user authentication, and incorporating user feedback to enhance responses, allowing for easy installation and advanced personalization.

Benefits of technology

Enables users to easily install and configure smart speakers with advanced customization, providing a tailored experience that meets individual needs and adapts to user preferences and feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016221000001_ABST
    Figure 2026016221000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system including means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing a smart speaker function based on the generated profile, means for inputting a voice command by a user, means for voice-recognizing the input voice command and converting the voice command into text data, means for analyzing the converted text data to generate an appropriate response, and means for outputting the generated response as voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The current smart speaker market faces challenges such as low adoption rates and technical barriers, making installation and setup cumbersome for average consumers. Furthermore, it is difficult to differentiate smart speakers from competitors' products, and there is a lack of customization to meet individual user needs. The present invention aims to solve these challenges and provide users with a simple, highly customized smart speaker experience. [Means for solving the problem]

[0005] The present invention relates to a system including a means for acquiring user data, a means for analyzing the user data to generate a profile, a means for customizing smart speaker functions based on the generated profile, a means for a user to input a voice command, a means for speech recognition of the input voice command and converting it into text data, a means for analyzing the converted text data to generate an appropriate response, and a means for outputting the generated response as speech. The system also includes a means for replacing a communication device with a new communication device, a means for acquiring initial setup information and performing user authentication, and a means for activating smart speaker functions after successful authentication. The system also includes a means for acquiring user feedback and a means for analyzing the acquired feedback and reflecting it in future responses. This system allows even ordinary consumers to easily install and configure smart speakers and enables advanced customization to meet individual needs.

[0006] "User data" refers to information including a user's behavioral history, preferences, and personal information, and is data used to customize smart speaker functionality.

[0007] "Analysis" is the process of analyzing acquired user data to identify the user's behavioral patterns and preferences and generate a profile.

[0008] A "profile" is setting information that reflects each user's characteristics and preferences based on user data, and is information used to customize the functions of a smart speaker.

[0009] The "smart speaker function" is a function that allows you to operate the device using voice commands, provide information, control home automation, and more.

[0010] A "voice command" refers to an instruction or request that a user makes verbally to a smart speaker.

[0011] "Speech recognition" is a technology that converts voice commands into text data, and is the process of analyzing voice and recognizing its content as text information.

[0012] A "communication device" is a device for connecting to the Internet and transmitting data, and is an equipment that is integrated with smart speaker functionality.

[0013] "User authentication" is the process of identifying a user based on initial configuration information and verifying their eligibility to use a service.

[0014] "Feedback" refers to the ratings and opinions users provide in response to smart speaker responses.

[0015] "Response" refers to the smart speaker's reaction or reply to the user's voice command, including providing information or executing operational instructions via voice. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention relates to a system for providing customized smart speaker functionality based on the collection and analysis of user data.

[0038] Explanation of program processing

[0039] Collection and analysis of user data

[0040] The server collects user data, including user behavior history, preferences, and personal information. For example, the server analyzes the user's interests based on the user's internet search history and application usage data. From the results of this analysis, a customized profile is created for each user.

[0041] Replacing and configuring communication devices

[0042] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user performs basic settings using an initial setup wizard. The initial setup includes Wi-Fi settings and user ID entry.

[0043] Enabling the smart speaker feature

[0044] The server receives the initial setup information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. This applies the customized profile to the communication device. The user can immediately use the smart speaker function.

[0045] Processing voice commands

[0046] The user inputs a voice command into the smart speaker. For example, they might say, "Tell me what the weather will be tomorrow." This voice command is recognized by the device and converted into text data.

[0047] The server receives the text data that has been recognized by speech recognition and analyzes the command. Based on the analysis results, it generates an appropriate response. For example, it may generate a response such as "Tomorrow's weather will be sunny." This response data is then sent to the terminal and output as voice.

[0048] Handling User Feedback

[0049] Users can provide feedback on the smart speaker's response, for example by stating their opinion such as "Please be more specific." This feedback is recognized by the device and then sent to the server.

[0050] The server has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the smart speaker's responses to evolve to more closely match user expectations.

[0051] Specific examples

[0052] After User A installs new communication equipment and completes the initial setup, the following scenario is assumed:

[0053] 1. User A: "What's in the news today?"

[0054] 2. Terminal: Recognizes speech and converts it into text.

[0055] 3. Server: Analyzes the text data, collects news information, and generates appropriate responses.

[0056] 4. Response: "The major news stories today are the stock market boom and the announcement of new government policies."

[0057] 5. Terminal: Provides the response as audio to the user.

[0058] In this way, the present invention provides users with an optimized smart speaker experience and realizes a system that is easy to use and intuitive.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] The server retrieves user data from LINE and Yahoo! databases, including user message history, search history, and app usage data.

[0062] Step 2:

[0063] The server analyzes the acquired user data and uses machine learning algorithms to identify user behavior patterns and preferences, generating a user profile from the analysis results.

[0064] Step 3:

[0065] The device is a new communication device (with smart speaker functionality) that is installed to replace the existing communication device. The new device requires basic settings to establish a network connection.

[0066] Step 4:

[0067] Users go through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID.

[0068] Step 5:

[0069] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[0070] Step 6:

[0071] The server then sends the generated user profile to the device, reflecting the individual settings, and the user is ready to use the customized smart speaker functions.

[0072] Step 7:

[0073] A user inputs a voice command to a smart speaker, for example, "Tell me the weather tomorrow."

[0074] Step 8:

[0075] The device converts the received voice commands into text data using a voice recognition algorithm, which is then sent to the server.

[0076] Step 9:

[0077] The server analyzes the text data and generates an appropriate response, for example, "Tomorrow's weather will be sunny."

[0078] Step 10:

[0079] The server sends the generated response data to the terminal, which runs a speech synthesis algorithm to convert the data into speech data.

[0080] Step 11:

[0081] The terminal outputs the audio data from a speaker to provide it to the user.

[0082] Step 12:

[0083] The user provides feedback on the smart speaker's response, such as saying, "The response was accurate" or "I wish it was shorter."

[0084] Step 13:

[0085] The device recognizes the user's voice and sends the feedback to the server, which analyzes it and uses it in future responses.

[0086] Example 1

[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0088] Conventional smart devices have been unable to adequately adapt to individual users' needs and habits. It has been particularly difficult to effectively analyze user data and use smart speakers in a more personalized way. Initial setup and device replacement are also burdensome for users, making it difficult to provide responses that meet user expectations. Furthermore, there has been a lack of mechanisms for utilizing user feedback to improve responses.

[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0090] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a personal profile, and means for customizing smart device functions based on the generated personal profile, thereby enabling the provision of personalized smart speaker functions according to the individual needs and habits of the user.

[0091] "User data" refers to information including a user's behavioral history, internet search history, data on applications used, personal information, and preferences.

[0092] A "personal profile" is information that indicates a user's interests, lifestyle habits, etc., generated as a result of analyzing user data.

[0093] "Smart device functionality" is the functionality of a digital device that provides voice recognition and voice response and acts based on user instructions.

[0094] A "voice instruction" is a voice command issued by a user to a smart device.

[0095] "Text information" is data in a document format that is generated by converting voice instructions using a voice recognition system.

[0096] A "response" is a reply in text or audio format that the server generates as a result of analyzing a voice command.

[0097] A "communications device" is an electronic device that allows for the transmission and reception of data.

[0098] "User authentication" is the process of identifying a particular user and verifying their permissions.

[0099] "Feedback" refers to opinions and suggestions that a user provides in response to a smart device's response.

[0100] The present invention is a system that provides customized smart device functions based on the collection and analysis of user data. This system collects data on user behavior history and applications used, and generates a customized profile for each user, thereby providing a personalized experience tailored to individual needs.

[0101] Hardware and software used

[0102] The main hardware required to implement the system includes a smart speaker, an internet-enabled device, and a server, while the software includes a voice recognition system, a data analysis engine, and a user profile generation algorithm.

[0103] Collection and analysis of user data

[0104] The server collects data about users' internet browsing history and the applications they use. This data comes from a variety of sources, including cookie information and log data. The collected data is analyzed using algorithms, such as machine learning models, to generate a personalized profile for each user. This profile reflects the user's interests, preferences, and preferences.

[0105] Replacing and initial setup of communication equipment

[0106] The terminal replaces the existing communication device with a new one. The new communication device has smart device functionality integrated into it. After the device is installed, the user uses an initial setup wizard to connect to a Wi-Fi network and enter their user ID. This configuration information is sent to the server.

[0107] Enabling the smart speaker feature

[0108] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device functions are enabled and the customized profile is applied, allowing the user to immediately start using the smart speaker.

[0109] Processing voice instructions

[0110] A user gives a voice command to a smart speaker. For example, they might say, "Tell me what the weather will be like tomorrow." This voice is recorded on the device and converted into text information in real time by a voice recognition system. This text information is sent to a server and analyzed. Based on the analysis results, an appropriate response is generated. For example, a response such as "Tomorrow's weather will be sunny" is generated, and this information is sent to the device and provided to the user as voice.

[0111] Handling User Feedback

[0112] Users can provide feedback on the smart speaker's response by voice. For example, they can say, "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server implements an algorithm that analyzes this feedback and reflects it in future responses. This allows the smart speaker's responses to better meet user expectations.

[0113] Examples of prompt statements

[0114] Examples of prompts for generative AI models include:

[0115] "Please tell me the details of the program processing for the user data collection and analysis system. Please provide a detailed explanation, including the names of the hardware and software used, and the details of data processing and calculation."

[0116] As described above, the present invention provides users with an optimized smart device experience and realizes a system that is easy to use and intuitive.

[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0118] Step 1: Collect user data

[0119] The server collects user data such as internet search history, application usage data, and behavioral history. Specifically, it collects cookie information and log data. It also collects personal data that the user has authorized.

[0120] Input: Internet logs, cookie data, application usage data

[0121] Data processing: Filtering log data and removing unnecessary data

[0122] Output: The curated user dataset

[0123] Step 2: Analyze data and generate profiles

[0124] The server analyzes the collected user data, specifically using machine learning algorithms to classify the user's interests and concerns, and generates a personalized profile for each user based on the results of this analysis.

[0125] Input: Curated user dataset

[0126] Data Computing: Analysis based on machine learning models

[0127] Output: A customized personal profile

[0128] Step 3: Replace the communication device

[0129] The terminal replaces the old communication device with a new one, which has integrated smart device functionality, and a technician visits the customer to physically install the device and connect the wiring.

[0130] Input: Existing communication equipment, replacement communication equipment

[0131] Data processing: None (physical exchange work)

[0132] Output: New communications device installed

[0133] Step 4: Perform initial configuration

[0134] After installing a new communication device, the user uses the initial setup wizard to perform basic settings, such as connecting to a Wi-Fi network and entering a user ID. This setting information is then sent to the server.

[0135] Input: Wi-Fi information, user ID

[0136] Data processing: Format conversion of setting information

[0137] Output: Send initial setting information to server

[0138] Step 5: Enable the Smart Speaker feature

[0139] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device function is enabled and the customized profile is applied. The user receives a notification that "Smart speaker has been enabled."

[0140] Input: Initial setting information

[0141] Data calculation: User authentication

[0142] Output: Enable smart device features, send notifications

[0143] Step 6: Processing voice instructions

[0144] The user gives voice commands to the smart speaker. The voice is recorded on the device and converted into text information in real time by a speech recognition system. This text information is sent to the server and analyzed. Based on the analysis results, an appropriate response is generated and provided to the user as voice.

[0145] Input: Voice commands

[0146] Data processing: Text conversion using speech recognition

[0147] Output: Response based on analysis results, providing voice response

[0148] Step 7: Processing user feedback

[0149] The user provides feedback on the smart speaker's response. For example, they may state their opinion, such as "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server analyzes this feedback and improves the response algorithm. This allows future responses to more faithfully meet the user's expectations.

[0150] Input: Feedback voice

[0151] Data processing: Text conversion by voice recognition, feedback analysis

[0152] Output: Improved response algorithm

[0153] In this way, the system provides a personalized smart device experience tailored to the user's individual needs.

[0154] (Application example 1)

[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0156] Conventional automotive voice assistant systems are unable to fully reflect the user's individual preferences and behavioral history, resulting in uniform information and services that make it difficult to provide an optimal experience for the user. Furthermore, adjusting the in-car environment requires a lot of manual settings, placing a heavy burden on the user. The present invention aims to solve these problems and realize a personalized in-car assistant based on the user's preferences and behavioral history.

[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0158] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing the voice assistant function based on the generated profile, means for the user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for analyzing the converted text data to generate an appropriate response, means for outputting the generated response as voice, means for generating and transmitting a control signal for adjusting the environment inside the vehicle, and means for automatically adjusting the in-vehicle environment settings (temperature, lighting, music, etc.) based on the user's preferences. This enables personalized information provision based on the user's preferences and behavioral history, and automatic adjustment of the in-vehicle environment.

[0159] "User data" refers to data such as a user's behavioral history, preferences, and personal information.

[0160] A "profile" is a collection of information that indicates a user's characteristics and preferences, generated by analyzing user data.

[0161] "Voice assistant function" refers to the function of receiving voice commands, generating appropriate responses, and outputting them in voice.

[0162] "Voice command" refers to instructions or questions that a user enters by voice.

[0163] "Speech recognition" refers to the technology of converting voice data into text data.

[0164] "Text data" refers to text information generated by speech recognition.

[0165] A "response" refers to information or instructions provided to a user that are generated based on the analyzed text data.

[0166] "Output" refers to conveying the generated response to the user as voice.

[0167] "Control signal" refers to a command signal sent to control a device within a vehicle.

[0168] "Environmental settings" refers to settings such as temperature, lighting, and music inside the car.

[0169] "Preferences" refer to the user's tastes and interests.

[0170] "Initial setting information" refers to the setting information required to activate a new communication device.

[0171] "User authentication" is a procedure for verifying that a user is a legitimate user.

[0172] "Feedback" refers to evaluations and opinions regarding responses and services provided by users.

[0173] "Analysis" refers to the process of analyzing acquired data to extract useful information.

[0174] The present invention relates to a system for providing personalized information and automatically adjusting the in-vehicle environment based on the user's preferences and behavioral history in an autonomous vehicle.

[0175] The server acquires user data, including the user's behavioral history, preferences, and personal information. The data is acquired using internet search history and application usage data. By analyzing this data, a customized profile for each user is created.

[0176] The server customizes the voice assistant function based on the generated profile, enabling it to provide optimal information to the user. Additionally, when a user enters a voice command, the terminal (communication equipment inside the autonomous vehicle) uses voice recognition technology to convert the voice command into text data, using the Google Cloud Speech-to-Text API.

[0177] The server analyzes the converted text data and generates an appropriate response using the Python data analysis library scikit-learn. The generated response is then converted into audio using the Google Cloud Text-to-Speech API, which is then output to the user through the car's speakers.

[0178] Furthermore, the terminal generates and transmits control signals to adjust the vehicle's interior environment, which are transmitted via the MQTT protocol and control various devices (such as temperature control, lighting, and audio) via the vehicle's CAN bus (in-vehicle communication network).

[0179] Users can provide feedback on the provided responses by voice, which is then re-audio-recognized and sent to the server, which has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the voice assistant to continually improve its accuracy.

[0180] For example, if a user says to a smart speaker installed in the car, "Tell me today's news," the personal in-car assistant will automatically retrieve the news and respond aloud, "Today's main news stories are a sharp rise in stock prices and the announcement of new government policies."

[0181] Next, say "Make the car warmer," and the car's temperature will be adjusted to the appropriate level based on the user's profile.

[0182] Examples of prompt sentences that are important in the present invention include the following:

[0183] 1. "Collect user data and generate user profiles"

[0184] 2. "Analyzing user preferences based on internet search history"

[0185] 3. "Accept voice commands from a smart speaker and convert the voice to text."

[0186] 4. "Improve system response based on user feedback"

[0187] This invention makes it possible to provide information optimized for the user in an autonomous vehicle and automatically adjust the in-vehicle environment to be comfortable.

[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0189] Step 1:

[0190] User Data Collection:

[0191] The server collects data such as user behavior, preferences, and personal information. This includes internet browsing history and application usage data. This data is obtained by using APIs to collect data from external services. It may also be obtained through sensors or other interfaces. The input for this step is internet browsing history and application usage data, and the output is raw user data for analysis.

[0192] Step 2:

[0193] User data analysis:

[0194] The server analyzes the collected user data to understand user preferences and behavioral patterns. It uses Python and data analysis libraries such as scikit-learn to perform the analysis. It preprocesses the raw data and generates profiles using techniques such as feature extraction, clustering, and classification. The input for this step is raw user data, and the output is a profile customized for each user.

[0195] Step 3:

[0196] Creating profile-based customizations:

[0197] Based on the generated user profile, the server customizes the voice assistant's functionality by specifying the type of music, news, and information the user prefers, and also applies algorithms to provide appropriate feedback and adjustments. The input for this step is the generated profile, and the output is customized voice assistant settings.

[0198] Step 4:

[0199] Receiving and recognizing voice commands:

[0200] When a user inputs a voice command into the smart speaker in the car, the device captures the speech and converts it into text data using the Google Cloud Speech-to-Text API. The input of this step is the user's voice command, and the output is text data.

[0201] Step 5:

[0202] Voice command analysis:

[0203] The server analyzes the speech-recognized text data and generates an appropriate response to the user's request. The analysis uses natural language processing and rule-based algorithms. The server also takes into account the user's profile to provide the most appropriate information and services. The input for this step is text data, and the output is response data.

[0204] Step 6:

[0205] Response transcription and output:

[0206] The server converts the generated response data into speech using the Google Cloud Text-to-Speech API. The device then outputs the converted speech to the user through the car speaker. The input of this step is the response data, and the output is a speech response.

[0207] Step 7:

[0208] Control signal generation and transmission:

[0209] The server generates control signals to adjust the in-car environment based on the user's preference profile and sends them to the terminal via the MQTT protocol. The terminal receives these control signals and controls each device (temperature, lighting, audio, etc.) via the vehicle's CAN bus. The input of this step is the user profile and control request, and the output is the control signal sent to each device.

[0210] Step 8:

[0211] Collecting and analyzing user feedback:

[0212] Users can provide feedback on the voice assistant's responses, which are then re-audio-recognized and sent to the server. The server then has an algorithm that analyzes the feedback and incorporates it into future responses, improving the accuracy of the voice assistant. The input for this step is the user's voice feedback, and the output is the analysis results and the system adjustments based on them.

[0213] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0214] The present invention relates to a system that provides customized smart speaker functions based on the collection and analysis of user data, and further combines an emotion engine to respond to the user's emotional state.

[0215] Explanation of program processing

[0216] Collection and analysis of user data

[0217] The server retrieves user data from the internet service database, including the user's message history, search history, and app usage data, and uses this data to analyze the user's behavioral patterns and preferences and create a profile.

[0218] As a specific example, the server analyzes the news articles that the user has viewed and the product data that the user has purchased, and can provide news and product information through the speaker according to the user's preferences.

[0219] Replacing and configuring communication devices

[0220] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user goes through the initial setup wizard to perform basic setup, including Wi-Fi settings and entering a user ID.

[0221] Enabling the smart speaker feature

[0222] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. The server then sends the generated user profile to the device and reflects the individual settings. This allows the user to use customized smart speaker functions.

[0223] Voice commands and emotion recognition

[0224] The user inputs a voice command into the smart speaker, for example, "Tell me today's news." The voice command is recognized by the device and converted into text data.

[0225] Here, the emotion engine recognizes the user's emotion based on the voice command. For example, the emotion engine determines whether the user is excited or calm based on the tone and speed of the voice.

[0226] Response generation and customization

[0227] The server receives the speech-recognized text data and performs data analysis, including emotional data from the emotion engine. Based on the analysis results, an appropriate response is generated. For example, if the recognized emotion is joy, the server generates a response such as, "It's nice weather today, enjoy going outside."

[0228] The response data is sent to the terminal, which then runs a speech synthesis algorithm to convert the data into voice data, which is finally output from the terminal's speaker and presented to the user.

[0229] Handling User Feedback

[0230] The user can provide feedback on the smart speaker's response, for example, by saying, "Please be more specific." This feedback is also recognized by voice on the device and sent to the server.

[0231] The server has algorithms that analyze the feedback and incorporate it into future responses. The emotion engine also integrates emotional data into the user profile, which is updated based on the user's preferences and emotional state.

[0232] Specific examples

[0233] After User B installs new communication equipment and completes the initial setup, the following scenario is assumed:

[0234] 1. User B: "What's in the news today?"

[0235] 2. Terminal: Recognizes speech and converts it into text.

[0236] 3. Emotion engine: Recognizes the user's excited emotions from their voice.

[0237] 4. Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[0238] 5. Terminal: Provides the response as audio to the user.

[0239] In this way, the present invention provides a customized smart speaker experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[0240] The processing flow will be explained below.

[0241] Step 1:

[0242] The server retrieves user data, including the user's behavioral history, preferences, and personal information, from the database of the internet service, such as message history, search history, and application usage data.

[0243] Step 2:

[0244] The server analyzes the acquired user data and uses machine learning algorithms to identify the user's behavioral patterns and preferences, then generates a user profile based on that. For example, a user's cooking interests can be identified from their search history and a profile created based on that.

[0245] Step 3:

[0246] The terminal will install a new communication device (with smart speaker functionality), remove the old communication device, and connect the new device to the network.

[0247] Step 4:

[0248] The user goes through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID. For example, the user selects their home Wi-Fi network and enters the password.

[0249] Step 5:

[0250] The server receives the initial setup information sent from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[0251] Step 6:

[0252] The server then sends the generated user profile to the device and reflects the user's personal settings. For example, if the user is interested in cooking, a function to suggest related recipes is enabled.

[0253] Step 7:

[0254] Users input voice commands into the smart speaker, for example, "Tell me the weather today."

[0255] Step 8:

[0256] The device converts the voice command into text data using a speech recognition algorithm, and the voice command "Tell me the weather today" is sent to the server as text data.

[0257] Step 9:

[0258] The server analyzes the text data and generates an appropriate response, for example, by retrieving the latest weather information from a weather forecast database and generating a response such as "Today's weather is sunny."

[0259] Step 10:

[0260] The emotion engine recognizes the user's emotions based on their voice commands, determining whether they are happy or sad based on the tone and speed of their voice.

[0261] Step 11:

[0262] The server takes into account the data from the emotion engine and generates a response that matches the user's emotional state. For example, if the user is feeling down, the server generates a response like "It's a sunny and pleasant day today."

[0263] Step 12:

[0264] The server sends the generated response data to the terminal, which converts the response into voice data using a voice synthesis algorithm.

[0265] Step 13:

[0266] The terminal outputs the converted voice data from a speaker to provide it to the user.

[0267] Step 14:

[0268] Users provide feedback on the smart speaker's response, for example, by saying, "I'd like more detailed weather information."

[0269] Step 15:

[0270] The terminal recognizes the user's feedback through speech recognition and transmits it to the server as text data.

[0271] Step 16:

[0272] The server analyzes the feedback and reflects it in future responses. It also integrates the emotional data recognized by the emotion engine into the user profile and updates it, allowing it to provide services that are more tailored to the user's preferences and emotional state.

[0273] In this way, a system is realized that provides users with an optimized smart speaker experience through user data collection, analysis, profile generation, emotion recognition, and feedback processing.

[0274] Example 2

[0275] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0276] Although existing smart devices are customized based on user behavior patterns and preferences, they lack the ability to consider the user's emotional state, making them unable to fully meet user needs. In addition, there are insufficient means to effectively utilize user feedback to improve the device's responsiveness.

[0277] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0278] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart device functions based on the generated profile, means for analyzing the user's emotional state using an emotion engine, means for acquiring user feedback, analyzing the acquired feedback and reflecting it in future responses, and means for integrating the emotion data acquired from the emotion engine into the profile and updating the profile based on the user's preferences and emotional state. This enables customization that takes into account the user's emotional state in addition to their behavioral patterns and preferences, and further enables continuous improvement of the system through user feedback.

[0279] "User data" refers to data obtained from Internet services, such as a user's message history, search history, and application usage data.

[0280] A "profile" is information that indicates a user's behavioral patterns and preferences, and is generated by analyzing user data.

[0281] "Smart device features" are features of smart speakers and similar devices that are customized based on a user profile.

[0282] A "voice command" is a voice input that includes instructions or questions spoken by a user to a smart device.

[0283] "Voice recognition" is a technology that analyzes input voice commands and converts them into text data.

[0284] The "emotion engine" is a technology for analyzing the user's emotional state from the tone and speed of their voice.

[0285] An "appropriate response" is a response message to the user that is generated based on the voice command and the user's emotional state.

[0286] "Feedback" refers to information including opinions and reactions provided by a user in response to a smart device's response.

[0287] "User authentication" is the process that takes place to verify a user's identity.

[0288] The present invention relates to a system that provides customized smart device functions based on the collection and analysis of user data, and furthermore, combines an emotion engine to respond to the user's emotional state.

[0289] First, the server retrieves user data from the Internet service database. This user data includes the user's message history, search history, and application usage data. Specifically, the server uses a software module to collect this data and analyzes behavioral patterns and preferences. A profile is generated from the analyzed data. This profile is unique to each user and reflects their individual preferences and behavioral patterns.

[0290] Next, the terminal replaces its existing communication device with a new communication device, which has smart device functionality integrated into it. After installing the terminal, the user configures Wi-Fi and enters their user ID through an initial setup wizard. All initial setup information is sent to the server, where user authentication is performed. If authentication is successful, the smart device functionality is enabled, and the server sends the generated profile to the terminal, reflecting the customized settings.

[0291] Next, the user inputs a voice command into the smart device. For example, they might say, "Tell me today's news." The device receives this voice command and converts it into text data using speech recognition software. The emotion engine then works to recognize emotions from the tone and speed of the user's voice. This allows the device to understand the user's emotional state, such as whether they are excited or calm.

[0292] The server analyzes the text data and emotion data generated by the speech recognition. Based on the analysis results, an appropriate response is generated. For example, if the user is excited, the server can respond with, "It's beautiful weather today, so please enjoy some activities." This response data is then sent back to the device, which uses a speech synthesis algorithm to output the response as voice data. Finally, the response is provided to the user as voice from the device.

[0293] When a user provides feedback on a smart device's response, for example by saying, "Please be more specific," this feedback is also recognized by voice on the device and sent to the server. The server analyzes the feedback and reflects it in future responses. During this process, the emotion engine also operates, integrating the acquired emotional data into the profile. This allows the profile to be updated more precisely to reflect the user's preferences and emotional state.

[0294] As a concrete example, a scenario will be shown in which user B has installed a new communication device and has completed the initial setup.

[0295] User B: "What's in the news today?"

[0296] Device: Recognizes speech and converts it to text.

[0297] Emotion engine: Recognizes the user's excited emotions from their voice.

[0298] Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[0299] Terminal: Provides the response as audio to the user.

[0300] In this way, the present invention provides a customized smart device experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[0301] Example prompt sentence:

[0302] Design a smart device system that delivers news based on a user's message and search history. The system will have voice command and emotion recognition capabilities, and will exchange data between the server and the device. Please explain the specific processing steps and technologies used, including examples of the system.

[0303] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0304] Step 1:

[0305] The user performs the initial setup of the smart device. Specifically, the user sets up Wi-Fi on the device and enters a user ID to complete the basic environment setup. This setup information is sent from the device to the server (input: Wi-Fi settings, user ID; output: sending initial setup information to the server).

[0306] Step 2:

[0307] The server receives the initial setting information and performs user authentication. If the user authentication is successful, the server enables the smart device function (input: initial setting information, output: user authentication result and enablement of smart device function).

[0308] Step 3:

[0309] The server retrieves user data from the Internet Services database, including the user's message history, search history, and application usage data (input: user database, output: retrieved user data).

[0310] Step 4:

[0311] The server analyzes the acquired user data and generates a profile based on the user's behavioral patterns and preferences (input: acquired user data, output: generated user profile).

[0312] Step 5:

[0313] The server transmits the generated user profile to the terminal, and reflects the individual settings on the terminal (input: generated user profile, output: user profile transmitted to the terminal).

[0314] Step 6:

[0315] The user inputs a voice command into the smart device, for example, "Tell me today's news" (input: voice command, output: voice input on the device).

[0316] Step 7:

[0317] The terminal converts the voice command into text data using voice recognition software (input: voice command, output: text data).

[0318] Step 8:

[0319] The emotion engine works by analyzing the tone and speed of the voice to recognize the user's emotional state (input: text data, output: user's emotional state).

[0320] Step 9:

[0321] The server analyzes the speech-recognized text data and emotion data and generates an appropriate response. For example, if the user is excited, it generates a response containing positive content (input: text data and emotion data, output: generated response).

[0322] Step 10:

[0323] The server sends the generated response to the terminal (input: generated response, output: response sent to terminal).

[0324] Step 11:

[0325] The device uses a speech synthesis algorithm to output the response as voice data (input: transmitted response, output: voice data). The device provides the voice data to the user through a speaker (input: voice data, output: voice provided to the user).

[0326] Step 12:

[0327] The user provides feedback on the smart device's response, for example, by saying "Please be more specific" (input: feedback speech, output: speech input on device).

[0328] Step 13:

[0329] The terminal recognizes the feedback by voice and sends it to the server (input: feedback voice, output: feedback text data).

[0330] Step 14:

[0331] The server analyzes the feedback and reflects it in future responses. The emotion data acquired by the emotion engine is also integrated into the profile, which is updated based on the user's preferences and emotional state (input: feedback text data and emotion data, output: updated user profile).

[0332] (Application example 2)

[0333] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0334] Conventional smart speaker systems have difficulty providing personalized responses and information based on user preferences, and they do not take into account the user's emotional state. This results in a uniform user experience, making it impossible to provide optimal services for each individual user. It is also difficult to effectively incorporate user feedback. This has led to a need for effective means to improve customer satisfaction in commercial environments such as brick-and-mortar stores.

[0335] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0336] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart speaker functions based on the generated profile, means for a user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for performing emotion analysis, means for analyzing the converted text data and emotion data to generate an appropriate response, means for outputting the generated response as voice, means for replacing the communication device with a new one, means for acquiring initial setting information and performing user authentication, means for activating the customized smart speaker function after successful authentication, means for acquiring user feedback, and means for analyzing the acquired feedback and integrating emotion data into the user profile to update it. This enables personalized responses and information provision according to the user's preferences and emotional state, thereby improving customer satisfaction in physical stores.

[0337] "User data" is a general term for information such as a user's behavior history, preference information, and application usage data.

[0338] A "profile" is information generated based on user data that analyzes and records a user's behavioral patterns and preferences.

[0339] "Smart speaker functions" refers to all functions of smart speakers, including voice recognition, voice response, and information provision.

[0340] A "voice command" is an instruction or question that a user speaks to a smart speaker.

[0341] "Speech recognition" refers to the general process of analyzing input voice data and converting it into text data.

[0342] "Emotion analysis" is a technology that recognizes and determines a user's emotional state from the tone and speed of their voice.

[0343] An "appropriate response" is a reply or information provided to the user that is generated based on the voice command and emotion data.

[0344] "Communication equipment" refers to any hardware device used to connect to the Internet or other devices.

[0345] "Initial setting information" refers to the setting information required to use new communication devices or services.

[0346] "User authentication" refers to the overall process of verifying a user's identity and granting access permissions.

[0347] "Feedback" refers to information such as responses, opinions, and suggestions obtained from users.

[0348] The present invention relates to a smart speaker system for providing personalized information and services to customers in physical stores. The specific configuration and operation of the system are described below.

[0349] System configuration and operation

[0350] Collection and analysis of user data

[0351] The server acquires user data, which includes behavioral history, preference information, and application usage data. The server analyzes this data to generate a profile that records the user's behavioral patterns and preferences. For example, the server analyzes the user's past purchase history to identify the user's favorite product categories.

[0352] Smart speaker setup and authentication

[0353] The device replaces the existing communication device with a new one. The user enters the initial setup information, and the server authenticates the user. If authentication is successful, the customized smart speaker function is activated.

[0354] Voice Commands and Sentiment Analysis

[0355] A user inputs a voice command into a smart speaker, for example, "Tell me about new promotions." The voice command is recognized by the device and converted into text data. Emotion analysis is then performed to recognize the user's emotional state from the tone and speed of the voice.

[0356] Generating and serving the response

[0357] The server analyzes the converted text data and emotional data to generate an appropriate response. For example, if the server recognizes that the user is excited, it generates a customized response such as, "We have new products at 20% off today!" The device outputs this response as speech.

[0358] Get feedback and update your profile

[0359] The user provides feedback on the smart speaker's response, for example, by saying, "Please be more specific." The device recognizes this feedback and sends it to the server, which analyzes it and updates the user profile, including emotional data.

[0360] Hardware and software used

[0361] This system uses the following hardware and software:

[0362] Speech Recognition Library: Convert speech to text using speech_recognition.

[0363] Sentiment Analysis Library: Performs sentiment analysis using a custom emotion_recognition library.

[0364] Text-to-speech: The generated response is converted into audio using the Google Text-to-Speech (gTTS) library.

[0365] Communications Equipment: Use common internet-connected devices for user authentication and data communication.

[0366] Specific examples

[0367] For example, if a user asks a smart speaker, "What are the new promotions?", the speech is converted into text and the user's emotional state is recognized through emotion analysis. Based on the user's profile and emotion data, the server generates a response such as, "We're offering 20% ​​off new products today!", which the device then delivers to the user as audio.

[0368] Prompt Sentence Examples

[0369] User ID: user123

[0370] User Emotion: Delight

[0371] User preferences: Interested in the latest gadgets

[0372] Response: New items are being offered at 20% off today!”

[0373] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0374] Step 1:

[0375] The server obtains the user data.

[0376] The user's behavioral history, preference information, and application usage data are obtained from the database of the internet service. This data is sent to the server and used as input data for analysis.

[0377] Step 2:

[0378] The server analyzes the user data and generates a profile.

[0379] The acquired user data is analyzed to analyze behavioral patterns and preferences, and a user profile is generated. This profile records the user's preferred product and service categories. Based on the analysis results, profile data is generated and passed on to the next process.

[0380] Step 3:

[0381] The terminal replaces the existing communication device with a new communication device.

[0382] The user goes through the initial setup wizard to configure Wi-Fi and enter their user ID, and the initial setup information is entered into the device, allowing the communication device to adapt to the new environment and saving the initial setup information to the device.

[0383] Step 4:

[0384] The server obtains the initial setting information and performs user authentication.

[0385] The server authenticates the user based on the initial setting information sent from the device. If the user ID and password are verified and authentication is successful, the customized smart speaker function is activated.

[0386] Step 5:

[0387] The user enters voice commands into the smart speaker.

[0388] For example, say, "Tell me about new promotions." This voice input is captured through the device's microphone.

[0389] Step 6:

[0390] The terminal recognizes the input voice command and converts it into text data.

[0391] The speech_recognition library is used to convert the audio data into text data, which is then sent to the server.

[0392] Step 7:

[0393] The terminal performs emotion analysis.

[0394] The emotion_recognition library is used to analyze the user's emotional state from the tone and rate of speech. Emotion data is generated and sent to the server along with the text data.

[0395] Step 8:

[0396] The server analyzes the converted text data and emotion data to generate an appropriate response.

[0397] Based on this data, the server uses a generative AI model to generate a customized response. For example, if the emotion of joy is recognized, the server generates a response such as, "We have new products at 20% off today!" The generated response is sent to the device for speech synthesis.

[0398] Step 9:

[0399] The terminal outputs the generated response as voice.

[0400] The Google Text-to-Speech (gTTS) library is used to convert the text response into speech, which is then output to the speaker and presented to the user.

[0401] Step 10:

[0402] The user provides feedback on the smart speaker's response.

[0403] For example, say, "Tell me more specifically." This feedback is captured through the device's microphone.

[0404] Step 11:

[0405] The terminal recognizes the feedback by voice and transmits it to the server.

[0406] Again, we use the speech_recognition library to convert the audio feedback into text data and send it to the server.

[0407] Step 12:

[0408] The server analyzes the feedback and integrates the emotional data into the user profile to update it.

[0409] Update user profiles based on feedback and sentiment data, leading to more personalized responses and a better user experience.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0411] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0413] [Second embodiment]

[0414] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0415] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0416] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0418] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0420] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0421] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0422] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0425] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0426] The present invention relates to a system for providing customized smart speaker functionality based on the collection and analysis of user data.

[0427] Explanation of program processing

[0428] Collection and analysis of user data

[0429] The server collects user data, including user behavior history, preferences, and personal information. For example, the server analyzes the user's interests based on the user's internet search history and application usage data. From the results of this analysis, a customized profile is created for each user.

[0430] Replacing and configuring communication devices

[0431] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user performs basic settings using an initial setup wizard. The initial setup includes Wi-Fi settings and user ID entry.

[0432] Enabling the smart speaker feature

[0433] The server receives the initial setup information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. This applies the customized profile to the communication device. The user can immediately use the smart speaker function.

[0434] Processing voice commands

[0435] The user inputs a voice command into the smart speaker. For example, they might say, "Tell me what the weather will be tomorrow." This voice command is recognized by the device and converted into text data.

[0436] The server receives the text data that has been recognized by speech recognition and analyzes the command. Based on the analysis results, it generates an appropriate response. For example, it may generate a response such as "Tomorrow's weather will be sunny." This response data is then sent to the terminal and output as voice.

[0437] Handling User Feedback

[0438] Users can provide feedback on the smart speaker's response, for example by stating their opinion such as "Please be more specific." This feedback is recognized by the device and then sent to the server.

[0439] The server has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the smart speaker's responses to evolve to more closely match user expectations.

[0440] Specific examples

[0441] After User A installs new communication equipment and completes the initial setup, the following scenario is assumed:

[0442] 1. User A: "What's in the news today?"

[0443] 2. Terminal: Recognizes speech and converts it into text.

[0444] 3. Server: Analyzes the text data, collects news information, and generates appropriate responses.

[0445] 4. Response: "The major news stories today are the stock market boom and the announcement of new government policies."

[0446] 5. Terminal: Provides the response as audio to the user.

[0447] In this way, the present invention provides users with an optimized smart speaker experience and realizes a system that is easy to use and intuitive.

[0448] The processing flow will be explained below.

[0449] Step 1:

[0450] The server retrieves user data from LINE and Yahoo! databases, including user message history, search history, and app usage data.

[0451] Step 2:

[0452] The server analyzes the acquired user data and uses machine learning algorithms to identify user behavior patterns and preferences, generating a user profile from the analysis results.

[0453] Step 3:

[0454] The device is a new communication device (with smart speaker functionality) that is installed to replace the existing communication device. The new device requires basic settings to establish a network connection.

[0455] Step 4:

[0456] Users go through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID.

[0457] Step 5:

[0458] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[0459] Step 6:

[0460] The server then sends the generated user profile to the device, reflecting the individual settings, and the user is ready to use the customized smart speaker functions.

[0461] Step 7:

[0462] A user inputs a voice command to a smart speaker, for example, "Tell me the weather tomorrow."

[0463] Step 8:

[0464] The device converts the received voice commands into text data using a voice recognition algorithm, which is then sent to the server.

[0465] Step 9:

[0466] The server analyzes the text data and generates an appropriate response, for example, "Tomorrow's weather will be sunny."

[0467] Step 10:

[0468] The server sends the generated response data to the terminal, which runs a speech synthesis algorithm to convert the data into speech data.

[0469] Step 11:

[0470] The terminal outputs the audio data from a speaker to provide it to the user.

[0471] Step 12:

[0472] The user provides feedback on the smart speaker's response, such as saying, "The response was accurate" or "I wish it was shorter."

[0473] Step 13:

[0474] The device recognizes the user's voice and sends the feedback to the server, which analyzes it and uses it in future responses.

[0475] Example 1

[0476] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0477] Conventional smart devices have been unable to adequately adapt to individual users' needs and habits. It has been particularly difficult to effectively analyze user data and use smart speakers in a more personalized way. Initial setup and device replacement are also burdensome for users, making it difficult to provide responses that meet user expectations. Furthermore, there has been a lack of mechanisms for utilizing user feedback to improve responses.

[0478] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0479] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a personal profile, and means for customizing smart device functions based on the generated personal profile, thereby enabling the provision of personalized smart speaker functions according to the individual needs and habits of the user.

[0480] "User data" refers to information including a user's behavioral history, internet search history, data on applications used, personal information, and preferences.

[0481] A "personal profile" is information that indicates a user's interests, lifestyle habits, etc., generated as a result of analyzing user data.

[0482] "Smart device functionality" is the functionality of a digital device that provides voice recognition and voice response and acts based on user instructions.

[0483] A "voice instruction" is a voice command issued by a user to a smart device.

[0484] "Text information" is data in a document format that is generated by converting voice instructions using a voice recognition system.

[0485] A "response" is a reply in text or audio format that the server generates as a result of analyzing a voice command.

[0486] A "communications device" is an electronic device that allows for the transmission and reception of data.

[0487] "User authentication" is the process of identifying a particular user and verifying their permissions.

[0488] "Feedback" refers to opinions and suggestions that a user provides in response to a smart device's response.

[0489] The present invention is a system that provides customized smart device functions based on the collection and analysis of user data. This system collects data on user behavior history and applications used, and generates a customized profile for each user, thereby providing a personalized experience tailored to individual needs.

[0490] Hardware and software used

[0491] The main hardware required to implement the system includes a smart speaker, an internet-enabled device, and a server, while the software includes a voice recognition system, a data analysis engine, and a user profile generation algorithm.

[0492] Collection and analysis of user data

[0493] The server collects data about users' internet browsing history and the applications they use. This data comes from a variety of sources, including cookie information and log data. The collected data is analyzed using algorithms, such as machine learning models, to generate a personalized profile for each user. This profile reflects the user's interests, preferences, and preferences.

[0494] Replacing and initial setup of communication equipment

[0495] The terminal replaces the existing communication device with a new one. The new communication device has smart device functionality integrated into it. After the device is installed, the user uses an initial setup wizard to connect to a Wi-Fi network and enter their user ID. This configuration information is sent to the server.

[0496] Enabling the smart speaker feature

[0497] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device functions are enabled and the customized profile is applied, allowing the user to immediately start using the smart speaker.

[0498] Processing voice instructions

[0499] A user gives a voice command to a smart speaker. For example, they might say, "Tell me what the weather will be like tomorrow." This voice is recorded on the device and converted into text information in real time by a voice recognition system. This text information is sent to a server and analyzed. Based on the analysis results, an appropriate response is generated. For example, a response such as "Tomorrow's weather will be sunny" is generated, and this information is sent to the device and provided to the user as voice.

[0500] Handling User Feedback

[0501] Users can provide feedback on the smart speaker's response by voice. For example, they can say, "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server implements an algorithm that analyzes this feedback and reflects it in future responses. This allows the smart speaker's responses to better meet user expectations.

[0502] Examples of prompt statements

[0503] Examples of prompts for generative AI models include:

[0504] "Please tell me the details of the program processing for the user data collection and analysis system. Please provide a detailed explanation, including the names of the hardware and software used, and the details of data processing and calculation."

[0505] As described above, the present invention provides users with an optimized smart device experience and realizes a system that is easy to use and intuitive.

[0506] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0507] Step 1: Collect user data

[0508] The server collects user data such as internet search history, application usage data, and behavioral history. Specifically, it collects cookie information and log data. It also collects personal data that the user has authorized.

[0509] Input: Internet logs, cookie data, application usage data

[0510] Data processing: Filtering log data and removing unnecessary data

[0511] Output: The curated user dataset

[0512] Step 2: Analyze data and generate profiles

[0513] The server analyzes the collected user data, specifically using machine learning algorithms to classify the user's interests and concerns, and generates a personalized profile for each user based on the results of this analysis.

[0514] Input: Curated user dataset

[0515] Data Computing: Analysis based on machine learning models

[0516] Output: A customized personal profile

[0517] Step 3: Replace the communication device

[0518] The terminal replaces the old communication device with a new one, which has integrated smart device functionality, and a technician visits the customer to physically install the device and connect the wiring.

[0519] Input: Existing communication equipment, replacement communication equipment

[0520] Data processing: None (physical exchange work)

[0521] Output: New communications device installed

[0522] Step 4: Perform initial configuration

[0523] After installing a new communication device, the user uses the initial setup wizard to perform basic settings, such as connecting to a Wi-Fi network and entering a user ID. This setting information is then sent to the server.

[0524] Input: Wi-Fi information, user ID

[0525] Data processing: Format conversion of setting information

[0526] Output: Send initial setting information to server

[0527] Step 5: Enable the Smart Speaker feature

[0528] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device function is enabled and the customized profile is applied. The user receives a notification that "Smart speaker has been enabled."

[0529] Input: Initial setting information

[0530] Data calculation: User authentication

[0531] Output: Enable smart device features, send notifications

[0532] Step 6: Processing voice instructions

[0533] The user gives voice commands to the smart speaker. The voice is recorded on the device and converted into text information in real time by a speech recognition system. This text information is sent to the server and analyzed. Based on the analysis results, an appropriate response is generated and provided to the user as voice.

[0534] Input: Voice commands

[0535] Data processing: Text conversion using speech recognition

[0536] Output: Response based on analysis results, providing voice response

[0537] Step 7: Processing user feedback

[0538] The user provides feedback on the smart speaker's response. For example, they may state their opinion, such as "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server analyzes this feedback and improves the response algorithm. This allows future responses to more faithfully meet the user's expectations.

[0539] Input: Feedback voice

[0540] Data processing: Text conversion by voice recognition, feedback analysis

[0541] Output: Improved response algorithm

[0542] In this way, the system provides a personalized smart device experience tailored to the user's individual needs.

[0543] (Application example 1)

[0544] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0545] Conventional automotive voice assistant systems are unable to fully reflect the user's individual preferences and behavioral history, resulting in uniform information and services that make it difficult to provide an optimal experience for the user. Furthermore, adjusting the in-car environment requires a lot of manual settings, placing a heavy burden on the user. The present invention aims to solve these problems and realize a personalized in-car assistant based on the user's preferences and behavioral history.

[0546] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0547] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing the voice assistant function based on the generated profile, means for the user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for analyzing the converted text data to generate an appropriate response, means for outputting the generated response as voice, means for generating and transmitting a control signal for adjusting the environment inside the vehicle, and means for automatically adjusting the in-vehicle environment settings (temperature, lighting, music, etc.) based on the user's preferences. This enables personalized information provision based on the user's preferences and behavioral history, and automatic adjustment of the in-vehicle environment.

[0548] "User data" refers to data such as a user's behavioral history, preferences, and personal information.

[0549] A "profile" is a collection of information that indicates a user's characteristics and preferences, generated by analyzing user data.

[0550] "Voice assistant function" refers to the function of receiving voice commands, generating appropriate responses, and outputting them in voice.

[0551] "Voice command" refers to instructions or questions that a user enters by voice.

[0552] "Speech recognition" refers to the technology of converting voice data into text data.

[0553] "Text data" refers to text information generated by speech recognition.

[0554] A "response" refers to information or instructions provided to a user that are generated based on the analyzed text data.

[0555] "Output" refers to conveying the generated response to the user as voice.

[0556] "Control signal" refers to a command signal sent to control a device within a vehicle.

[0557] "Environmental settings" refers to settings such as temperature, lighting, and music inside the car.

[0558] "Preferences" refer to the user's tastes and interests.

[0559] "Initial setting information" refers to the setting information required to activate a new communication device.

[0560] "User authentication" is a procedure for verifying that a user is a legitimate user.

[0561] "Feedback" refers to evaluations and opinions regarding responses and services provided by users.

[0562] "Analysis" refers to the process of analyzing acquired data to extract useful information.

[0563] The present invention relates to a system for providing personalized information and automatically adjusting the in-vehicle environment based on the user's preferences and behavioral history in an autonomous vehicle.

[0564] The server acquires user data, including the user's behavioral history, preferences, and personal information. The data is acquired using internet search history and application usage data. By analyzing this data, a customized profile for each user is created.

[0565] The server customizes the voice assistant function based on the generated profile, enabling it to provide optimal information to the user. Additionally, when a user enters a voice command, the terminal (communication equipment inside the autonomous vehicle) uses voice recognition technology to convert the voice command into text data, using the Google Cloud Speech-to-Text API.

[0566] The server analyzes the converted text data and generates an appropriate response using the Python data analysis library scikit-learn. The generated response is then converted into audio using the Google Cloud Text-to-Speech API, which is then output to the user through the car's speakers.

[0567] Furthermore, the terminal generates and transmits control signals to adjust the vehicle's interior environment, which are transmitted via the MQTT protocol and control various devices (such as temperature control, lighting, and audio) via the vehicle's CAN bus (in-vehicle communication network).

[0568] Users can provide feedback on the provided responses by voice, which is then re-audio-recognized and sent to the server, which has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the voice assistant to continually improve its accuracy.

[0569] For example, if a user says to a smart speaker installed in the car, "Tell me today's news," the personal in-car assistant will automatically retrieve the news and respond aloud, "Today's main news stories are a sharp rise in stock prices and the announcement of new government policies."

[0570] Next, say "Make the car warmer," and the car's temperature will be adjusted to the appropriate level based on the user's profile.

[0571] Examples of prompt sentences that are important in the present invention include the following:

[0572] 1. "Collect user data and generate user profiles"

[0573] 2. "Analyzing user preferences based on internet search history"

[0574] 3. "Accept voice commands from a smart speaker and convert the voice to text."

[0575] 4. "Improve system response based on user feedback"

[0576] This invention makes it possible to provide information optimized for the user in an autonomous vehicle and automatically adjust the in-vehicle environment to be comfortable.

[0577] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0578] Step 1:

[0579] User Data Collection:

[0580] The server collects data such as user behavior, preferences, and personal information. This includes internet browsing history and application usage data. This data is obtained by using APIs to collect data from external services. It may also be obtained through sensors or other interfaces. The input for this step is internet browsing history and application usage data, and the output is raw user data for analysis.

[0581] Step 2:

[0582] User data analysis:

[0583] The server analyzes the collected user data to understand user preferences and behavioral patterns. It uses Python and data analysis libraries such as scikit-learn to perform the analysis. It preprocesses the raw data and generates profiles using techniques such as feature extraction, clustering, and classification. The input for this step is raw user data, and the output is a profile customized for each user.

[0584] Step 3:

[0585] Creating profile-based customizations:

[0586] Based on the generated user profile, the server customizes the voice assistant's functionality by specifying the type of music, news, and information the user prefers, and also applies algorithms to provide appropriate feedback and adjustments. The input for this step is the generated profile, and the output is customized voice assistant settings.

[0587] Step 4:

[0588] Receiving and recognizing voice commands:

[0589] When a user inputs a voice command into the smart speaker in the car, the device captures the speech and converts it into text data using the Google Cloud Speech-to-Text API. The input of this step is the user's voice command, and the output is text data.

[0590] Step 5:

[0591] Voice command analysis:

[0592] The server analyzes the speech-recognized text data and generates an appropriate response to the user's request. The analysis uses natural language processing and rule-based algorithms. The server also takes into account the user's profile to provide the most appropriate information and services. The input for this step is text data, and the output is response data.

[0593] Step 6:

[0594] Response transcription and output:

[0595] The server converts the generated response data into speech using the Google Cloud Text-to-Speech API. The device then outputs the converted speech to the user through the car speaker. The input of this step is the response data, and the output is a speech response.

[0596] Step 7:

[0597] Control signal generation and transmission:

[0598] The server generates control signals to adjust the in-car environment based on the user's preference profile and sends them to the terminal via the MQTT protocol. The terminal receives these control signals and controls each device (temperature, lighting, audio, etc.) via the vehicle's CAN bus. The input of this step is the user profile and control request, and the output is the control signal sent to each device.

[0599] Step 8:

[0600] Collecting and analyzing user feedback:

[0601] Users can provide feedback on the voice assistant's responses, which are then re-audio-recognized and sent to the server. The server then has an algorithm that analyzes the feedback and incorporates it into future responses, improving the accuracy of the voice assistant. The input for this step is the user's voice feedback, and the output is the analysis results and the system adjustments based on them.

[0602] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0603] The present invention relates to a system that provides customized smart speaker functions based on the collection and analysis of user data, and further combines an emotion engine to respond to the user's emotional state.

[0604] Explanation of program processing

[0605] Collection and analysis of user data

[0606] The server retrieves user data from the internet service database, including the user's message history, search history, and app usage data, and uses this data to analyze the user's behavioral patterns and preferences and create a profile.

[0607] As a specific example, the server analyzes the news articles that the user has viewed and the product data that the user has purchased, and can provide news and product information through the speaker according to the user's preferences.

[0608] Replacing and configuring communication devices

[0609] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user goes through the initial setup wizard to perform basic setup, including Wi-Fi settings and entering a user ID.

[0610] Enabling the smart speaker feature

[0611] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. The server then sends the generated user profile to the device and reflects the individual settings. This allows the user to use customized smart speaker functions.

[0612] Voice commands and emotion recognition

[0613] The user inputs a voice command into the smart speaker, for example, "Tell me today's news." The voice command is recognized by the device and converted into text data.

[0614] Here, the emotion engine recognizes the user's emotion based on the voice command. For example, the emotion engine determines whether the user is excited or calm based on the tone and speed of the voice.

[0615] Response generation and customization

[0616] The server receives the speech-recognized text data and performs data analysis, including emotional data from the emotion engine. Based on the analysis results, an appropriate response is generated. For example, if the recognized emotion is joy, the server generates a response such as, "It's nice weather today, enjoy going outside."

[0617] The response data is sent to the terminal, which then runs a speech synthesis algorithm to convert the data into voice data, which is finally output from the terminal's speaker and presented to the user.

[0618] Handling User Feedback

[0619] The user can provide feedback on the smart speaker's response, for example, by saying, "Please be more specific." This feedback is also recognized by voice on the device and sent to the server.

[0620] The server has algorithms that analyze the feedback and incorporate it into future responses. The emotion engine also integrates emotional data into the user profile, which is updated based on the user's preferences and emotional state.

[0621] Specific examples

[0622] After User B installs new communication equipment and completes the initial setup, the following scenario is assumed:

[0623] 1. User B: "What's in the news today?"

[0624] 2. Terminal: Recognizes speech and converts it into text.

[0625] 3. Emotion engine: Recognizes the user's excited emotions from their voice.

[0626] 4. Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[0627] 5. Terminal: Provides the response as audio to the user.

[0628] In this way, the present invention provides a customized smart speaker experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[0629] The processing flow will be explained below.

[0630] Step 1:

[0631] The server retrieves user data, including the user's behavioral history, preferences, and personal information, from the database of the internet service, such as message history, search history, and application usage data.

[0632] Step 2:

[0633] The server analyzes the acquired user data and uses machine learning algorithms to identify the user's behavioral patterns and preferences, then generates a user profile based on that. For example, a user's cooking interests can be identified from their search history and a profile created based on that.

[0634] Step 3:

[0635] The terminal will install a new communication device (with smart speaker functionality), remove the old communication device, and connect the new device to the network.

[0636] Step 4:

[0637] The user goes through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID. For example, the user selects their home Wi-Fi network and enters the password.

[0638] Step 5:

[0639] The server receives the initial setup information sent from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[0640] Step 6:

[0641] The server then sends the generated user profile to the device and reflects the user's personal settings. For example, if the user is interested in cooking, a function to suggest related recipes is enabled.

[0642] Step 7:

[0643] Users input voice commands into the smart speaker, for example, "Tell me the weather today."

[0644] Step 8:

[0645] The device converts the voice command into text data using a speech recognition algorithm, and the voice command "Tell me the weather today" is sent to the server as text data.

[0646] Step 9:

[0647] The server analyzes the text data and generates an appropriate response, for example, by retrieving the latest weather information from a weather forecast database and generating a response such as "Today's weather is sunny."

[0648] Step 10:

[0649] The emotion engine recognizes the user's emotions based on their voice commands, determining whether they are happy or sad based on the tone and speed of their voice.

[0650] Step 11:

[0651] The server takes into account the data from the emotion engine and generates a response that matches the user's emotional state. For example, if the user is feeling down, the server generates a response like "It's a sunny and pleasant day today."

[0652] Step 12:

[0653] The server sends the generated response data to the terminal, which converts the response into voice data using a voice synthesis algorithm.

[0654] Step 13:

[0655] The terminal outputs the converted voice data from a speaker to provide it to the user.

[0656] Step 14:

[0657] Users provide feedback on the smart speaker's response, for example, by saying, "I'd like more detailed weather information."

[0658] Step 15:

[0659] The terminal recognizes the user's feedback through speech recognition and transmits it to the server as text data.

[0660] Step 16:

[0661] The server analyzes the feedback and reflects it in future responses. It also integrates the emotional data recognized by the emotion engine into the user profile and updates it, allowing it to provide services that are more tailored to the user's preferences and emotional state.

[0662] In this way, a system is realized that provides users with an optimized smart speaker experience through user data collection, analysis, profile generation, emotion recognition, and feedback processing.

[0663] Example 2

[0664] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0665] Although existing smart devices are customized based on user behavior patterns and preferences, they lack the ability to consider the user's emotional state, making them unable to fully meet user needs. In addition, there are insufficient means to effectively utilize user feedback to improve the device's responsiveness.

[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0667] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart device functions based on the generated profile, means for analyzing the user's emotional state using an emotion engine, means for acquiring user feedback, analyzing the acquired feedback and reflecting it in future responses, and means for integrating the emotion data acquired from the emotion engine into the profile and updating the profile based on the user's preferences and emotional state. This enables customization that takes into account the user's emotional state in addition to their behavioral patterns and preferences, and further enables continuous improvement of the system through user feedback.

[0668] "User data" refers to data obtained from Internet services, such as a user's message history, search history, and application usage data.

[0669] A "profile" is information that indicates a user's behavioral patterns and preferences, and is generated by analyzing user data.

[0670] "Smart device features" are features of smart speakers and similar devices that are customized based on a user profile.

[0671] A "voice command" is a voice input that includes instructions or questions spoken by a user to a smart device.

[0672] "Voice recognition" is a technology that analyzes input voice commands and converts them into text data.

[0673] The "emotion engine" is a technology for analyzing the user's emotional state from the tone and speed of their voice.

[0674] An "appropriate response" is a response message to the user that is generated based on the voice command and the user's emotional state.

[0675] "Feedback" refers to information including opinions and reactions provided by a user in response to a smart device's response.

[0676] "User authentication" is the process that takes place to verify a user's identity.

[0677] The present invention relates to a system that provides customized smart device functions based on the collection and analysis of user data, and furthermore, combines an emotion engine to respond to the user's emotional state.

[0678] First, the server retrieves user data from the Internet service database. This user data includes the user's message history, search history, and application usage data. Specifically, the server uses a software module to collect this data and analyzes behavioral patterns and preferences. A profile is generated from the analyzed data. This profile is unique to each user and reflects their individual preferences and behavioral patterns.

[0679] Next, the terminal replaces its existing communication device with a new communication device, which has smart device functionality integrated into it. After installing the terminal, the user configures Wi-Fi and enters their user ID through an initial setup wizard. All initial setup information is sent to the server, where user authentication is performed. If authentication is successful, the smart device functionality is enabled, and the server sends the generated profile to the terminal, reflecting the customized settings.

[0680] Next, the user inputs a voice command into the smart device. For example, they might say, "Tell me today's news." The device receives this voice command and converts it into text data using speech recognition software. The emotion engine then works to recognize emotions from the tone and speed of the user's voice. This allows the device to understand the user's emotional state, such as whether they are excited or calm.

[0681] The server analyzes the text data and emotion data generated by the speech recognition. Based on the analysis results, an appropriate response is generated. For example, if the user is excited, the server can respond with, "It's beautiful weather today, so please enjoy some activities." This response data is then sent back to the device, which uses a speech synthesis algorithm to output the response as voice data. Finally, the response is provided to the user as voice from the device.

[0682] When a user provides feedback on a smart device's response, for example by saying, "Please be more specific," this feedback is also recognized by voice on the device and sent to the server. The server analyzes the feedback and reflects it in future responses. During this process, the emotion engine also operates, integrating the acquired emotional data into the profile. This allows the profile to be updated more precisely to reflect the user's preferences and emotional state.

[0683] As a concrete example, a scenario will be shown in which user B has installed a new communication device and has completed the initial setup.

[0684] User B: "What's in the news today?"

[0685] Device: Recognizes speech and converts it to text.

[0686] Emotion engine: Recognizes the user's excited emotions from their voice.

[0687] Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[0688] Terminal: Provides the response as audio to the user.

[0689] In this way, the present invention provides a customized smart device experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[0690] Example prompt sentence:

[0691] Design a smart device system that delivers news based on a user's message and search history. The system will have voice command and emotion recognition capabilities, and will exchange data between the server and the device. Please explain the specific processing steps and technologies used, including examples of the system.

[0692] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0693] Step 1:

[0694] The user performs the initial setup of the smart device. Specifically, the user sets up Wi-Fi on the device and enters a user ID to complete the basic environment setup. This setup information is sent from the device to the server (input: Wi-Fi settings, user ID; output: sending initial setup information to the server).

[0695] Step 2:

[0696] The server receives the initial setting information and performs user authentication. If the user authentication is successful, the server enables the smart device function (input: initial setting information, output: user authentication result and enablement of smart device function).

[0697] Step 3:

[0698] The server retrieves user data from the Internet Services database, including the user's message history, search history, and application usage data (input: user database, output: retrieved user data).

[0699] Step 4:

[0700] The server analyzes the acquired user data and generates a profile based on the user's behavioral patterns and preferences (input: acquired user data, output: generated user profile).

[0701] Step 5:

[0702] The server transmits the generated user profile to the terminal, and reflects the individual settings on the terminal (input: generated user profile, output: user profile transmitted to the terminal).

[0703] Step 6:

[0704] The user inputs a voice command into the smart device, for example, "Tell me today's news" (input: voice command, output: voice input on the device).

[0705] Step 7:

[0706] The terminal converts the voice command into text data using voice recognition software (input: voice command, output: text data).

[0707] Step 8:

[0708] The emotion engine works by analyzing the tone and speed of the voice to recognize the user's emotional state (input: text data, output: user's emotional state).

[0709] Step 9:

[0710] The server analyzes the speech-recognized text data and emotion data and generates an appropriate response. For example, if the user is excited, it generates a response containing positive content (input: text data and emotion data, output: generated response).

[0711] Step 10:

[0712] The server sends the generated response to the terminal (input: generated response, output: response sent to terminal).

[0713] Step 11:

[0714] The device uses a speech synthesis algorithm to output the response as voice data (input: transmitted response, output: voice data). The device provides the voice data to the user through a speaker (input: voice data, output: voice provided to the user).

[0715] Step 12:

[0716] The user provides feedback on the smart device's response, for example, by saying "Please be more specific" (input: feedback speech, output: speech input on device).

[0717] Step 13:

[0718] The terminal recognizes the feedback by voice and sends it to the server (input: feedback voice, output: feedback text data).

[0719] Step 14:

[0720] The server analyzes the feedback and reflects it in future responses. The emotion data acquired by the emotion engine is also integrated into the profile, which is updated based on the user's preferences and emotional state (input: feedback text data and emotion data, output: updated user profile).

[0721] (Application example 2)

[0722] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0723] Conventional smart speaker systems have difficulty providing personalized responses and information based on user preferences, and they do not take into account the user's emotional state. This results in a uniform user experience, making it impossible to provide optimal services for each individual user. It is also difficult to effectively incorporate user feedback. This has led to a need for effective means to improve customer satisfaction in commercial environments such as brick-and-mortar stores.

[0724] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0725] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart speaker functions based on the generated profile, means for a user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for performing emotion analysis, means for analyzing the converted text data and emotion data to generate an appropriate response, means for outputting the generated response as voice, means for replacing the communication device with a new one, means for acquiring initial setting information and performing user authentication, means for activating the customized smart speaker function after successful authentication, means for acquiring user feedback, and means for analyzing the acquired feedback and integrating emotion data into the user profile to update it. This enables personalized responses and information provision according to the user's preferences and emotional state, thereby improving customer satisfaction in physical stores.

[0726] "User data" is a general term for information such as a user's behavior history, preference information, and application usage data.

[0727] A "profile" is information generated based on user data that analyzes and records a user's behavioral patterns and preferences.

[0728] "Smart speaker functions" refers to all functions of smart speakers, including voice recognition, voice response, and information provision.

[0729] A "voice command" is an instruction or question that a user speaks to a smart speaker.

[0730] "Speech recognition" refers to the general process of analyzing input voice data and converting it into text data.

[0731] "Emotion analysis" is a technology that recognizes and determines a user's emotional state from the tone and speed of their voice.

[0732] An "appropriate response" is a reply or information provided to the user that is generated based on the voice command and emotion data.

[0733] "Communication equipment" refers to any hardware device used to connect to the Internet or other devices.

[0734] "Initial setting information" refers to the setting information required to use new communication devices or services.

[0735] "User authentication" refers to the overall process of verifying a user's identity and granting access permissions.

[0736] "Feedback" refers to information such as responses, opinions, and suggestions obtained from users.

[0737] The present invention relates to a smart speaker system for providing personalized information and services to customers in physical stores. The specific configuration and operation of the system are described below.

[0738] System configuration and operation

[0739] Collection and analysis of user data

[0740] The server acquires user data, which includes behavioral history, preference information, and application usage data. The server analyzes this data to generate a profile that records the user's behavioral patterns and preferences. For example, the server analyzes the user's past purchase history to identify the user's favorite product categories.

[0741] Smart speaker setup and authentication

[0742] The device replaces the existing communication device with a new one. The user enters the initial setup information, and the server authenticates the user. If authentication is successful, the customized smart speaker function is activated.

[0743] Voice Commands and Sentiment Analysis

[0744] A user inputs a voice command into a smart speaker, for example, "Tell me about new promotions." The voice command is recognized by the device and converted into text data. Emotion analysis is then performed to recognize the user's emotional state from the tone and speed of the voice.

[0745] Generating and serving the response

[0746] The server analyzes the converted text data and emotional data to generate an appropriate response. For example, if the server recognizes that the user is excited, it generates a customized response such as, "We have new products at 20% off today!" The device outputs this response as speech.

[0747] Get feedback and update your profile

[0748] The user provides feedback on the smart speaker's response, for example, by saying, "Please be more specific." The device recognizes this feedback and sends it to the server, which analyzes it and updates the user profile, including emotional data.

[0749] Hardware and software used

[0750] This system uses the following hardware and software:

[0751] Speech Recognition Library: Convert speech to text using speech_recognition.

[0752] Sentiment Analysis Library: Performs sentiment analysis using a custom emotion_recognition library.

[0753] Text-to-speech: The generated response is converted into audio using the Google Text-to-Speech (gTTS) library.

[0754] Communications Equipment: Use common internet-connected devices for user authentication and data communication.

[0755] Specific examples

[0756] For example, if a user asks a smart speaker, "What are the new promotions?", the speech is converted into text and the user's emotional state is recognized through emotion analysis. Based on the user's profile and emotion data, the server generates a response such as, "We're offering 20% ​​off new products today!", which the device then delivers to the user as audio.

[0757] Prompt Sentence Examples

[0758] User ID: user123

[0759] User Emotion: Delight

[0760] User preferences: Interested in the latest gadgets

[0761] Response: New items are being offered at 20% off today!”

[0762] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0763] Step 1:

[0764] The server obtains the user data.

[0765] The user's behavioral history, preference information, and application usage data are obtained from the database of the internet service. This data is sent to the server and used as input data for analysis.

[0766] Step 2:

[0767] The server analyzes the user data and generates a profile.

[0768] The acquired user data is analyzed to analyze behavioral patterns and preferences, and a user profile is generated. This profile records the user's preferred product and service categories. Based on the analysis results, profile data is generated and passed on to the next process.

[0769] Step 3:

[0770] The terminal replaces the existing communication device with a new communication device.

[0771] The user goes through the initial setup wizard to configure Wi-Fi and enter their user ID, and the initial setup information is entered into the device, allowing the communication device to adapt to the new environment and saving the initial setup information to the device.

[0772] Step 4:

[0773] The server obtains the initial setting information and performs user authentication.

[0774] The server authenticates the user based on the initial setting information sent from the device. If the user ID and password are verified and authentication is successful, the customized smart speaker function is activated.

[0775] Step 5:

[0776] The user enters voice commands into the smart speaker.

[0777] For example, say, "Tell me about new promotions." This voice input is captured through the device's microphone.

[0778] Step 6:

[0779] The terminal recognizes the input voice command and converts it into text data.

[0780] The speech_recognition library is used to convert the audio data into text data, which is then sent to the server.

[0781] Step 7:

[0782] The terminal performs emotion analysis.

[0783] The emotion_recognition library is used to analyze the user's emotional state from the tone and rate of speech. Emotion data is generated and sent to the server along with the text data.

[0784] Step 8:

[0785] The server analyzes the converted text data and emotion data to generate an appropriate response.

[0786] Based on this data, the server uses a generative AI model to generate a customized response. For example, if the emotion of joy is recognized, the server generates a response such as, "We have new products at 20% off today!" The generated response is sent to the device for speech synthesis.

[0787] Step 9:

[0788] The terminal outputs the generated response as voice.

[0789] The Google Text-to-Speech (gTTS) library is used to convert the text response into speech, which is then output to the speaker and presented to the user.

[0790] Step 10:

[0791] The user provides feedback on the smart speaker's response.

[0792] For example, say, "Tell me more specifically." This feedback is captured through the device's microphone.

[0793] Step 11:

[0794] The terminal recognizes the feedback by voice and transmits it to the server.

[0795] Again, we use the speech_recognition library to convert the audio feedback into text data and send it to the server.

[0796] Step 12:

[0797] The server analyzes the feedback and integrates the emotional data into the user profile to update it.

[0798] Update user profiles based on feedback and sentiment data, leading to more personalized responses and a better user experience.

[0799] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0800] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0801] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0802] [Third embodiment]

[0803] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0804] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0805] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0806] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0807] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0808] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0809] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0810] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0811] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0812] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0813] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0814] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0815] The present invention relates to a system for providing customized smart speaker functionality based on the collection and analysis of user data.

[0816] Explanation of program processing

[0817] Collection and analysis of user data

[0818] The server collects user data, including user behavior history, preferences, and personal information. For example, the server analyzes the user's interests based on the user's internet search history and application usage data. From the results of this analysis, a customized profile is created for each user.

[0819] Replacing and configuring communication devices

[0820] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user performs basic settings using an initial setup wizard. The initial setup includes Wi-Fi settings and user ID entry.

[0821] Enabling the smart speaker feature

[0822] The server receives the initial setup information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. This applies the customized profile to the communication device. The user can immediately use the smart speaker function.

[0823] Processing voice commands

[0824] The user inputs a voice command into the smart speaker. For example, they might say, "Tell me what the weather will be tomorrow." This voice command is recognized by the device and converted into text data.

[0825] The server receives the text data that has been recognized by speech recognition and analyzes the command. Based on the analysis results, it generates an appropriate response. For example, it may generate a response such as "Tomorrow's weather will be sunny." This response data is then sent to the terminal and output as voice.

[0826] Handling User Feedback

[0827] Users can provide feedback on the smart speaker's response, for example by stating their opinion such as "Please be more specific." This feedback is recognized by the device and then sent to the server.

[0828] The server has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the smart speaker's responses to evolve to more closely match user expectations.

[0829] Specific examples

[0830] After User A installs new communication equipment and completes the initial setup, the following scenario is assumed:

[0831] 1. User A: "What's in the news today?"

[0832] 2. Terminal: Recognizes speech and converts it into text.

[0833] 3. Server: Analyzes the text data, collects news information, and generates appropriate responses.

[0834] 4. Response: "The major news stories today are the stock market boom and the announcement of new government policies."

[0835] 5. Terminal: Provides the response as audio to the user.

[0836] In this way, the present invention provides users with an optimized smart speaker experience and realizes a system that is easy to use and intuitive.

[0837] The processing flow will be explained below.

[0838] Step 1:

[0839] The server retrieves user data from LINE and Yahoo! databases, including user message history, search history, and app usage data.

[0840] Step 2:

[0841] The server analyzes the acquired user data and uses machine learning algorithms to identify user behavior patterns and preferences, generating a user profile from the analysis results.

[0842] Step 3:

[0843] The device is a new communication device (with smart speaker functionality) that is installed to replace the existing communication device. The new device requires basic settings to establish a network connection.

[0844] Step 4:

[0845] Users go through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID.

[0846] Step 5:

[0847] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[0848] Step 6:

[0849] The server then sends the generated user profile to the device, reflecting the individual settings, and the user is ready to use the customized smart speaker functions.

[0850] Step 7:

[0851] A user inputs a voice command to a smart speaker, for example, "Tell me the weather tomorrow."

[0852] Step 8:

[0853] The device converts the received voice commands into text data using a voice recognition algorithm, which is then sent to the server.

[0854] Step 9:

[0855] The server analyzes the text data and generates an appropriate response, for example, "Tomorrow's weather will be sunny."

[0856] Step 10:

[0857] The server sends the generated response data to the terminal, which runs a speech synthesis algorithm to convert the data into speech data.

[0858] Step 11:

[0859] The terminal outputs the audio data from a speaker to provide it to the user.

[0860] Step 12:

[0861] The user provides feedback on the smart speaker's response, such as saying, "The response was accurate" or "I wish it was shorter."

[0862] Step 13:

[0863] The device recognizes the user's voice and sends the feedback to the server, which analyzes it and uses it in future responses.

[0864] Example 1

[0865] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0866] Conventional smart devices have been unable to adequately adapt to individual users' needs and habits. It has been particularly difficult to effectively analyze user data and use smart speakers in a more personalized way. Initial setup and device replacement are also burdensome for users, making it difficult to provide responses that meet user expectations. Furthermore, there has been a lack of mechanisms for utilizing user feedback to improve responses.

[0867] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0868] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a personal profile, and means for customizing smart device functions based on the generated personal profile, thereby enabling the provision of personalized smart speaker functions according to the individual needs and habits of the user.

[0869] "User data" refers to information including a user's behavioral history, internet search history, data on applications used, personal information, and preferences.

[0870] A "personal profile" is information that indicates a user's interests, lifestyle habits, etc., generated as a result of analyzing user data.

[0871] "Smart device functionality" is the functionality of a digital device that provides voice recognition and voice response and acts based on user instructions.

[0872] A "voice instruction" is a voice command issued by a user to a smart device.

[0873] "Text information" is data in a document format that is generated by converting voice instructions using a voice recognition system.

[0874] A "response" is a reply in text or audio format that the server generates as a result of analyzing a voice command.

[0875] A "communications device" is an electronic device that allows for the transmission and reception of data.

[0876] "User authentication" is the process of identifying a particular user and verifying their permissions.

[0877] "Feedback" refers to opinions and suggestions that a user provides in response to a smart device's response.

[0878] The present invention is a system that provides customized smart device functions based on the collection and analysis of user data. This system collects data on user behavior history and applications used, and generates a customized profile for each user, thereby providing a personalized experience tailored to individual needs.

[0879] Hardware and software used

[0880] The main hardware required to implement the system includes a smart speaker, an internet-enabled device, and a server, while the software includes a voice recognition system, a data analysis engine, and a user profile generation algorithm.

[0881] Collection and analysis of user data

[0882] The server collects data about users' internet browsing history and the applications they use. This data comes from a variety of sources, including cookie information and log data. The collected data is analyzed using algorithms, such as machine learning models, to generate a personalized profile for each user. This profile reflects the user's interests, preferences, and preferences.

[0883] Replacing and initial setup of communication equipment

[0884] The terminal replaces the existing communication device with a new one. The new communication device has smart device functionality integrated into it. After the device is installed, the user uses an initial setup wizard to connect to a Wi-Fi network and enter their user ID. This configuration information is sent to the server.

[0885] Enabling the smart speaker feature

[0886] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device functions are enabled and the customized profile is applied, allowing the user to immediately start using the smart speaker.

[0887] Processing voice instructions

[0888] A user gives a voice command to a smart speaker. For example, they might say, "Tell me what the weather will be like tomorrow." This voice is recorded on the device and converted into text information in real time by a voice recognition system. This text information is sent to a server and analyzed. Based on the analysis results, an appropriate response is generated. For example, a response such as "Tomorrow's weather will be sunny" is generated, and this information is sent to the device and provided to the user as voice.

[0889] Handling User Feedback

[0890] Users can provide feedback on the smart speaker's response by voice. For example, they can say, "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server implements an algorithm that analyzes this feedback and reflects it in future responses. This allows the smart speaker's responses to better meet user expectations.

[0891] Examples of prompt statements

[0892] Examples of prompts for generative AI models include:

[0893] "Please tell me the details of the program processing for the user data collection and analysis system. Please provide a detailed explanation, including the names of the hardware and software used, and the details of data processing and calculation."

[0894] As described above, the present invention provides users with an optimized smart device experience and realizes a system that is easy to use and intuitive.

[0895] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0896] Step 1: Collect user data

[0897] The server collects user data such as internet search history, application usage data, and behavioral history. Specifically, it collects cookie information and log data. It also collects personal data that the user has authorized.

[0898] Input: Internet logs, cookie data, application usage data

[0899] Data processing: Filtering log data and removing unnecessary data

[0900] Output: The curated user dataset

[0901] Step 2: Analyze data and generate profiles

[0902] The server analyzes the collected user data, specifically using machine learning algorithms to classify the user's interests and concerns, and generates a personalized profile for each user based on the results of this analysis.

[0903] Input: Curated user dataset

[0904] Data Computing: Analysis based on machine learning models

[0905] Output: A customized personal profile

[0906] Step 3: Replace the communication device

[0907] The terminal replaces the old communication device with a new one, which has integrated smart device functionality, and a technician visits the customer to physically install the device and connect the wiring.

[0908] Input: Existing communication equipment, replacement communication equipment

[0909] Data processing: None (physical exchange work)

[0910] Output: New communications device installed

[0911] Step 4: Perform initial configuration

[0912] After installing a new communication device, the user uses the initial setup wizard to perform basic settings, such as connecting to a Wi-Fi network and entering a user ID. This setting information is then sent to the server.

[0913] Input: Wi-Fi information, user ID

[0914] Data processing: Format conversion of setting information

[0915] Output: Send initial setting information to server

[0916] Step 5: Enable the Smart Speaker feature

[0917] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device function is enabled and the customized profile is applied. The user receives a notification that "Smart speaker has been enabled."

[0918] Input: Initial setting information

[0919] Data calculation: User authentication

[0920] Output: Enable smart device features, send notifications

[0921] Step 6: Processing voice instructions

[0922] The user gives voice commands to the smart speaker. The voice is recorded on the device and converted into text information in real time by a speech recognition system. This text information is sent to the server and analyzed. Based on the analysis results, an appropriate response is generated and provided to the user as voice.

[0923] Input: Voice commands

[0924] Data processing: Text conversion using speech recognition

[0925] Output: Response based on analysis results, providing voice response

[0926] Step 7: Processing user feedback

[0927] The user provides feedback on the smart speaker's response. For example, they may state their opinion, such as "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server analyzes this feedback and improves the response algorithm. This allows future responses to more faithfully meet the user's expectations.

[0928] Input: Feedback voice

[0929] Data processing: Text conversion by voice recognition, feedback analysis

[0930] Output: Improved response algorithm

[0931] In this way, the system provides a personalized smart device experience tailored to the user's individual needs.

[0932] (Application example 1)

[0933] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0934] Conventional automotive voice assistant systems are unable to fully reflect the user's individual preferences and behavioral history, resulting in uniform information and services that make it difficult to provide an optimal experience for the user. Furthermore, adjusting the in-car environment requires a lot of manual settings, placing a heavy burden on the user. The present invention aims to solve these problems and realize a personalized in-car assistant based on the user's preferences and behavioral history.

[0935] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0936] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing the voice assistant function based on the generated profile, means for the user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for analyzing the converted text data to generate an appropriate response, means for outputting the generated response as voice, means for generating and transmitting a control signal for adjusting the environment inside the vehicle, and means for automatically adjusting the in-vehicle environment settings (temperature, lighting, music, etc.) based on the user's preferences. This enables personalized information provision based on the user's preferences and behavioral history, and automatic adjustment of the in-vehicle environment.

[0937] "User data" refers to data such as a user's behavioral history, preferences, and personal information.

[0938] A "profile" is a collection of information that indicates a user's characteristics and preferences, generated by analyzing user data.

[0939] "Voice assistant function" refers to the function of receiving voice commands, generating appropriate responses, and outputting them in voice.

[0940] "Voice command" refers to instructions or questions that a user enters by voice.

[0941] "Speech recognition" refers to the technology of converting voice data into text data.

[0942] "Text data" refers to text information generated by speech recognition.

[0943] A "response" refers to information or instructions provided to a user that are generated based on the analyzed text data.

[0944] "Output" refers to conveying the generated response to the user as voice.

[0945] "Control signal" refers to a command signal sent to control a device within a vehicle.

[0946] "Environmental settings" refers to settings such as temperature, lighting, and music inside the car.

[0947] "Preferences" refer to the user's tastes and interests.

[0948] "Initial setting information" refers to the setting information required to activate a new communication device.

[0949] "User authentication" is a procedure for verifying that a user is a legitimate user.

[0950] "Feedback" refers to evaluations and opinions regarding responses and services provided by users.

[0951] "Analysis" refers to the process of analyzing acquired data to extract useful information.

[0952] The present invention relates to a system for providing personalized information and automatically adjusting the in-vehicle environment based on the user's preferences and behavioral history in an autonomous vehicle.

[0953] The server acquires user data, including the user's behavioral history, preferences, and personal information. The data is acquired using internet search history and application usage data. By analyzing this data, a customized profile for each user is created.

[0954] The server customizes the voice assistant function based on the generated profile, enabling it to provide optimal information to the user. Additionally, when a user enters a voice command, the terminal (communication equipment inside the autonomous vehicle) uses voice recognition technology to convert the voice command into text data, using the Google Cloud Speech-to-Text API.

[0955] The server analyzes the converted text data and generates an appropriate response using the Python data analysis library scikit-learn. The generated response is then converted into audio using the Google Cloud Text-to-Speech API, which is then output to the user through the car's speakers.

[0956] Furthermore, the terminal generates and transmits control signals to adjust the vehicle's interior environment, which are transmitted via the MQTT protocol and control various devices (such as temperature control, lighting, and audio) via the vehicle's CAN bus (in-vehicle communication network).

[0957] Users can provide feedback on the provided responses by voice, which is then re-audio-recognized and sent to the server, which has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the voice assistant to continually improve its accuracy.

[0958] For example, if a user says to a smart speaker installed in the car, "Tell me today's news," the personal in-car assistant will automatically retrieve the news and respond aloud, "Today's main news stories are a sharp rise in stock prices and the announcement of new government policies."

[0959] Next, say "Make the car warmer," and the car's temperature will be adjusted to the appropriate level based on the user's profile.

[0960] Examples of prompt sentences that are important in the present invention include the following:

[0961] 1. "Collect user data and generate user profiles"

[0962] 2. "Analyzing user preferences based on internet search history"

[0963] 3. "Accept voice commands from a smart speaker and convert the voice to text."

[0964] 4. "Improve system response based on user feedback"

[0965] This invention makes it possible to provide information optimized for the user in an autonomous vehicle and automatically adjust the in-vehicle environment to be comfortable.

[0966] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0967] Step 1:

[0968] User Data Collection:

[0969] The server collects data such as user behavior, preferences, and personal information. This includes internet browsing history and application usage data. This data is obtained by using APIs to collect data from external services. It may also be obtained through sensors or other interfaces. The input for this step is internet browsing history and application usage data, and the output is raw user data for analysis.

[0970] Step 2:

[0971] User data analysis:

[0972] The server analyzes the collected user data to understand user preferences and behavioral patterns. It uses Python and data analysis libraries such as scikit-learn to perform the analysis. It preprocesses the raw data and generates profiles using techniques such as feature extraction, clustering, and classification. The input for this step is raw user data, and the output is a profile customized for each user.

[0973] Step 3:

[0974] Creating profile-based customizations:

[0975] Based on the generated user profile, the server customizes the voice assistant's functionality by specifying the type of music, news, and information the user prefers, and also applies algorithms to provide appropriate feedback and adjustments. The input for this step is the generated profile, and the output is customized voice assistant settings.

[0976] Step 4:

[0977] Receiving and recognizing voice commands:

[0978] When a user inputs a voice command into the smart speaker in the car, the device captures the speech and converts it into text data using the Google Cloud Speech-to-Text API. The input of this step is the user's voice command, and the output is text data.

[0979] Step 5:

[0980] Voice command analysis:

[0981] The server analyzes the speech-recognized text data and generates an appropriate response to the user's request. The analysis uses natural language processing and rule-based algorithms. The server also takes into account the user's profile to provide the most appropriate information and services. The input for this step is text data, and the output is response data.

[0982] Step 6:

[0983] Response transcription and output:

[0984] The server converts the generated response data into speech using the Google Cloud Text-to-Speech API. The device then outputs the converted speech to the user through the car speaker. The input of this step is the response data, and the output is a speech response.

[0985] Step 7:

[0986] Control signal generation and transmission:

[0987] The server generates control signals to adjust the in-car environment based on the user's preference profile and sends them to the terminal via the MQTT protocol. The terminal receives these control signals and controls each device (temperature, lighting, audio, etc.) via the vehicle's CAN bus. The input of this step is the user profile and control request, and the output is the control signal sent to each device.

[0988] Step 8:

[0989] Collecting and analyzing user feedback:

[0990] Users can provide feedback on the voice assistant's responses, which are then re-audio-recognized and sent to the server. The server then has an algorithm that analyzes the feedback and incorporates it into future responses, improving the accuracy of the voice assistant. The input for this step is the user's voice feedback, and the output is the analysis results and the system adjustments based on them.

[0991] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0992] The present invention relates to a system that provides customized smart speaker functions based on the collection and analysis of user data, and further combines an emotion engine to respond to the user's emotional state.

[0993] Explanation of program processing

[0994] Collection and analysis of user data

[0995] The server retrieves user data from the internet service database, including the user's message history, search history, and app usage data, and uses this data to analyze the user's behavioral patterns and preferences and create a profile.

[0996] As a specific example, the server analyzes the news articles that the user has viewed and the product data that the user has purchased, and can provide news and product information through the speaker according to the user's preferences.

[0997] Replacing and configuring communication devices

[0998] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user goes through the initial setup wizard to perform basic setup, including Wi-Fi settings and entering a user ID.

[0999] Enabling the smart speaker feature

[1000] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. The server then sends the generated user profile to the device and reflects the individual settings. This allows the user to use customized smart speaker functions.

[1001] Voice commands and emotion recognition

[1002] The user inputs a voice command into the smart speaker, for example, "Tell me today's news." The voice command is recognized by the device and converted into text data.

[1003] Here, the emotion engine recognizes the user's emotion based on the voice command. For example, the emotion engine determines whether the user is excited or calm based on the tone and speed of the voice.

[1004] Response generation and customization

[1005] The server receives the speech-recognized text data and performs data analysis, including emotional data from the emotion engine. Based on the analysis results, an appropriate response is generated. For example, if the recognized emotion is joy, the server generates a response such as, "It's nice weather today, enjoy going outside."

[1006] The response data is sent to the terminal, which then runs a speech synthesis algorithm to convert the data into voice data, which is finally output from the terminal's speaker and presented to the user.

[1007] Handling User Feedback

[1008] The user can provide feedback on the smart speaker's response, for example, by saying, "Please be more specific." This feedback is also recognized by voice on the device and sent to the server.

[1009] The server has algorithms that analyze the feedback and incorporate it into future responses. The emotion engine also integrates emotional data into the user profile, which is updated based on the user's preferences and emotional state.

[1010] Specific examples

[1011] After User B installs new communication equipment and completes the initial setup, the following scenario is assumed:

[1012] 1. User B: "What's in the news today?"

[1013] 2. Terminal: Recognizes speech and converts it into text.

[1014] 3. Emotion engine: Recognizes the user's excited emotions from their voice.

[1015] 4. Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[1016] 5. Terminal: Provides the response as audio to the user.

[1017] In this way, the present invention provides a customized smart speaker experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[1018] The processing flow will be explained below.

[1019] Step 1:

[1020] The server retrieves user data, including the user's behavioral history, preferences, and personal information, from the database of the internet service, such as message history, search history, and application usage data.

[1021] Step 2:

[1022] The server analyzes the acquired user data and uses machine learning algorithms to identify the user's behavioral patterns and preferences, then generates a user profile based on that. For example, a user's cooking interests can be identified from their search history and a profile created based on that.

[1023] Step 3:

[1024] The terminal will install a new communication device (with smart speaker functionality), remove the old communication device, and connect the new device to the network.

[1025] Step 4:

[1026] The user goes through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID. For example, the user selects their home Wi-Fi network and enters the password.

[1027] Step 5:

[1028] The server receives the initial setup information sent from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[1029] Step 6:

[1030] The server then sends the generated user profile to the device and reflects the user's personal settings. For example, if the user is interested in cooking, a function to suggest related recipes is enabled.

[1031] Step 7:

[1032] Users input voice commands into the smart speaker, for example, "Tell me the weather today."

[1033] Step 8:

[1034] The device converts the voice command into text data using a speech recognition algorithm, and the voice command "Tell me the weather today" is sent to the server as text data.

[1035] Step 9:

[1036] The server analyzes the text data and generates an appropriate response, for example, by retrieving the latest weather information from a weather forecast database and generating a response such as "Today's weather is sunny."

[1037] Step 10:

[1038] The emotion engine recognizes the user's emotions based on their voice commands, determining whether they are happy or sad based on the tone and speed of their voice.

[1039] Step 11:

[1040] The server takes into account the data from the emotion engine and generates a response that matches the user's emotional state. For example, if the user is feeling down, the server generates a response like "It's a sunny and pleasant day today."

[1041] Step 12:

[1042] The server sends the generated response data to the terminal, which converts the response into voice data using a voice synthesis algorithm.

[1043] Step 13:

[1044] The terminal outputs the converted voice data from a speaker to provide it to the user.

[1045] Step 14:

[1046] Users provide feedback on the smart speaker's response, for example, by saying, "I'd like more detailed weather information."

[1047] Step 15:

[1048] The terminal recognizes the user's feedback through speech recognition and transmits it to the server as text data.

[1049] Step 16:

[1050] The server analyzes the feedback and reflects it in future responses. It also integrates the emotional data recognized by the emotion engine into the user profile and updates it, allowing it to provide services that are more tailored to the user's preferences and emotional state.

[1051] In this way, a system is realized that provides users with an optimized smart speaker experience through user data collection, analysis, profile generation, emotion recognition, and feedback processing.

[1052] Example 2

[1053] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1054] Although existing smart devices are customized based on user behavior patterns and preferences, they lack the ability to consider the user's emotional state, making them unable to fully meet user needs. In addition, there are insufficient means to effectively utilize user feedback to improve the device's responsiveness.

[1055] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1056] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart device functions based on the generated profile, means for analyzing the user's emotional state using an emotion engine, means for acquiring user feedback, analyzing the acquired feedback and reflecting it in future responses, and means for integrating the emotion data acquired from the emotion engine into the profile and updating the profile based on the user's preferences and emotional state. This enables customization that takes into account the user's emotional state in addition to their behavioral patterns and preferences, and further enables continuous improvement of the system through user feedback.

[1057] "User data" refers to data obtained from Internet services, such as a user's message history, search history, and application usage data.

[1058] A "profile" is information that indicates a user's behavioral patterns and preferences, and is generated by analyzing user data.

[1059] "Smart device features" are features of smart speakers and similar devices that are customized based on a user profile.

[1060] A "voice command" is a voice input that includes instructions or questions spoken by a user to a smart device.

[1061] "Voice recognition" is a technology that analyzes input voice commands and converts them into text data.

[1062] The "emotion engine" is a technology for analyzing the user's emotional state from the tone and speed of their voice.

[1063] An "appropriate response" is a response message to the user that is generated based on the voice command and the user's emotional state.

[1064] "Feedback" refers to information including opinions and reactions provided by a user in response to a smart device's response.

[1065] "User authentication" is the process that takes place to verify a user's identity.

[1066] The present invention relates to a system that provides customized smart device functions based on the collection and analysis of user data, and furthermore, combines an emotion engine to respond to the user's emotional state.

[1067] First, the server retrieves user data from the Internet service database. This user data includes the user's message history, search history, and application usage data. Specifically, the server uses a software module to collect this data and analyzes behavioral patterns and preferences. A profile is generated from the analyzed data. This profile is unique to each user and reflects their individual preferences and behavioral patterns.

[1068] Next, the terminal replaces its existing communication device with a new communication device, which has smart device functionality integrated into it. After installing the terminal, the user configures Wi-Fi and enters their user ID through an initial setup wizard. All initial setup information is sent to the server, where user authentication is performed. If authentication is successful, the smart device functionality is enabled, and the server sends the generated profile to the terminal, reflecting the customized settings.

[1069] Next, the user inputs a voice command into the smart device. For example, they might say, "Tell me today's news." The device receives this voice command and converts it into text data using speech recognition software. The emotion engine then works to recognize emotions from the tone and speed of the user's voice. This allows the device to understand the user's emotional state, such as whether they are excited or calm.

[1070] The server analyzes the text data and emotion data generated by the speech recognition. Based on the analysis results, an appropriate response is generated. For example, if the user is excited, the server can respond with, "It's beautiful weather today, so please enjoy some activities." This response data is then sent back to the device, which uses a speech synthesis algorithm to output the response as voice data. Finally, the response is provided to the user as voice from the device.

[1071] When a user provides feedback on a smart device's response, for example by saying, "Please be more specific," this feedback is also recognized by voice on the device and sent to the server. The server analyzes the feedback and reflects it in future responses. During this process, the emotion engine also operates, integrating the acquired emotional data into the profile. This allows the profile to be updated more precisely to reflect the user's preferences and emotional state.

[1072] As a concrete example, a scenario will be shown in which user B has installed a new communication device and has completed the initial setup.

[1073] User B: "What's in the news today?"

[1074] Device: Recognizes speech and converts it to text.

[1075] Emotion engine: Recognizes the user's excited emotions from their voice.

[1076] Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[1077] Terminal: Provides the response as audio to the user.

[1078] In this way, the present invention provides a customized smart device experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[1079] Example prompt sentence:

[1080] Design a smart device system that delivers news based on a user's message and search history. The system will have voice command and emotion recognition capabilities, and will exchange data between the server and the device. Please explain the specific processing steps and technologies used, including examples of the system.

[1081] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1082] Step 1:

[1083] The user performs the initial setup of the smart device. Specifically, the user sets up Wi-Fi on the device and enters a user ID to complete the basic environment setup. This setup information is sent from the device to the server (input: Wi-Fi settings, user ID; output: sending initial setup information to the server).

[1084] Step 2:

[1085] The server receives the initial setting information and performs user authentication. If the user authentication is successful, the server enables the smart device function (input: initial setting information, output: user authentication result and enablement of smart device function).

[1086] Step 3:

[1087] The server retrieves user data from the Internet Services database, including the user's message history, search history, and application usage data (input: user database, output: retrieved user data).

[1088] Step 4:

[1089] The server analyzes the acquired user data and generates a profile based on the user's behavioral patterns and preferences (input: acquired user data, output: generated user profile).

[1090] Step 5:

[1091] The server transmits the generated user profile to the terminal, and reflects the individual settings on the terminal (input: generated user profile, output: user profile transmitted to the terminal).

[1092] Step 6:

[1093] The user inputs a voice command into the smart device, for example, "Tell me today's news" (input: voice command, output: voice input on the device).

[1094] Step 7:

[1095] The terminal converts the voice command into text data using voice recognition software (input: voice command, output: text data).

[1096] Step 8:

[1097] The emotion engine works by analyzing the tone and speed of the voice to recognize the user's emotional state (input: text data, output: user's emotional state).

[1098] Step 9:

[1099] The server analyzes the speech-recognized text data and emotion data and generates an appropriate response. For example, if the user is excited, it generates a response containing positive content (input: text data and emotion data, output: generated response).

[1100] Step 10:

[1101] The server sends the generated response to the terminal (input: generated response, output: response sent to terminal).

[1102] Step 11:

[1103] The device uses a speech synthesis algorithm to output the response as voice data (input: transmitted response, output: voice data). The device provides the voice data to the user through a speaker (input: voice data, output: voice provided to the user).

[1104] Step 12:

[1105] The user provides feedback on the smart device's response, for example, by saying "Please be more specific" (input: feedback speech, output: speech input on device).

[1106] Step 13:

[1107] The terminal recognizes the feedback by voice and sends it to the server (input: feedback voice, output: feedback text data).

[1108] Step 14:

[1109] The server analyzes the feedback and reflects it in future responses. The emotion data acquired by the emotion engine is also integrated into the profile, which is updated based on the user's preferences and emotional state (input: feedback text data and emotion data, output: updated user profile).

[1110] (Application example 2)

[1111] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1112] Conventional smart speaker systems have difficulty providing personalized responses and information based on user preferences, and they do not take into account the user's emotional state. This results in a uniform user experience, making it impossible to provide optimal services for each individual user. It is also difficult to effectively incorporate user feedback. This has led to a need for effective means to improve customer satisfaction in commercial environments such as brick-and-mortar stores.

[1113] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1114] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart speaker functions based on the generated profile, means for a user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for performing emotion analysis, means for analyzing the converted text data and emotion data to generate an appropriate response, means for outputting the generated response as voice, means for replacing the communication device with a new one, means for acquiring initial setting information and performing user authentication, means for activating the customized smart speaker function after successful authentication, means for acquiring user feedback, and means for analyzing the acquired feedback and integrating emotion data into the user profile to update it. This enables personalized responses and information provision according to the user's preferences and emotional state, thereby improving customer satisfaction in physical stores.

[1115] "User data" is a general term for information such as a user's behavior history, preference information, and application usage data.

[1116] A "profile" is information generated based on user data that analyzes and records a user's behavioral patterns and preferences.

[1117] "Smart speaker functions" refers to all functions of smart speakers, including voice recognition, voice response, and information provision.

[1118] A "voice command" is an instruction or question that a user speaks to a smart speaker.

[1119] "Speech recognition" refers to the general process of analyzing input voice data and converting it into text data.

[1120] "Emotion analysis" is a technology that recognizes and determines a user's emotional state from the tone and speed of their voice.

[1121] An "appropriate response" is a reply or information provided to the user that is generated based on the voice command and emotion data.

[1122] "Communication equipment" refers to any hardware device used to connect to the Internet or other devices.

[1123] "Initial setting information" refers to the setting information required to use new communication devices or services.

[1124] "User authentication" refers to the overall process of verifying a user's identity and granting access permissions.

[1125] "Feedback" refers to information such as responses, opinions, and suggestions obtained from users.

[1126] The present invention relates to a smart speaker system for providing personalized information and services to customers in physical stores. The specific configuration and operation of the system are described below.

[1127] System configuration and operation

[1128] Collection and analysis of user data

[1129] The server acquires user data, which includes behavioral history, preference information, and application usage data. The server analyzes this data to generate a profile that records the user's behavioral patterns and preferences. For example, the server analyzes the user's past purchase history to identify the user's favorite product categories.

[1130] Smart speaker setup and authentication

[1131] The device replaces the existing communication device with a new one. The user enters the initial setup information, and the server authenticates the user. If authentication is successful, the customized smart speaker function is activated.

[1132] Voice Commands and Sentiment Analysis

[1133] A user inputs a voice command into a smart speaker, for example, "Tell me about new promotions." The voice command is recognized by the device and converted into text data. Emotion analysis is then performed to recognize the user's emotional state from the tone and speed of the voice.

[1134] Generating and serving the response

[1135] The server analyzes the converted text data and emotional data to generate an appropriate response. For example, if the server recognizes that the user is excited, it generates a customized response such as, "We have new products at 20% off today!" The device outputs this response as speech.

[1136] Get feedback and update your profile

[1137] The user provides feedback on the smart speaker's response, for example, by saying, "Please be more specific." The device recognizes this feedback and sends it to the server, which analyzes it and updates the user profile, including emotional data.

[1138] Hardware and software used

[1139] This system uses the following hardware and software:

[1140] Speech Recognition Library: Convert speech to text using speech_recognition.

[1141] Sentiment Analysis Library: Performs sentiment analysis using a custom emotion_recognition library.

[1142] Text-to-speech: The generated response is converted into audio using the Google Text-to-Speech (gTTS) library.

[1143] Communications Equipment: Use common internet-connected devices for user authentication and data communication.

[1144] Specific examples

[1145] For example, if a user asks a smart speaker, "What are the new promotions?", the speech is converted into text and the user's emotional state is recognized through emotion analysis. Based on the user's profile and emotion data, the server generates a response such as, "We're offering 20% ​​off new products today!", which the device then delivers to the user as audio.

[1146] Prompt Sentence Examples

[1147] User ID: user123

[1148] User Emotion: Delight

[1149] User preferences: Interested in the latest gadgets

[1150] Response: New items are being offered at 20% off today!”

[1151] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1152] Step 1:

[1153] The server obtains the user data.

[1154] The user's behavioral history, preference information, and application usage data are obtained from the database of the internet service. This data is sent to the server and used as input data for analysis.

[1155] Step 2:

[1156] The server analyzes the user data and generates a profile.

[1157] The acquired user data is analyzed to analyze behavioral patterns and preferences, and a user profile is generated. This profile records the user's preferred product and service categories. Based on the analysis results, profile data is generated and passed on to the next process.

[1158] Step 3:

[1159] The terminal replaces the existing communication device with a new communication device.

[1160] The user goes through the initial setup wizard to configure Wi-Fi and enter their user ID, and the initial setup information is entered into the device, allowing the communication device to adapt to the new environment and saving the initial setup information to the device.

[1161] Step 4:

[1162] The server obtains the initial setting information and performs user authentication.

[1163] The server authenticates the user based on the initial setting information sent from the device. If the user ID and password are verified and authentication is successful, the customized smart speaker function is activated.

[1164] Step 5:

[1165] The user enters voice commands into the smart speaker.

[1166] For example, say, "Tell me about new promotions." This voice input is captured through the device's microphone.

[1167] Step 6:

[1168] The terminal recognizes the input voice command and converts it into text data.

[1169] The speech_recognition library is used to convert the audio data into text data, which is then sent to the server.

[1170] Step 7:

[1171] The terminal performs emotion analysis.

[1172] The emotion_recognition library is used to analyze the user's emotional state from the tone and rate of speech. Emotion data is generated and sent to the server along with the text data.

[1173] Step 8:

[1174] The server analyzes the converted text data and emotion data to generate an appropriate response.

[1175] Based on this data, the server uses a generative AI model to generate a customized response. For example, if the emotion of joy is recognized, the server generates a response such as, "We have new products at 20% off today!" The generated response is sent to the device for speech synthesis.

[1176] Step 9:

[1177] The terminal outputs the generated response as voice.

[1178] The Google Text-to-Speech (gTTS) library is used to convert the text response into speech, which is then output to the speaker and presented to the user.

[1179] Step 10:

[1180] The user provides feedback on the smart speaker's response.

[1181] For example, say, "Tell me more specifically." This feedback is captured through the device's microphone.

[1182] Step 11:

[1183] The terminal recognizes the feedback by voice and transmits it to the server.

[1184] Again, we use the speech_recognition library to convert the audio feedback into text data and send it to the server.

[1185] Step 12:

[1186] The server analyzes the feedback and integrates the emotional data into the user profile to update it.

[1187] Update user profiles based on feedback and sentiment data, leading to more personalized responses and a better user experience.

[1188] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1189] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1190] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1191] [Fourth embodiment]

[1192] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1193] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1194] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1195] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1196] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1198] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1199] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1200] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1201] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1202] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1203] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1204] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1205] The present invention relates to a system for providing customized smart speaker functionality based on the collection and analysis of user data.

[1206] Explanation of program processing

[1207] Collection and analysis of user data

[1208] The server collects user data, including user behavior history, preferences, and personal information. For example, the server analyzes the user's interests based on the user's internet search history and application usage data. From the results of this analysis, a customized profile is created for each user.

[1209] Replacing and configuring communication devices

[1210] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user performs basic settings using an initial setup wizard. The initial setup includes Wi-Fi settings and user ID entry.

[1211] Enabling the smart speaker feature

[1212] The server receives the initial setup information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. This applies the customized profile to the communication device. The user can immediately use the smart speaker function.

[1213] Processing voice commands

[1214] The user inputs a voice command into the smart speaker. For example, they might say, "Tell me what the weather will be tomorrow." This voice command is recognized by the device and converted into text data.

[1215] The server receives the text data that has been recognized by speech recognition and analyzes the command. Based on the analysis results, it generates an appropriate response. For example, it may generate a response such as "Tomorrow's weather will be sunny." This response data is then sent to the terminal and output as voice.

[1216] Handling User Feedback

[1217] Users can provide feedback on the smart speaker's response, for example by stating their opinion such as "Please be more specific." This feedback is recognized by the device and then sent to the server.

[1218] The server has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the smart speaker's responses to evolve to more closely match user expectations.

[1219] Specific examples

[1220] After User A installs new communication equipment and completes the initial setup, the following scenario is assumed:

[1221] 1. User A: "What's in the news today?"

[1222] 2. Terminal: Recognizes speech and converts it into text.

[1223] 3. Server: Analyzes the text data, collects news information, and generates appropriate responses.

[1224] 4. Response: "The major news stories today are the stock market boom and the announcement of new government policies."

[1225] 5. Terminal: Provides the response as audio to the user.

[1226] In this way, the present invention provides users with an optimized smart speaker experience and realizes a system that is easy to use and intuitive.

[1227] The processing flow will be explained below.

[1228] Step 1:

[1229] The server retrieves user data from LINE and Yahoo! databases, including user message history, search history, and app usage data.

[1230] Step 2:

[1231] The server analyzes the acquired user data and uses machine learning algorithms to identify user behavior patterns and preferences, generating a user profile from the analysis results.

[1232] Step 3:

[1233] The device is a new communication device (with smart speaker functionality) that is installed to replace the existing communication device. The new device requires basic settings to establish a network connection.

[1234] Step 4:

[1235] Users go through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID.

[1236] Step 5:

[1237] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[1238] Step 6:

[1239] The server then sends the generated user profile to the device, reflecting the individual settings, and the user is ready to use the customized smart speaker functions.

[1240] Step 7:

[1241] A user inputs a voice command to a smart speaker, for example, "Tell me the weather tomorrow."

[1242] Step 8:

[1243] The device converts the received voice commands into text data using a voice recognition algorithm, which is then sent to the server.

[1244] Step 9:

[1245] The server analyzes the text data and generates an appropriate response, for example, "Tomorrow's weather will be sunny."

[1246] Step 10:

[1247] The server sends the generated response data to the terminal, which runs a speech synthesis algorithm to convert the data into speech data.

[1248] Step 11:

[1249] The terminal outputs the audio data from a speaker to provide it to the user.

[1250] Step 12:

[1251] The user provides feedback on the smart speaker's response, such as saying, "The response was accurate" or "I wish it was shorter."

[1252] Step 13:

[1253] The device recognizes the user's voice and sends the feedback to the server, which analyzes it and uses it in future responses.

[1254] Example 1

[1255] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1256] Conventional smart devices have been unable to adequately adapt to individual users' needs and habits. It has been particularly difficult to effectively analyze user data and use smart speakers in a more personalized way. Initial setup and device replacement are also burdensome for users, making it difficult to provide responses that meet user expectations. Furthermore, there has been a lack of mechanisms for utilizing user feedback to improve responses.

[1257] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1258] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a personal profile, and means for customizing smart device functions based on the generated personal profile, thereby enabling the provision of personalized smart speaker functions according to the individual needs and habits of the user.

[1259] "User data" refers to information including a user's behavioral history, internet search history, data on applications used, personal information, and preferences.

[1260] A "personal profile" is information that indicates a user's interests, lifestyle habits, etc., generated as a result of analyzing user data.

[1261] "Smart device functionality" is the functionality of a digital device that provides voice recognition and voice response and acts based on user instructions.

[1262] A "voice instruction" is a voice command issued by a user to a smart device.

[1263] "Text information" is data in a document format that is generated by converting voice instructions using a voice recognition system.

[1264] A "response" is a reply in text or audio format that the server generates as a result of analyzing a voice command.

[1265] A "communications device" is an electronic device that allows for the transmission and reception of data.

[1266] "User authentication" is the process of identifying a particular user and verifying their permissions.

[1267] "Feedback" refers to opinions and suggestions that a user provides in response to a smart device's response.

[1268] The present invention is a system that provides customized smart device functions based on the collection and analysis of user data. This system collects data on user behavior history and applications used, and generates a customized profile for each user, thereby providing a personalized experience tailored to individual needs.

[1269] Hardware and software used

[1270] The main hardware required to implement the system includes a smart speaker, an internet-enabled device, and a server, while the software includes a voice recognition system, a data analysis engine, and a user profile generation algorithm.

[1271] Collection and analysis of user data

[1272] The server collects data about users' internet browsing history and the applications they use. This data comes from a variety of sources, including cookie information and log data. The collected data is analyzed using algorithms, such as machine learning models, to generate a personalized profile for each user. This profile reflects the user's interests, preferences, and preferences.

[1273] Replacing and initial setup of communication equipment

[1274] The terminal replaces the existing communication device with a new one. The new communication device has smart device functionality integrated into it. After the device is installed, the user uses an initial setup wizard to connect to a Wi-Fi network and enter their user ID. This configuration information is sent to the server.

[1275] Enabling the smart speaker feature

[1276] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device functions are enabled and the customized profile is applied, allowing the user to immediately start using the smart speaker.

[1277] Processing voice instructions

[1278] A user gives a voice command to a smart speaker. For example, they might say, "Tell me what the weather will be like tomorrow." This voice is recorded on the device and converted into text information in real time by a voice recognition system. This text information is sent to a server and analyzed. Based on the analysis results, an appropriate response is generated. For example, a response such as "Tomorrow's weather will be sunny" is generated, and this information is sent to the device and provided to the user as voice.

[1279] Handling User Feedback

[1280] Users can provide feedback on the smart speaker's response by voice. For example, they can say, "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server implements an algorithm that analyzes this feedback and reflects it in future responses. This allows the smart speaker's responses to better meet user expectations.

[1281] Examples of prompt statements

[1282] Examples of prompts for generative AI models include:

[1283] "Please tell me the details of the program processing for the user data collection and analysis system. Please provide a detailed explanation, including the names of the hardware and software used, and the details of data processing and calculation."

[1284] As described above, the present invention provides users with an optimized smart device experience and realizes a system that is easy to use and intuitive.

[1285] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1286] Step 1: Collect user data

[1287] The server collects user data such as internet search history, application usage data, and behavioral history. Specifically, it collects cookie information and log data. It also collects personal data that the user has authorized.

[1288] Input: Internet logs, cookie data, application usage data

[1289] Data processing: Filtering log data and removing unnecessary data

[1290] Output: The curated user dataset

[1291] Step 2: Analyze data and generate profiles

[1292] The server analyzes the collected user data, specifically using machine learning algorithms to classify the user's interests and concerns, and generates a personalized profile for each user based on the results of this analysis.

[1293] Input: Curated user dataset

[1294] Data Computing: Analysis based on machine learning models

[1295] Output: A customized personal profile

[1296] Step 3: Replace the communication device

[1297] The terminal replaces the old communication device with a new one, which has integrated smart device functionality, and a technician visits the customer to physically install the device and connect the wiring.

[1298] Input: Existing communication equipment, replacement communication equipment

[1299] Data processing: None (physical exchange work)

[1300] Output: New communications device installed

[1301] Step 4: Perform initial configuration

[1302] After installing a new communication device, the user uses the initial setup wizard to perform basic settings, such as connecting to a Wi-Fi network and entering a user ID. This setting information is then sent to the server.

[1303] Input: Wi-Fi information, user ID

[1304] Data processing: Format conversion of setting information

[1305] Output: Send initial setting information to server

[1306] Step 5: Enable the Smart Speaker feature

[1307] The server authenticates the user based on the received initial configuration information. If authentication is successful, the smart device function is enabled and the customized profile is applied. The user receives a notification that "Smart speaker has been enabled."

[1308] Input: Initial setting information

[1309] Data calculation: User authentication

[1310] Output: Enable smart device features, send notifications

[1311] Step 6: Processing voice instructions

[1312] The user gives voice commands to the smart speaker. The voice is recorded on the device and converted into text information in real time by a speech recognition system. This text information is sent to the server and analyzed. Based on the analysis results, an appropriate response is generated and provided to the user as voice.

[1313] Input: Voice commands

[1314] Data processing: Text conversion using speech recognition

[1315] Output: Response based on analysis results, providing voice response

[1316] Step 7: Processing user feedback

[1317] The user provides feedback on the smart speaker's response. For example, they may state their opinion, such as "Please be more specific." This feedback is recognized by voice on the device and sent as text information to the server. The server analyzes this feedback and improves the response algorithm. This allows future responses to more faithfully meet the user's expectations.

[1318] Input: Feedback voice

[1319] Data processing: Text conversion by voice recognition, feedback analysis

[1320] Output: Improved response algorithm

[1321] In this way, the system provides a personalized smart device experience tailored to the user's individual needs.

[1322] (Application example 1)

[1323] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1324] Conventional automotive voice assistant systems are unable to fully reflect the user's individual preferences and behavioral history, resulting in uniform information and services that make it difficult to provide an optimal experience for the user. Furthermore, adjusting the in-car environment requires a lot of manual settings, placing a heavy burden on the user. The present invention aims to solve these problems and realize a personalized in-car assistant based on the user's preferences and behavioral history.

[1325] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1326] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing the voice assistant function based on the generated profile, means for the user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for analyzing the converted text data to generate an appropriate response, means for outputting the generated response as voice, means for generating and transmitting a control signal for adjusting the environment inside the vehicle, and means for automatically adjusting the in-vehicle environment settings (temperature, lighting, music, etc.) based on the user's preferences. This enables personalized information provision based on the user's preferences and behavioral history, and automatic adjustment of the in-vehicle environment.

[1327] "User data" refers to data such as a user's behavioral history, preferences, and personal information.

[1328] A "profile" is a collection of information that indicates a user's characteristics and preferences, generated by analyzing user data.

[1329] "Voice assistant function" refers to the function of receiving voice commands, generating appropriate responses, and outputting them in voice.

[1330] "Voice command" refers to instructions or questions that a user enters by voice.

[1331] "Speech recognition" refers to the technology of converting voice data into text data.

[1332] "Text data" refers to text information generated by speech recognition.

[1333] A "response" refers to information or instructions provided to a user that are generated based on the analyzed text data.

[1334] "Output" refers to conveying the generated response to the user as voice.

[1335] "Control signal" refers to a command signal sent to control a device within a vehicle.

[1336] "Environmental settings" refers to settings such as temperature, lighting, and music inside the car.

[1337] "Preferences" refer to the user's tastes and interests.

[1338] "Initial setting information" refers to the setting information required to activate a new communication device.

[1339] "User authentication" is a procedure for verifying that a user is a legitimate user.

[1340] "Feedback" refers to evaluations and opinions regarding responses and services provided by users.

[1341] "Analysis" refers to the process of analyzing acquired data to extract useful information.

[1342] The present invention relates to a system for providing personalized information and automatically adjusting the in-vehicle environment based on the user's preferences and behavioral history in an autonomous vehicle.

[1343] The server acquires user data, including the user's behavioral history, preferences, and personal information. The data is acquired using internet search history and application usage data. By analyzing this data, a customized profile for each user is created.

[1344] The server customizes the voice assistant function based on the generated profile, enabling it to provide optimal information to the user. Additionally, when a user enters a voice command, the terminal (communication equipment inside the autonomous vehicle) uses voice recognition technology to convert the voice command into text data, using the Google Cloud Speech-to-Text API.

[1345] The server analyzes the converted text data and generates an appropriate response using the Python data analysis library scikit-learn. The generated response is then converted into audio using the Google Cloud Text-to-Speech API, which is then output to the user through the car's speakers.

[1346] Furthermore, the terminal generates and transmits control signals to adjust the vehicle's interior environment, which are transmitted via the MQTT protocol and control various devices (such as temperature control, lighting, and audio) via the vehicle's CAN bus (in-vehicle communication network).

[1347] Users can provide feedback on the provided responses by voice, which is then re-audio-recognized and sent to the server, which has an algorithm that analyzes the feedback and incorporates it into future responses, allowing the voice assistant to continually improve its accuracy.

[1348] For example, if a user says to a smart speaker installed in the car, "Tell me today's news," the personal in-car assistant will automatically retrieve the news and respond aloud, "Today's main news stories are a sharp rise in stock prices and the announcement of new government policies."

[1349] Next, say "Make the car warmer," and the car's temperature will be adjusted to the appropriate level based on the user's profile.

[1350] Examples of prompt sentences that are important in the present invention include the following:

[1351] 1. "Collect user data and generate user profiles"

[1352] 2. "Analyzing user preferences based on internet search history"

[1353] 3. "Accept voice commands from a smart speaker and convert the voice to text."

[1354] 4. "Improve system response based on user feedback"

[1355] This invention makes it possible to provide information optimized for the user in an autonomous vehicle and automatically adjust the in-vehicle environment to be comfortable.

[1356] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1357] Step 1:

[1358] User Data Collection:

[1359] The server collects data such as user behavior, preferences, and personal information. This includes internet browsing history and application usage data. This data is obtained by using APIs to collect data from external services. It may also be obtained through sensors or other interfaces. The input for this step is internet browsing history and application usage data, and the output is raw user data for analysis.

[1360] Step 2:

[1361] User data analysis:

[1362] The server analyzes the collected user data to understand user preferences and behavioral patterns. It uses Python and data analysis libraries such as scikit-learn to perform the analysis. It preprocesses the raw data and generates profiles using techniques such as feature extraction, clustering, and classification. The input for this step is raw user data, and the output is a profile customized for each user.

[1363] Step 3:

[1364] Creating profile-based customizations:

[1365] Based on the generated user profile, the server customizes the voice assistant's functionality by specifying the type of music, news, and information the user prefers, and also applies algorithms to provide appropriate feedback and adjustments. The input for this step is the generated profile, and the output is customized voice assistant settings.

[1366] Step 4:

[1367] Receiving and recognizing voice commands:

[1368] When a user inputs a voice command into the smart speaker in the car, the device captures the speech and converts it into text data using the Google Cloud Speech-to-Text API. The input of this step is the user's voice command, and the output is text data.

[1369] Step 5:

[1370] Voice command analysis:

[1371] The server analyzes the speech-recognized text data and generates an appropriate response to the user's request. The analysis uses natural language processing and rule-based algorithms. The server also takes into account the user's profile to provide the most appropriate information and services. The input for this step is text data, and the output is response data.

[1372] Step 6:

[1373] Response transcription and output:

[1374] The server converts the generated response data into speech using the Google Cloud Text-to-Speech API. The device then outputs the converted speech to the user through the car speaker. The input of this step is the response data, and the output is a speech response.

[1375] Step 7:

[1376] Control signal generation and transmission:

[1377] The server generates control signals to adjust the in-car environment based on the user's preference profile and sends them to the terminal via the MQTT protocol. The terminal receives these control signals and controls each device (temperature, lighting, audio, etc.) via the vehicle's CAN bus. The input of this step is the user profile and control request, and the output is the control signal sent to each device.

[1378] Step 8:

[1379] Collecting and analyzing user feedback:

[1380] Users can provide feedback on the voice assistant's responses, which are then re-audio-recognized and sent to the server. The server then has an algorithm that analyzes the feedback and incorporates it into future responses, improving the accuracy of the voice assistant. The input for this step is the user's voice feedback, and the output is the analysis results and the system adjustments based on them.

[1381] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1382] The present invention relates to a system that provides customized smart speaker functions based on the collection and analysis of user data, and further combines an emotion engine to respond to the user's emotional state.

[1383] Explanation of program processing

[1384] Collection and analysis of user data

[1385] The server retrieves user data from the internet service database, including the user's message history, search history, and app usage data, and uses this data to analyze the user's behavioral patterns and preferences and create a profile.

[1386] As a specific example, the server analyzes the news articles that the user has viewed and the product data that the user has purchased, and can provide news and product information through the speaker according to the user's preferences.

[1387] Replacing and configuring communication devices

[1388] The device replaces the existing communication device with a new one. The new communication device has smart speaker functionality integrated into it. After installation, the user goes through the initial setup wizard to perform basic setup, including Wi-Fi settings and entering a user ID.

[1389] Enabling the smart speaker feature

[1390] The server receives the initial setting information from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled. The server then sends the generated user profile to the device and reflects the individual settings. This allows the user to use customized smart speaker functions.

[1391] Voice commands and emotion recognition

[1392] The user inputs a voice command into the smart speaker, for example, "Tell me today's news." The voice command is recognized by the device and converted into text data.

[1393] Here, the emotion engine recognizes the user's emotion based on the voice command. For example, the emotion engine determines whether the user is excited or calm based on the tone and speed of the voice.

[1394] Response generation and customization

[1395] The server receives the speech-recognized text data and performs data analysis, including emotional data from the emotion engine. Based on the analysis results, an appropriate response is generated. For example, if the recognized emotion is joy, the server generates a response such as, "It's nice weather today, enjoy going outside."

[1396] The response data is sent to the terminal, which then runs a speech synthesis algorithm to convert the data into voice data, which is finally output from the terminal's speaker and presented to the user.

[1397] Handling User Feedback

[1398] The user can provide feedback on the smart speaker's response, for example, by saying, "Please be more specific." This feedback is also recognized by voice on the device and sent to the server.

[1399] The server has algorithms that analyze the feedback and incorporate it into future responses. The emotion engine also integrates emotional data into the user profile, which is updated based on the user's preferences and emotional state.

[1400] Specific examples

[1401] After User B installs new communication equipment and completes the initial setup, the following scenario is assumed:

[1402] 1. User B: "What's in the news today?"

[1403] 2. Terminal: Recognizes speech and converts it into text.

[1404] 3. Emotion engine: Recognizes the user's excited emotions from their voice.

[1405] 4. Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[1406] 5. Terminal: Provides the response as audio to the user.

[1407] In this way, the present invention provides a customized smart speaker experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[1408] The processing flow will be explained below.

[1409] Step 1:

[1410] The server retrieves user data, including the user's behavioral history, preferences, and personal information, from the database of the internet service, such as message history, search history, and application usage data.

[1411] Step 2:

[1412] The server analyzes the acquired user data and uses machine learning algorithms to identify the user's behavioral patterns and preferences, then generates a user profile based on that. For example, a user's cooking interests can be identified from their search history and a profile created based on that.

[1413] Step 3:

[1414] The terminal will install a new communication device (with smart speaker functionality), remove the old communication device, and connect the new device to the network.

[1415] Step 4:

[1416] The user goes through the device's initial setup wizard to configure basic settings, including Wi-Fi settings and entering a user ID. For example, the user selects their home Wi-Fi network and enters the password.

[1417] Step 5:

[1418] The server receives the initial setup information sent from the device and performs user authentication. If authentication is successful, the smart speaker function is enabled.

[1419] Step 6:

[1420] The server then sends the generated user profile to the device and reflects the user's personal settings. For example, if the user is interested in cooking, a function to suggest related recipes is enabled.

[1421] Step 7:

[1422] Users input voice commands into the smart speaker, for example, "Tell me the weather today."

[1423] Step 8:

[1424] The device converts the voice command into text data using a speech recognition algorithm, and the voice command "Tell me the weather today" is sent to the server as text data.

[1425] Step 9:

[1426] The server analyzes the text data and generates an appropriate response, for example, by retrieving the latest weather information from a weather forecast database and generating a response such as "Today's weather is sunny."

[1427] Step 10:

[1428] The emotion engine recognizes the user's emotions based on their voice commands, determining whether they are happy or sad based on the tone and speed of their voice.

[1429] Step 11:

[1430] The server takes into account the data from the emotion engine and generates a response that matches the user's emotional state. For example, if the user is feeling down, the server generates a response like "It's a sunny and pleasant day today."

[1431] Step 12:

[1432] The server sends the generated response data to the terminal, which converts the response into voice data using a voice synthesis algorithm.

[1433] Step 13:

[1434] The terminal outputs the converted voice data from a speaker to provide it to the user.

[1435] Step 14:

[1436] Users provide feedback on the smart speaker's response, for example, by saying, "I'd like more detailed weather information."

[1437] Step 15:

[1438] The terminal recognizes the user's feedback through speech recognition and transmits it to the server as text data.

[1439] Step 16:

[1440] The server analyzes the feedback and reflects it in future responses. It also integrates the emotional data recognized by the emotion engine into the user profile and updates it, allowing it to provide services that are more tailored to the user's preferences and emotional state.

[1441] In this way, a system is realized that provides users with an optimized smart speaker experience through user data collection, analysis, profile generation, emotion recognition, and feedback processing.

[1442] Example 2

[1443] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1444] Although existing smart devices are customized based on user behavior patterns and preferences, they lack the ability to consider the user's emotional state, making them unable to fully meet user needs. In addition, there are insufficient means to effectively utilize user feedback to improve the device's responsiveness.

[1445] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1446] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart device functions based on the generated profile, means for analyzing the user's emotional state using an emotion engine, means for acquiring user feedback, analyzing the acquired feedback and reflecting it in future responses, and means for integrating the emotion data acquired from the emotion engine into the profile and updating the profile based on the user's preferences and emotional state. This enables customization that takes into account the user's emotional state in addition to their behavioral patterns and preferences, and further enables continuous improvement of the system through user feedback.

[1447] "User data" refers to data obtained from Internet services, such as a user's message history, search history, and application usage data.

[1448] A "profile" is information that indicates a user's behavioral patterns and preferences, and is generated by analyzing user data.

[1449] "Smart device features" are features of smart speakers and similar devices that are customized based on a user profile.

[1450] A "voice command" is a voice input that includes instructions or questions spoken by a user to a smart device.

[1451] "Voice recognition" is a technology that analyzes input voice commands and converts them into text data.

[1452] The "emotion engine" is a technology for analyzing the user's emotional state from the tone and speed of their voice.

[1453] An "appropriate response" is a response message to the user that is generated based on the voice command and the user's emotional state.

[1454] "Feedback" refers to information including opinions and reactions provided by a user in response to a smart device's response.

[1455] "User authentication" is the process that takes place to verify a user's identity.

[1456] The present invention relates to a system that provides customized smart device functions based on the collection and analysis of user data, and furthermore, combines an emotion engine to respond to the user's emotional state.

[1457] First, the server retrieves user data from the Internet service database. This user data includes the user's message history, search history, and application usage data. Specifically, the server uses a software module to collect this data and analyzes behavioral patterns and preferences. A profile is generated from the analyzed data. This profile is unique to each user and reflects their individual preferences and behavioral patterns.

[1458] Next, the terminal replaces its existing communication device with a new communication device, which has smart device functionality integrated into it. After installing the terminal, the user configures Wi-Fi and enters their user ID through an initial setup wizard. All initial setup information is sent to the server, where user authentication is performed. If authentication is successful, the smart device functionality is enabled, and the server sends the generated profile to the terminal, reflecting the customized settings.

[1459] Next, the user inputs a voice command into the smart device. For example, they might say, "Tell me today's news." The device receives this voice command and converts it into text data using speech recognition software. The emotion engine then works to recognize emotions from the tone and speed of the user's voice. This allows the device to understand the user's emotional state, such as whether they are excited or calm.

[1460] The server analyzes the text data and emotion data generated by the speech recognition. Based on the analysis results, an appropriate response is generated. For example, if the user is excited, the server can respond with, "It's beautiful weather today, so please enjoy some activities." This response data is then sent back to the device, which uses a speech synthesis algorithm to output the response as voice data. Finally, the response is provided to the user as voice from the device.

[1461] When a user provides feedback on a smart device's response, for example by saying, "Please be more specific," this feedback is also recognized by voice on the device and sent to the server. The server analyzes the feedback and reflects it in future responses. During this process, the emotion engine also operates, integrating the acquired emotional data into the profile. This allows the profile to be updated more precisely to reflect the user's preferences and emotional state.

[1462] As a concrete example, a scenario will be shown in which user B has installed a new communication device and has completed the initial setup.

[1463] User B: "What's in the news today?"

[1464] Device: Recognizes speech and converts it to text.

[1465] Emotion engine: Recognizes the user's excited emotions from their voice.

[1466] Server: Analyzes the text and sentiment data and generates a response such as, "Today's main news is the results of a major sporting event and the announcement of a new technology!"

[1467] Terminal: Provides the response as audio to the user.

[1468] In this way, the present invention provides a customized smart device experience that takes into account the user's emotions, realizing a system that is easy to use and intuitive.

[1469] Example prompt sentence:

[1470] Design a smart device system that delivers news based on a user's message and search history. The system will have voice command and emotion recognition capabilities, and will exchange data between the server and the device. Please explain the specific processing steps and technologies used, including examples of the system.

[1471] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1472] Step 1:

[1473] The user performs the initial setup of the smart device. Specifically, the user sets up Wi-Fi on the device and enters a user ID to complete the basic environment setup. This setup information is sent from the device to the server (input: Wi-Fi settings, user ID; output: sending initial setup information to the server).

[1474] Step 2:

[1475] The server receives the initial setting information and performs user authentication. If the user authentication is successful, the server enables the smart device function (input: initial setting information, output: user authentication result and enablement of smart device function).

[1476] Step 3:

[1477] The server retrieves user data from the Internet Services database, including the user's message history, search history, and application usage data (input: user database, output: retrieved user data).

[1478] Step 4:

[1479] The server analyzes the acquired user data and generates a profile based on the user's behavioral patterns and preferences (input: acquired user data, output: generated user profile).

[1480] Step 5:

[1481] The server transmits the generated user profile to the terminal, and reflects the individual settings on the terminal (input: generated user profile, output: user profile transmitted to the terminal).

[1482] Step 6:

[1483] The user inputs a voice command into the smart device, for example, "Tell me today's news" (input: voice command, output: voice input on the device).

[1484] Step 7:

[1485] The terminal converts the voice command into text data using voice recognition software (input: voice command, output: text data).

[1486] Step 8:

[1487] The emotion engine works by analyzing the tone and speed of the voice to recognize the user's emotional state (input: text data, output: user's emotional state).

[1488] Step 9:

[1489] The server analyzes the speech-recognized text data and emotion data and generates an appropriate response. For example, if the user is excited, it generates a response containing positive content (input: text data and emotion data, output: generated response).

[1490] Step 10:

[1491] The server sends the generated response to the terminal (input: generated response, output: response sent to terminal).

[1492] Step 11:

[1493] The device uses a speech synthesis algorithm to output the response as voice data (input: transmitted response, output: voice data). The device provides the voice data to the user through a speaker (input: voice data, output: voice provided to the user).

[1494] Step 12:

[1495] The user provides feedback on the smart device's response, for example, by saying "Please be more specific" (input: feedback speech, output: speech input on device).

[1496] Step 13:

[1497] The terminal recognizes the feedback by voice and sends it to the server (input: feedback voice, output: feedback text data).

[1498] Step 14:

[1499] The server analyzes the feedback and reflects it in future responses. The emotion data acquired by the emotion engine is also integrated into the profile, which is updated based on the user's preferences and emotional state (input: feedback text data and emotion data, output: updated user profile).

[1500] (Application example 2)

[1501] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1502] Conventional smart speaker systems have difficulty providing personalized responses and information based on user preferences, and they do not take into account the user's emotional state. This results in a uniform user experience, making it impossible to provide optimal services for each individual user. It is also difficult to effectively incorporate user feedback. This has led to a need for effective means to improve customer satisfaction in commercial environments such as brick-and-mortar stores.

[1503] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1504] In this invention, the server includes means for acquiring user data, means for analyzing the user data to generate a profile, means for customizing smart speaker functions based on the generated profile, means for a user to input a voice command, means for speech recognition of the input voice command and converting it into text data, means for performing emotion analysis, means for analyzing the converted text data and emotion data to generate an appropriate response, means for outputting the generated response as voice, means for replacing the communication device with a new one, means for acquiring initial setting information and performing user authentication, means for activating the customized smart speaker function after successful authentication, means for acquiring user feedback, and means for analyzing the acquired feedback and integrating emotion data into the user profile to update it. This enables personalized responses and information provision according to the user's preferences and emotional state, thereby improving customer satisfaction in physical stores.

[1505] "User data" is a general term for information such as a user's behavior history, preference information, and application usage data.

[1506] A "profile" is information generated based on user data that analyzes and records a user's behavioral patterns and preferences.

[1507] "Smart speaker functions" refers to all functions of smart speakers, including voice recognition, voice response, and information provision.

[1508] A "voice command" is an instruction or question that a user speaks to a smart speaker.

[1509] "Speech recognition" refers to the general process of analyzing input voice data and converting it into text data.

[1510] "Emotion analysis" is a technology that recognizes and determines a user's emotional state from the tone and speed of their voice.

[1511] An "appropriate response" is a reply or information provided to the user that is generated based on the voice command and emotion data.

[1512] "Communication equipment" refers to any hardware device used to connect to the Internet or other devices.

[1513] "Initial setting information" refers to the setting information required to use new communication devices or services.

[1514] "User authentication" refers to the overall process of verifying a user's identity and granting access permissions.

[1515] "Feedback" refers to information such as responses, opinions, and suggestions obtained from users.

[1516] The present invention relates to a smart speaker system for providing personalized information and services to customers in physical stores. The specific configuration and operation of the system are described below.

[1517] System configuration and operation

[1518] Collection and analysis of user data

[1519] The server acquires user data, which includes behavioral history, preference information, and application usage data. The server analyzes this data to generate a profile that records the user's behavioral patterns and preferences. For example, the server analyzes the user's past purchase history to identify the user's favorite product categories.

[1520] Smart speaker setup and authentication

[1521] The device replaces the existing communication device with a new one. The user enters the initial setup information, and the server authenticates the user. If authentication is successful, the customized smart speaker function is activated.

[1522] Voice Commands and Sentiment Analysis

[1523] A user inputs a voice command into a smart speaker, for example, "Tell me about new promotions." The voice command is recognized by the device and converted into text data. Emotion analysis is then performed to recognize the user's emotional state from the tone and speed of the voice.

[1524] Generating and serving the response

[1525] The server analyzes the converted text data and emotional data to generate an appropriate response. For example, if the server recognizes that the user is excited, it generates a customized response such as, "We have new products at 20% off today!" The device outputs this response as speech.

[1526] Get feedback and update your profile

[1527] The user provides feedback on the smart speaker's response, for example, by saying, "Please be more specific." The device recognizes this feedback and sends it to the server, which analyzes it and updates the user profile, including emotional data.

[1528] Hardware and software used

[1529] This system uses the following hardware and software:

[1530] Speech Recognition Library: Convert speech to text using speech_recognition.

[1531] Sentiment Analysis Library: Performs sentiment analysis using a custom emotion_recognition library.

[1532] Text-to-speech: The generated response is converted into audio using the Google Text-to-Speech (gTTS) library.

[1533] Communications Equipment: Use common internet-connected devices for user authentication and data communication.

[1534] Specific examples

[1535] For example, if a user asks a smart speaker, "What are the new promotions?", the speech is converted into text and the user's emotional state is recognized through emotion analysis. Based on the user's profile and emotion data, the server generates a response such as, "We're offering 20% ​​off new products today!", which the device then delivers to the user as audio.

[1536] Prompt Sentence Examples

[1537] User ID: user123

[1538] User Emotion: Delight

[1539] User preferences: Interested in the latest gadgets

[1540] Response: New items are being offered at 20% off today!”

[1541] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1542] Step 1:

[1543] The server obtains the user data.

[1544] The user's behavioral history, preference information, and application usage data are obtained from the database of the internet service. This data is sent to the server and used as input data for analysis.

[1545] Step 2:

[1546] The server analyzes the user data and generates a profile.

[1547] The acquired user data is analyzed to analyze behavioral patterns and preferences, and a user profile is generated. This profile records the user's preferred product and service categories. Based on the analysis results, profile data is generated and passed on to the next process.

[1548] Step 3:

[1549] The terminal replaces the existing communication device with a new communication device.

[1550] The user goes through the initial setup wizard to configure Wi-Fi and enter their user ID, and the initial setup information is entered into the device, allowing the communication device to adapt to the new environment and saving the initial setup information to the device.

[1551] Step 4:

[1552] The server obtains the initial setting information and performs user authentication.

[1553] The server authenticates the user based on the initial setting information sent from the device. If the user ID and password are verified and authentication is successful, the customized smart speaker function is activated.

[1554] Step 5:

[1555] The user enters voice commands into the smart speaker.

[1556] For example, say, "Tell me about new promotions." This voice input is captured through the device's microphone.

[1557] Step 6:

[1558] The terminal recognizes the input voice command and converts it into text data.

[1559] The speech_recognition library is used to convert the audio data into text data, which is then sent to the server.

[1560] Step 7:

[1561] The terminal performs emotion analysis.

[1562] The emotion_recognition library is used to analyze the user's emotional state from the tone and rate of speech. Emotion data is generated and sent to the server along with the text data.

[1563] Step 8:

[1564] The server analyzes the converted text data and emotion data to generate an appropriate response.

[1565] Based on this data, the server uses a generative AI model to generate a customized response. For example, if the emotion of joy is recognized, the server generates a response such as, "We have new products at 20% off today!" The generated response is sent to the device for speech synthesis.

[1566] Step 9:

[1567] The terminal outputs the generated response as voice.

[1568] The Google Text-to-Speech (gTTS) library is used to convert the text response into speech, which is then output to the speaker and presented to the user.

[1569] Step 10:

[1570] The user provides feedback on the smart speaker's response.

[1571] For example, say, "Tell me more specifically." This feedback is captured through the device's microphone.

[1572] Step 11:

[1573] The terminal recognizes the feedback by voice and transmits it to the server.

[1574] Again, we use the speech_recognition library to convert the audio feedback into text data and send it to the server.

[1575] Step 12:

[1576] The server analyzes the feedback and integrates the emotional data into the user profile to update it.

[1577] Update user profiles based on feedback and sentiment data, leading to more personalized responses and a better user experience.

[1578] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1579] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1580] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1581] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1582] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1583] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1584] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1585] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1586] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1587] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1588] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1589] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1590] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1591] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1592] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1593] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1594] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1595] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1596] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1597] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1598] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1599] The following is further disclosed regarding the above embodiment.

[1600] (Claim 1)

[1601] a means for obtaining user data;

[1602] means for analyzing user data to generate a profile;

[1603] means for customizing smart speaker functionality based on the generated profile; and

[1604] a means for a user to input voice commands;

[1605] A means for recognizing an input voice command and converting it into text data;

[1606] means for analyzing the converted text data to generate an appropriate response;

[1607] A means to output the generated response as speech

[1608] A system including:

[1609] (Claim 2)

[1610] a means of replacing the communications equipment with new communications equipment;

[1611] A means for acquiring initial setting information and performing user authentication;

[1612] and further including means for activating a smart speaker function after successful authentication.

[1613] 10. The system of claim 1.

[1614] (Claim 3)

[1615] a means for obtaining user feedback;

[1616] and further comprising means for analyzing the obtained feedback and incorporating it into future responses.

[1617] 10. The system of claim 1.

[1618] "Example 1"

[1619] (Claim 1)

[1620] a means for obtaining user data;

[1621] means for analyzing user data to generate a personal profile;

[1622] means for customizing smart device functionality based on the generated personal profile;

[1623] means for a user to input voice instructions;

[1624] A means for recognizing input voice instructions and converting them into text information;

[1625] means for analyzing the converted text information to generate an appropriate response;

[1626] The system includes means for outputting the generated response as speech.

[1627] (Claim 2)

[1628] means for replacing the existing communication device with a new communication device;

[1629] A means for acquiring initial setting information and performing user authentication;

[1630] and means for enabling the smart device function after successful authentication.

[1631] 10. The system of claim 1.

[1632] (Claim 3)

[1633] means for obtaining feedback on the user's response;

[1634] and further comprising means for analyzing the obtained feedback and incorporating it into future responses.

[1635] 10. The system of claim 1.

[1636] "Application Example 1"

[1637] (Claim 1)

[1638] a means for obtaining user data;

[1639] means for analyzing user data to generate a profile;

[1640] a means for customizing voice assistant functionality based on the generated profile; and

[1641] a means for a user to input voice commands;

[1642] A means for recognizing an input voice command and converting it into text data;

[1643] means for analyzing the converted text data to generate an appropriate response;

[1644] means for outputting the generated response as speech;

[1645] means for generating and transmitting control signals for adjusting the environment within the vehicle;

[1646] A means to automatically adjust in-car environmental settings (temperature, lighting, music, etc.) based on user preferences

[1647] A system including:

[1648] (Claim 2)

[1649] a means of replacing the communications equipment with new communications equipment;

[1650] A means for acquiring initial setting information and performing user authentication;

[1651] and means for activating a voice assistant function after successful authentication.

[1652] 10. The system of claim 1.

[1653] (Claim 3)

[1654] a means for obtaining user feedback;

[1655] and further comprising means for analyzing the obtained feedback and incorporating it into future responses.

[1656] 10. The system of claim 1.

[1657] "Example 2: Combining Emotion Engines"

[1658] (Claim 1)

[1659] a means for obtaining user data;

[1660] means for analyzing user data to generate a profile;

[1661] means for customizing smart device functionality based on the generated profile;

[1662] a means for a user to input voice commands;

[1663] A means for recognizing an input voice command and converting it into text data;

[1664] means for analyzing the converted text data to generate appropriate responses and recognize emotional states;

[1665] means for outputting the generated response as speech;

[1666] A means of analyzing a user's emotional state using an emotion engine

[1667] A system including:

[1668] (Claim 2)

[1669] a means of replacing the communications equipment with new communications equipment;

[1670] The smart device further includes a means for acquiring initial setting information, performing user authentication, and enabling the smart device function after successful authentication.

[1671] 10. The system of claim 1.

[1672] (Claim 3)

[1673] a means for acquiring user feedback, analyzing the acquired feedback, and incorporating the feedback into future responses;

[1674] means for integrating the emotion data obtained from the emotion engine into the profile and updating the profile based on the user's preferences and emotional state;

[1675] 10. The system of claim 1.

[1676] "Application example 2 when combining emotion engines"

[1677] (Claim 1)

[1678] a means for obtaining user data;

[1679] means for analyzing user data to generate a profile;

[1680] means for customizing smart speaker functionality based on the generated profile; and

[1681] a means for a user to input voice commands;

[1682] A means for recognizing an input voice command and converting it into text data;

[1683] a means for performing sentiment analysis;

[1684] means for analyzing the converted text data and emotion data to generate an appropriate response;

[1685] A means to output the generated response as speech

[1686] A system including:

[1687] (Claim 2)

[1688] A means of replacing the communication equipment with new equipment;

[1689] A means for acquiring initial setting information and performing user authentication;

[1690] and further including means for activating customized smart speaker functionality after successful authentication.

[1691] 10. The system of claim 1.

[1692] (Claim 3)

[1693] a means for obtaining user feedback;

[1694] and means for analyzing the acquired feedback and integrating the emotion data into a user profile to update the user profile.

[1695] 10. The system of claim 1. [Explanation of symbols]

[1696] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for obtaining user data; means for analyzing user data to generate a profile; means for customizing smart speaker functionality based on the generated profile; and a means for a user to input voice commands; A means for recognizing an input voice command and converting it into text data; means for analyzing the converted text data to generate an appropriate response; A means to output the generated response as speech A system including:

2. a means of replacing the communications equipment with new communications equipment; A means for acquiring initial setting information and performing user authentication; and further including means for activating a smart speaker function after successful authentication. The system of claim 1 .

3. a means for obtaining user feedback; and further comprising means for analyzing the obtained feedback and incorporating it into future responses. The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A